IMAGE PROCESSING METHOD, APPARATUS, DEVICE, AND COMPUTER PROGRAM
An image processing method using a GAN model effectively addresses the challenge of high-accuracy and realistic avatar transformation by training the model to compute and retain attribute differences between original and target avatars.
Patent Information
- Application Number
- JP2023576240
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-03
- Filing Date
- 2021-07-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-07-26
AI Technical Summary
Existing image processing technologies struggle to achieve high accuracy and realism in avatar transformation, where an original avatar in an image is replaced with a target avatar while retaining the original avatar's attributes.
The development of an image processing method that trains a generative adversarial network (GAN) model by computing differences between target avatars and predicted avatars, allowing the model to replace avatars in input images with target avatars while retaining specific attributes.
The trained image processing model achieves high processing accuracy and realism in avatar replacement, enabling effective transformation across various scenarios while maintaining the original attributes.
Smart Images

Figure 0007689592000039 
Figure 0007689592000040 
Figure 0007689592000041
Abstract
Description
[Technical field]
[0001] The present application relates to the field of computer technology, and in particular to an image processing method, apparatus, device, and computer-readable storage medium.
[0002] This application claims priority to a Chinese patent application filed on June 3, 2021, bearing application number 2021106203828 and entitled "Image Processing Method, Apparatus, Equipment, and Computer-Readable Storage Medium," the entire contents of which are incorporated herein by reference. [Background technology]
[0003] With the continuous development of computer technology, image processing technology has been widely developed. Here, using image processing technology to realize avatar transformation is a relatively new attempt and application, where avatar transformation refers to the process of replacing an original avatar in an image with a target avatar. Summary of the Invention [Problem to be solved by the invention]
[0004] The embodiments of the present application provide an image processing method, apparatus, device, and computer-readable storage medium, which enable training of an image processing model, and the image processing model obtained by training has relatively high processing accuracy, the realism of the avatar obtained by replacement is relatively high, and the image processing model can be used in a wide range of scenarios. [Means for solving the problem]
[0005] In one aspect, embodiments of the present application provide a method for image processing, comprising: Invoke the first generative network in the image processing model and compute the first sample image x in the first sample set. ito obtain a first predicted image [Equation 1], the first predicted image [Equation 1] including a first predicted avatar, the first sample set including N first sample images, each of the first sample images including a target avatar corresponding to a same target person, N being a positive integer, i being a positive integer and i≦N; Invoke the first generative network and generate a second sample image y k to obtain a second predicted image [Equation 2], the second predicted image [Equation 2] including a second predicted avatar, the second sample set including M second sample images, each second sample image including a sample avatar, M being a positive integer, k being a positive integer and k≦M; The first sample image x i the difference between the target avatar and the first predicted avatar in the second sample image y k and training the image processing model based on a difference between a first type of attributes of the sample avatar in the input image and the first type of attributes of the second predicted avatar, wherein the image processing model is used to replace an avatar in an input image with the target avatar while retaining the first type of attributes of the avatar in the input image.
[0006]
number
number
[0007] In one aspect, embodiments of the present application provide an image processing apparatus, the image processing apparatus comprising: Invoke the first generative network in the image processing model and compute the first sample image x in the first sample set. ia first predicted image acquisition module for processing the first predicted image to obtain a first predicted image, the first predicted image including a first predicted avatar, the first sample set including N first sample images, each of the first sample images including a target avatar corresponding to a same target person, N being a positive integer, i being a positive integer, and i≦N; Invoke the first generative network and generate a second sample image y k a second predicted image acquisition module for processing the second predicted image to obtain a second predicted image, the second predicted image including a second predicted avatar, the second sample set including M second sample images, each of the second sample images including a sample avatar, M being a positive integer, k being a positive integer, and k≦M; The first sample image x i the difference between the target avatar and the first predicted avatar in the second sample image y k and a model training module used to train the image processing model based on a difference between a first type of attributes of the sample avatar in the input image and the first type of attributes of the second predicted avatar, wherein the image processing model is used to replace an avatar in an input image with the target avatar and retain the first type of attributes of the avatar in the input image.
[0008] In one aspect, the present application provides an image processing device, comprising: a storage device; and a processor; The storage device stores a computer program, The processor is used to load and execute the computer program to realize the image processing method.
[0009] In one aspect, the present application provides a computer readable storage medium having a computer program stored thereon, the computer program being suitable to be loaded by a processor and to execute the image processing method.
[0010] In one aspect, the present application provides a computer program product or a computer program comprising computer instructions stored in a computer readable storage medium, the computer instructions being read by a processor of a computing device from the computer readable storage medium and executed by the processor to cause the computing device to perform the image processing method provided in said various selectable implementations. Effect of the Invention
[0011] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0012] By training an image processing model based on the avatar difference between a first sample image and a first predicted image obtained by processing it with the first generation network, and the difference in a first type of attribute between a second sample image and a second predicted image obtained by processing it with the first generation network, the image processing model after training has learned the ability to replace an avatar in an input image with a target avatar and retain the first type of attributes of the avatar in the input image, thereby relatively enhancing the realism of the avatar after replacement, and improving the accuracy of image processing and the replacement effect.
[0013] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory and are not intended to limit the present application.
[0014] In order to more clearly describe the technical solutions in the embodiments of the present application, the following briefly introduces the drawings that need to be used in the description of the embodiments. It is obvious that the drawings in the following description are only some embodiments of the present application, and those skilled in the art can further obtain other drawings according to these drawings on the premise that no creative labor is required. [Brief description of the drawings]
[0015] [Figure 1] 1 shows a diagram of an image processing scene provided by one exemplary embodiment of the present application; [Diagram 2] 1 shows a flowchart of an image processing method provided by one exemplary embodiment of the present application. [Diagram 3] 4 shows a flowchart of another image processing method provided by an exemplary embodiment of the present application. [Figure 4] 1 shows a flowchart of generating a training sample set provided by one exemplary embodiment of the present application. [Diagram 5] 1 shows a structural schematic diagram of an encoder provided according to an exemplary embodiment of the present application; [Figure 6] 1 shows a structural schematic diagram of a first decoder provided according to an exemplary embodiment of the present application; [Figure 7] 4 illustrates a structural schematic diagram of a second decoder provided according to an exemplary embodiment of the present application; [Figure 8] 1 illustrates a schematic diagram of a process of determining a first loss function provided by an exemplary embodiment of the present application; [Figure 9] 1 illustrates a schematic diagram of a process of determining a first loss function provided by an exemplary embodiment of the present application; [Figure 10] 1 illustrates a schematic diagram of a process of determining a first loss function provided by an exemplary embodiment of the present application; [Figure 11] 1 shows a flowchart of a test video process provided by one exemplary embodiment of the present application. [Figure 12]1 illustrates a flowchart of a test image generation process by a first generation network provided by an exemplary embodiment of the present application. [Figure 13] 1 shows a flowchart of a test video process provided by one exemplary embodiment of the present application. [Figure 14] 1 shows a structural schematic diagram of an image processing device provided according to an exemplary embodiment of the present application; [Figure 15] 1 illustrates a structural schematic diagram of an image processing device provided by an exemplary embodiment of the present application; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Exemplary embodiments will now be described in detail, examples of which are illustrated in the drawings. When the following description refers to the drawings, the same numerals in different drawings refer to the same or similar elements unless otherwise noted. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of methods consistent with some aspects of the present application, as detailed in the appended claims.
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present application in combination with the drawings in the embodiments of the present application.
[0018] The embodiments of the present application relate to technologies such as artificial intelligence (AI) and machine learning (ML).
[0019] In addition, the embodiment of the present application further relates to an avatar transformation and a generator, where the avatar transformation refers to a process of replacing all or part of the avatar features in a first person image with a second person image, and in the embodiment of the present application, the second person image is input into an image processing model to obtain a predicted person image output from the image processing model. The predicted person image not only has the avatar of the first person image, but also has a first type of attribute of the second person image. Optionally, the predicted person image having the avatar of the first person image means that the predicted person image not only has the first person image avatar, such as facial features such as the five senses, hair, skin, glasses, etc., but also has a first type of attribute of the second person image, such as attribute features such as posture, facial expression, lighting, etc.
[0020] The generative network is a component of a generative adversarial network (GAN), which is a method of unsupervised learning and is composed of one generator network (Generator) and one discriminator network (Discriminator). The input of the discriminator network is a true sample image (i.e., a real image, where the real image refers to a non-model generated image) or a predicted image output from the generative network (i.e., a fake image, where the fake image refers to an image generated based on a model). The purpose of the discriminator network is to discriminate the veracity of the predicted image output from the generative network and the true sample image as much as possible, i.e., to be able to distinguish which is a real image and which is a predicted image. Meanwhile, the generative network makes the generated predicted image as unlikely to be identified by the discriminator network as much as possible, i.e., to make the predicted image as close to reality as possible. The two networks continue to compete against each other, adjusting their parameters (i.e., optimizing each other), until eventually the predicted images generated by the generative network are less likely to be judged as false by the discriminatory network, or the discrimination accuracy of the discriminatory network reaches a threshold.
[0021] Based on computer vision technology and machine learning technology in AI technology, an embodiment of the present application provides an image processing method, and trains an image processing model based on a generative adversarial network, so that the trained image processing model can transform any avatar into a target avatar, and the replaced avatar can retain a first type of attributes of any avatar (i.e., realize avatar transformation).
[0022] FIG. 1 shows a diagram of an image processing scene provided by one exemplary embodiment of the present application. As shown in FIG. 1, the image processing scene includes a terminal device 101 and a server 102. Here, the terminal device 101 is a device used by a user, and the terminal device may further have an image collection function or an interface display function, and the terminal device 101 may include, but is not limited to, a smartphone (such as an Android mobile phone, an iOS mobile phone, etc.), a tablet computer, a portable personal computer, a mobile Internet device (Mobile Internet Devices, MID), and other devices. A display device is disposed in the terminal device, and the display device may further be a display, a display screen, a touch screen, etc., and the touch screen may further be a touch control screen, a touch control panel, etc., and is not limited in the embodiment of the present application.
[0023] The server 102 refers to a background device that can train an image processing model according to the acquired samples. After obtaining the trained image processing model, the server 102 may return the trained image processing model to the terminal device 101, or may deploy the trained image processing model in the server. The server 102 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), big data, and artificial intelligence platforms. In addition, multiple servers may be combined to form a blockchain network, and each server is a node in the blockchain network and jointly trains, stores, or delivers the image processing model. The terminal device 101 and the server 102 may be directly or indirectly connected by wired communication or wireless communication, and the present application does not limit this.
[0024] It is necessary to explain that the number of terminal devices and servers in the model processing scene shown in FIG. 1 is only an example, for example, the number of terminal devices and servers may be multiple, and the present application does not limit the number of terminal devices and servers.
[0025] In the image processing scene shown in FIG. 1, the image processing process mainly includes the following steps:
[0026] (1) The server acquires a training sample set of the image processing model, and the training sample set may be acquired from a terminal device or may be acquired from a database. The training sample set includes a first sample set and a second sample set. The first sample set includes N first sample images, each of which includes a target avatar of a same target person, and the second sample set includes M second sample images, each of which includes one sample avatar, where M and N are both positive integers, and the image processing model includes a first generation network.
[0027] Here, the sample avatar included in the second sample image refers to an avatar of a person other than the target person. As an example, the first sample set includes N first sample images, each of which includes an avatar of user A, and the second sample set includes M second sample images, each of which includes a sample avatar of a person other than one user A. Optionally, the sample avatars in the M second sample images correspond to different people. The first generative network in the trained image processing model can generate a predicted image corresponding to user A, which has an avatar of user A in the predicted image, while retaining a first type of attribute of the avatar in the original input image.
[0028] (2) The server receives the first sample image x from the first sample set. i Then, we call the first generator network to generate the first sample image x i (the image is a real image) to obtain a first predicted image [Equation 1] (the image is a fake image), where i is a positive integer and i≦N, and a first generating network is invoked to generate a first sample image x i Processing refers to processing a first sample image x by an encoder in a first generating network. iand then decoding the first feature vector by a decoder in the first generation network to obtain a first predicted image [Equation 1].
[0029] (3) The server receives the second sample image y from the second sample set. k and call the first generator network to generate the second sample image y k (the image is a real image) to obtain a second predicted image [Equation 2] (the image is a fake image), where k is a positive integer and k≦M, and the first generating network is called to obtain a second sample image y k To perform a generating process on a second sample image y k and then decoding the second feature vector by a decoder in the first generation network to obtain a second predicted image [Equation 2].
[0030] (4) The server receives the first sample image x i the difference between the target avatar in the first predicted image (i.e., the first predicted avatar) and the avatar in the second sample image y k The image processing model is trained based on the difference between the first kind of attribute of the sample avatar in the second predicted image [Equation 2] and the first kind of attribute of the avatar (second predicted avatar) in the second predicted image [Equation 2].
[0031] In one embodiment, the server can calculate the difference between each image according to the constructed loss function. For example, the server can calculate the difference between a first sample image x i The difference between the target avatar and the first predicted avatar in kand determining a difference between a first type of attribute of the sample avatar in the second predicted image [Equation 2] and a first type of attribute in the second predicted image [Equation 2]; determining a function value of a target loss function based on the function value of the first loss function and the function value of the second loss function; and updating parameters of the image processing model based on the function value of the target loss function, thereby training the image processing model.
[0032] (5) Repeat steps (2) to (4) until the image processing model reaches a training completion condition. The training completion condition is that the loss value of the target loss function no longer decreases with the number of iterations, or the number of iterations reaches a number threshold, or the loss value of the target loss function is less than a loss threshold, etc. Here, the image processing model after completing training can replace the avatar in the input image with the target avatar and retain the first type of attribute of the image in the input avatar.
[0033] Based on the above description, the image processing method proposed in the embodiments of the present application will be introduced in detail below in conjunction with the drawings.
[0034] Taking one iteration of the image processing model as an example, Fig. 2 shows a flow chart of an image processing method provided by one exemplary embodiment of the present application, which can be executed by the server 102 shown in Fig. 1, and as shown in Fig. 2, the image processing method includes but is not limited to the following steps:
[0035] S201: Invoke a first generator network in the image processing model, and generate a first sample image x in a first sample set. i to obtain a first predicted image [Equation 1], which includes a first predicted avatar, and the first sample set includes N first sample images, each of which includes a target avatar corresponding to a same target person, where N is a positive integer, i is a positive integer, and i≦N.
[0036] Generally speaking, all of the N images in the first sample set include an avatar of user A, and the N first sample images may be images of user A in different scenes, for example, user A in different images in the first sample set may have different poses, such as raising his head, lowering his head, different facial expressions, such as laughing, crying, etc.
[0037] S202: Invoke the first generator network to generate a second sample image y in the second sample set. k to obtain a second predicted image [Equation 2], the second predicted image [Equation 2] including a second predicted avatar, the second sample set including M second sample images, each second sample image including a sample avatar, M being a positive integer, k being a positive integer, and k≦M.
[0038] Here, the M images in the second sample set include avatars of users other than user A. In general, the avatars in the second sample set may be images including avatars of any one or more users other than user A, and the second sample set may include images corresponding to the avatars of multiple different users in order to avoid overfitting in the training process. Optionally, the avatars in the second sample set may include images of any one or more users other than user A in different scenes.
[0039] S203: First sample image x i the difference between the target avatar and the first predicted avatar in the second sample image y k An image processing model is trained based on the difference between the first type of attributes of the sample avatar in and the first type of attributes of the second predicted avatar.
[0040] The server judges whether the updated image processing model reaches the training completion condition. If the updated image processing model does not reach the training completion condition, the above steps S201 to S203 are repeated until the image processing model reaches the training completion condition, and if the updated image processing model reaches the training completion condition, the training of the image processing model is terminated. The image processing model is used to replace the avatar in the input image with a target avatar and retain the first type of attribute of the avatar in the input image.
[0041] Here, replacing the avatar in the input image with the target avatar may refer to replacing a second type of attribute in the input image with a second type of attribute in the target avatar, the second type of attribute being an attribute other than the first type of attribute, that is, the replaced image not only retains the second type of attribute of the target avatar, but also retains the first type of attribute of the avatar in the input image.
[0042] After training, the image processing model can meet the avatar transformation needs of 1vN. The first generating network can realize the replacement of the target avatar of the target person into the person images of any other person in different scenes, that is, the first generating network can process the input image (the image containing the avatar of any other person) to obtain an image having a first kind of attribute of the input image and a second kind of attribute of the target avatar of the target person.
[0043] In the embodiment of the present application, the target user corresponding to the N images in the first sample set is a user who provides a second type of attribute of the avatar. After the training of the image processing model is completed, the avatar in the obtained predicted image not only has the second type of attribute of the avatar of the target user after the input image is processed by the first generation network in the trained image processing model, but also has the first type of attribute of the input image, in general, the user corresponding to the N images in the first sample set is user A, and when any image is processed by the first generation network in the image processing model, an avatar including the second type of attribute of the avatar of user A can be obtained.
[0044] The means shown in the embodiments of the present application can replace the avatar of any person with the avatar of a target person, thereby realizing the application of 1vN avatar replacement and expanding the application scenarios of avatar conversion.
[0045] Optionally, the first type of attribute is a non-identifying attribute, and the second type of attribute is an identifying attribute. In general, the identifying attribute can include facial features such as the senses, skin, hair, glasses, etc., and the non-identifying attribute can include characteristic attributes such as facial expression, posture, lighting, etc. Figure 3 shows a flowchart of another image processing method provided by one exemplary embodiment of the present application. The image processing method can be performed by the server 102 shown in Figure 1, and as shown in Figure 3, the image processing method includes but is not limited to the following steps:
[0046] S301: Obtain a training sample set for an image processing model.
[0047] In one possible implementation, in order to improve the stability and accuracy of model training, the images in the training sample set are images after pre-processing. The pre-processing includes at least one of the processes of avatar calibration, avatar region segmentation, etc., and generally, Fig. 4 shows a flowchart of the generation of the training sample set provided by one exemplary embodiment of the present application. As shown in Fig. 4, the process of obtaining the training sample set of the image processing model mainly includes steps S3011-S3013.
[0048] S3011: Data collection stage.
[0049] The server obtains a first original sample set and a second original sample set. Each image in the first original sample set includes a target avatar of a target person, and each image in the second original sample set includes a sample avatar of a person other than the target person. Optionally, the server can obtain images including corresponding avatars by performing image frame extraction on a video. Take the process of obtaining the first original sample set as an example, the server can obtain images including target avatars from a video corresponding to a target person. The video corresponding to the target person may be a video uploaded by a terminal device, and the video has a certain duration, and the server can extract X images including target avatars of the target person from the video by image frame extraction to obtain the first original sample set, and can further obtain a second original sample set according to the same manner. Alternatively, optionally, images corresponding to each user are pre-stored in the database, and similarly, the server obtaining the first original sample set and the second original sample set from the database includes: the server can obtain X images including a target avatar of a target person based on identity information corresponding to the target person to form the first original sample set, and obtain at least two images including sample avatars of other persons based on identity information other than the target person to form the second original sample set, where the different images in the second original sample set correspond to the same or different persons.
[0050] Optionally, the above two original sample set acquisition methods can be used in combination. In general, the server may acquire a first original sample set by performing image frame extraction on the video, and acquire a second original sample set from the database; or the server may acquire a second original sample set by performing image frame extraction on the video, and acquire the first original sample set from the database; or the server may acquire a part of the images in the first original sample set / second original sample set by performing image frame extraction on the video, and acquire another part of the images in the first original sample set / second original sample set from the database.
[0051] S3012: Avatar calibration stage. The server can perform avatar region detection on the images in the collected original sample set (including the first original sample set and the second original sample set) by a face detection algorithm (e.g., AdaBoost framework, Deformable Part Model (DMP) model, Cascade CNN, etc.), and calibrate the avatar region (e.g., adopt a face alignment algorithm based on regression tree), in order to locate the accurate shape of the avatar on the known avatar region. The server can also correct the avatar by an avatar pose correction algorithm such as a 3D Morphable Models (3DMM) algorithm to obtain a corrected avatar, that is, obtain a frontal avatar corresponding to the original avatar, and train an image processing model with the corrected avatar, which is favorable to improving the stability of model training. Optionally, the original sample image set after avatar calibration can be obtained as a training sample set, or the server can further perform avatar region segmentation on the original sample image set after avatar calibration to obtain a sample image set after avatar region segmentation as a training sample set. After obtaining the corrected training sample set, the server can directly start the training process of the image processing model, i.e., perform step S302, or continuously perform step S3013 (this step is an optional item).
[0052] S3013: Avatar region segmentation stage. In step S3012, the server has already determined the avatar region of each image in the original sample image set, so the server can trim each image in the original sample image set after correction and retain only the avatar region in each image. That is, before training the image processing model, avatar correction and avatar segmentation are performed in advance on the original sample image set. Compared with training the image processing model using the original sample image set as it is, training the image processing model using the sample image that retains only the avatar region can improve the training efficiency of the image processing model.
[0053] S302: Invoke a first generative network in the image processing model, and generate a first sample image x in the first sample set. i is processed to obtain a first predicted image [Equation 1].
[0054] The first sample image x i is an arbitrary first sample image in the first sample set.
[0055] In one embodiment, the first generating network includes an encoder and a first decoder, where the encoder is used for extracting image features to obtain a feature vector corresponding to a sample image, and the first decoder is used for generating a predicted image according to the feature vector.
[0056] In general, the process of obtaining the first predicted image includes: i After obtaining the first sample image x i Encode the first sample image x iand obtain a first feature vector corresponding to the first feature vector, and after obtaining the first feature vector, the server calls a first decoder to decode the first feature vector to obtain a first generated image and first region division information. The first region division information is used to indicate an avatar region in the first generated image, and the server further extracts a first predicted image [Equation 1] from the first generated image according to the first region division information.
[0057] 5 shows a structural schematic diagram of an encoder provided by one exemplary embodiment of the present application. As shown in FIG. 5, the encoder includes P feature extraction networks and one feature aggregation layer, where P is a positive integer, and each feature extraction network includes one downsampling layer. Therefore, there are P downsampling layers in one encoder, and the scale parameters of the P downsampling layers are different. For example, the dilation (scale parameter) in the first downsampling layer is 1, the dilation (scale parameter) in the second downsampling layer is 2, and the dilation (scale parameter) in the third downsampling layer is 4. Furthermore, each downsampling layer is constructed based on a Depth Separable Convolution Network (DSN), which includes a convolution function (Conv2d, k=1) and a depth convolution function (DepthConv2d, k=3, s=2, dilation=d). Based on this, the server calls the encoder to generate a first sample image x i to obtain a first feature vector, The first sample image x is obtained by downsampling layers in the P feature extraction networks (i.e., P downsampling layers). i Extract feature information under P different scale parameters of The feature aggregation layer extracts the first sample image x i The feature information under the P scale parameters is aggregated to obtain the first sample image x iThe first step is to obtain a first feature vector corresponding to
[0058] In one possible implementation, the first decoder includes a first feature transformation network, Q first image reconstruction networks, and a first convolution network, where Q is a positive integer, and each first image reconstruction network includes a first residual network and a first upsampling layer; The first feature transformation network is used to transform the feature vector input to the first decoder into a feature map; The Q first image reconstruction networks are used to perform a first feature restoration process on the feature map to obtain a fusion feature image.
[0059] The first convolution network is used to perform convolution processing on the fusion feature image, and output a generated image corresponding to the feature vector input to the first decoder. Schematically, FIG. 6 illustrates a structural schematic diagram of a first decoder provided by an exemplary embodiment of the present application. As shown in FIG. 6, the first decoder includes a first feature transformation network, Q first image reconstruction networks, and a first convolution network, where Q is a positive integer. Each of the first image reconstruction networks includes a first residual network and a first upsampling layer. Here, the first feature transformation network is used to reshape the feature vector input to the first decoder into a feature map. The first image reconstruction network performs a first feature restoration on the feature map, i.e., reshapes the size of the feature map into a first sample image x by Q first upsampling layers (Up Scale Block). i, and a first residual network is used to reduce the gradient vanishing problem that exists in the upsampling process, thereby obtaining a fusion feature image corresponding to the first sample image. The first convolution network is used to perform a convolution process on the fusion feature image corresponding to the first sample image, thereby obtaining a generated image corresponding to the feature vector input to the first decoder. Based on this, an embodiment in which the server calls the first decoder to decode the first feature vector and obtain the first generated image and the first region division information is as follows: Transform the first feature vector through a feature transformation network to obtain a first feature map; Perform a first feature reconstruction on the feature map using Q first image reconstruction networks to obtain a first sample image and a corresponding fusion feature image; A first convolutional network may perform a convolution process on the fusion feature image corresponding to the first sample image to obtain a first generated image corresponding to the first sample image, and first region division information.
[0060] S303: Invoke the first generator network and generate a second sample image y in the second sample set. k to obtain a second predicted image [Equation 2], which includes a second predicted avatar.
[0061] Similar to the process of the first generating network encoding the first sample image, the process of obtaining the second predicted image can be realized as follows: After the encoder in the first generating network obtains the second sample image, the encoder encodes the second sample image to obtain a second feature vector corresponding to the second sample image, and then the server calls the first decoder to decode the second feature vector to obtain a second generated image and second region division information. The second region division information is used to indicate the avatar region in the second generated image. Furthermore, the server extracts a second predicted image [Equation 2] from the second generated image according to the second region division information.
[0062] Now the server calls the encoder to generate the second sample image y k to obtain the second feature vector, The second sample image y k Extract feature information under P scale parameters of The feature aggregation layer extracts the second sample image y k Alternatively, an aggregation process may be performed on the feature information under the P scale parameters to obtain a second feature vector.
[0063] Similarly to the process in which the first generating network decodes the first feature vector, the embodiment in which the server calls the first decoder to decode the second feature vector to obtain the second generated image and the second region segmentation information can be implemented as follows: Transform the second feature vector through a feature transformation network to obtain a second feature map; Perform a first feature restoration process on the second feature map by Q first image reconstruction networks to obtain a fusion feature image corresponding to the second sample image; A first convolutional network may perform a convolution process on the fusion feature image corresponding to the second sample image to obtain a second generated image corresponding to the second sample image, and second region division information.
[0064] S304: Invoke a second generating network and generate a second sample image y k to obtain a third predicted image [Equation 3], which includes a third predicted avatar, and the second generating network has the same feature extraction unit as the first generating network.
[0065] Here, the same feature extraction unit in the second generating network and the first generating network is an encoder, and both have the same structure and parameters, and the second generating network is used to assist in training the first generating network, and further, the second generating network is used to assist in training the encoder in the first generating network.
[0066]
number
[0067] In one possible implementation, the second generating network includes an encoder, a second decoder, and an identity identification network.
[0068] Here, the encoder (i.e., the first generator network and the second generator network have the same feature extraction structure) is used to extract image features in the sample image to obtain a feature vector. The identity identification network is used to obtain image identification information based on the sample image or the feature vector corresponding to the sample image, which can be different IDs (identifications), encoding information, etc. corresponding to different images, and the second decoder is used to generate a predicted image according to the feature vector obtained by the encoder and the identification information provided by the identity identification network. In the process of generating the third predicted image [Equation 3], the server calls the encoder to extract the second sample image y k Encode the second sample image y k and obtain the second feature vector corresponding to the second sample image y k At the same time as encoding the second sample image y k After encoding, the server calls the identity network and receives the second sample image y k , or the second sample image y k and performing identity recognition based on the corresponding second feature vector to obtain the second sample image y k and the corresponding identification information (e.g., the second sample image yk The server then calls a second decoder to obtain a second sample image y k The second feature vector is decoded according to the identification information to obtain a third generated image and third region division information, which is used to indicate an avatar region in the third generated image, and the server further extracts a third predicted image [Equation 3] from the third generated image according to the third region division information.
[0069] Now the server calls the encoder to generate the second sample image y k to obtain the second feature vector, The second sample image y k Extract feature information under P scale parameters of The feature aggregation layer extracts the second sample image y k Alternatively, an aggregation process may be performed on the feature information under the P scale parameters to obtain a second feature vector.
[0070] 7 shows a structural schematic diagram of a second decoder provided by an exemplary embodiment of the present application. As shown in FIG. 7, the second decoder includes one second feature transformation network, Q second image reconstruction networks (corresponding to the quantity of the first image reconstruction networks in the first decoder), and one second convolution network, where Q is a positive integer, and each second image reconstruction network includes one second residual network, one second upsampling layer, and one self-adaptive module (AdaIN). Here, the second feature transformation network is used to transform the feature vector input to the second decoder into a feature map, and the second image reconstruction network performs a second feature restoration process on the feature map, i.e., the size of the feature map is increased by Q second upsampling layers to generate a second sample image y kThe self-adaptation module is used to restore the size of the second sample image y to match that of the second sample image y in the upsampling process by adding discrimination information corresponding to the feature vector input to the second decoder, so that the second decoder performs feature fusion based on the discrimination information, and uses the second residual network to alleviate the gradient vanishing problem that exists in the upsampling process, to obtain a fusion feature image corresponding to the second sample image. In other words, the self-adaptation module is used to restore the size of the second sample image y k The second convolutional network is used to perform a convolution process on the fusion feature image corresponding to the second sample image to obtain a generated image corresponding to the feature vector input to the second decoder according to the identification information of the second sample image, and the third feature vector is used to instruct the second decoder to decode the feature vector input to the second decoder. Based on this, an embodiment in which the server calls the second decoder to decode the second feature vector according to the identification information of the second sample image to obtain a third generated image and third region division information is as follows: Transform the second feature vector through a feature transformation network to obtain a second feature map; A second feature restoration process is performed on the second feature map by Q second image reconstruction networks to obtain a second sample image and a corresponding fusion feature image; A second convolutional network may perform a convolution process on the fusion feature image corresponding to the second sample image to obtain a third generated image corresponding to the second sample image, and third region division information.
[0071] S305: Determine a function value of a first loss function, the first loss function being a function value of a first sample image x i is used to indicate the difference between the target avatar and the first predicted avatar in
[0072] In one possible implementation, the process of determining the function value of the first loss function can be implemented as follows:
[0073] Call the first discrimination device and generate the first sample image x i , and a first predicted avatar are determined, respectively; 1st sample image x i and a second discrimination result of the first predicted avatar, a function value of a first branch function of the first loss function is determined based on the first discrimination result of the first sample image x i is used to indicate whether the first predicted avatar is a real image, and the second discrimination result is used to indicate whether the first predicted avatar is a real image; A function value of a second branch function of the first loss function is determined, the second branch function of the first loss function being a function value of the first sample image x i and indicating a difference between a first perceptual feature of the target avatar and a second perceptual feature of the first predicted avatar; The sum of the function value of the first branch function of the first loss function and the function value of the second branch function of the first loss function is determined as the function value of the first loss function.
[0074] FIG. 6 shows a schematic diagram of a process of determining a first loss function provided by an exemplary embodiment of the present application. As shown in FIG. 6, the server realizes the determination of the function value of the first loss function by calling a first discrimination device (a specific discrimination device) and a feature perception network. In general, the feature perception network can be realized as a Learned Perceptual Image Patch Similarity (LPIPS) network. When an image processing model is applied to a first sample image x i and generate a first sample image x by the first generating network. i For an embodiment of processing the first predicted image [Equation 1] to obtain the first predicted image [Equation 1], reference may be made to step S302, which will not be described in detail here.
[0075] After obtaining the first predicted image [Equation 1], the server, on the other hand, uses a first discrimination device to obtain a first sample image x i , and the first predicted image [Equation 1], i.e., the first sample image x i, and determine whether the first predicted image [Equation 1] is an actual image, and determine whether the first sample image x i Based on the first discrimination result of the first predicted image [Equation 1] and the second discrimination result of the first predicted image [Equation 1], a function value of a first branch function of the first loss function is determined, and the first branch function can be expressed as the following [Equation 4].
[0076]
number
[0077] Here, L GAN1 represents the first branch function of the first loss function of the first generative adversarial network GAN (including the first generative network (G1) and the first discriminator (D1)). [Equation 5] is the first predicted image [Equation 1] generated by the first generative network (G) and the first sample image x i The formula 6 represents the difference between the discrimination result of the first predicted image [Formula 1] by the first discrimination device and the first sample image x i The E(x) function is used to calculate the expected value of x, and D src (x) is used to represent the discrimination of x by using the first discrimination device, and I src is the first sample image, i.e., x i where Enc(x) is used to denote the encoding of x using an encoder, and Dec src (x) is used to denote the decoding of x by using the first decoder. From this, it can be deduced that D src (I src ) is a first discrimination device to generate a first sample image x i where [Equation 7] is the first predicted image, i.e., [Equation 1], and [Equation 8] is the first discriminator used to discriminate the first predicted image, i.e., [Equation 1].
[0078]
number
number
number
number
[0079] On the other hand, the server generates the first sample image x by the feature perception network (LPIPS network). i , and the first predicted image [Equation 1] are subjected to feature recognition to obtain the first sample image x i and a second perceptual feature corresponding to the first predicted image [Equation 1], and perform feature comparison between the first perceptual feature and the second perceptual feature to obtain a first feature comparison result. The first feature comparison result is i and the first predicted image [Equation 1]. After obtaining the first feature comparison result, the server determines a function value of a second branch function of the first loss function according to the first feature comparison result, and the second branch function can be expressed as follows [Equation 9]:
[0080]
number
[0081] Here, L LPIPS1 represents the second branch function of the first loss function corresponding to the feature perception network (LPIPS network), and LPIPS(x) represents feature perception for x by the feature perception network (LPIPS network). As is clear from the first branch function of the first loss function, I src is the first sample image, i.e., x i [Equation 10] is the first predicted image, i.e., [Equation 1]. Based on this, [Equation 11] represents feature perception of the first predicted image, i.e., [Equation 1], by the LPIPS network. LPIPS(I src ) is the first sample image x iIt represents feature perception.
[0082]
number
number
[0083] After obtaining the function value of the first branch function of the first loss function and the function value of the second branch function of the first loss function, the server determines the sum of the function value of the first branch function of the first loss function and the function value of the second branch function of the first loss function as the function value of the first loss function, and calculates the first loss function L 1 can be expressed as follows:
[0084]
number
[0085] S306: Determine a function value of a second loss function, the second loss function being a function value of the second sample image y k The second predicted avatar may be used to indicate a difference between the first type of attribute of the sample avatar and the first type of attribute of the second predicted avatar.
[0086] In one possible implementation, the process of determining the function value of the second loss function can be implemented as follows:
[0087] Call the first discrimination device to discriminate the second predicted avatar; determining a function value of a first branch function of the second loss function based on a third discrimination result of the second predicted avatar, the third discrimination result being used to indicate whether the second predicted avatar is a real image; Second sample image y k Perform attribute comparison between the first type of attribute of the sample avatar in and the first type of attribute of the second predicted avatar to obtain an attribute comparison result; determining a function value of a second split function of the second loss function based on the attribute comparison result; The sum of the function value of the first branch function of the second loss function and the function value of the second branch function of the second loss function is determined as the function value of the second loss function.
[0088] FIG. 9 shows a schematic diagram of a process of determining a second loss function provided by an exemplary embodiment of the present application. As shown in FIG. 9, the server realizes the determination of the function value of the second loss function by calling an attribute identification network and a first discrimination device. The attribute identification network can identify facial expression attributes used to indicate facial expressions, such as eye size, eyeball position, and mouth size, and the attribute identification network outputs a continuous value within a [0,1] range, for example, for eye size, 0 represents the eyes are closed, and 1 represents the eyes are fully open. For eye position, 0 represents the leftmost bias, and 1 represents the rightmost bias. For mouth size, 0 represents the mouth is closed, and 1 represents the mouth is fully open. When the image processing model processes a second sample image y k and the first generating network generates a second sample image y k For an embodiment of processing the above to obtain the second predicted image [Equation 2], reference may be made to step S303, which will not be described in detail here.
[0089] After obtaining the second predicted image [Equation 2], on the other hand, the server discriminates the second predicted image [Equation 2] by a first discrimination device, i.e., determines whether the second predicted image [Equation 2] is an actual image, and determines a function value of a first branch function of the second loss function based on the third discrimination result of the second predicted image [Equation 2], where the first branch function can be expressed as follows [Equation 13].
[0090]
number
[0091] Here, L GAN2represents the first branch function of the second loss function of the first generative adversarial network GAN (including the first generative network (G1) and the first discriminator (D1)). [Equation 14] represents the second predicted image [Equation 2] generated by the first generative network (G) and the second sample image y k The difference between the first prediction image and the second prediction image [Equation 2] by the first discrimination device is made as small as possible (to make the discrimination result of the second prediction image by the first discrimination device true). The E(x) function is used to calculate the expected value of x, and D src (x) is used to represent the discrimination of x by using the first discrimination device, and I other is the second sample image y k where Enc(x) is used to denote the encoding of x using an encoder, and Dec src (x) is used to represent the decoding of x by using the first decoder. As can be deduced from this, [Equation 15] is the second predicted image [Equation 2], and [Equation 16] represents the discrimination of the second predicted image [Equation 2] by using the first discrimination device.
[0092]
number
number
number
[0093] On the other hand, the server uses the attribute classification network to find the second sample image y k , and the second predicted image [Equation 2] to obtain a first type of attribute corresponding to the second sample image and a first type of attribute of the first predicted image, and perform attribute comparison between the first type of attribute of the second sample image and the first type of attribute of the first predicted image to obtain an attribute comparison result. k and the second predicted image [Equation 2], and determines the function value of the second loss function based on the attribute comparison result. The second loss function can be expressed as follows [Equation 17]:
[0094]
number
[0095] Here, L attri represents the loss function of the attribute classification network, and N attri (x) represents attribute classification for x using the attribute classification network, and as is clear from the first branch function of the second loss function, I other is the second sample image y k [Equation 18] is the second predicted image [Equation 2], and based on this, [Equation 19] represents the attribute feature extraction for the second predicted image [Equation 2] by the attribute classification network, and [Equation 20] represents the extraction of the second sample image y k This represents the extraction of attribute features for
[0096]
number
number
number
[0097] After obtaining the function value of the first branch function of the second loss function and the function value of the second branch function of the second loss function, the server determines the sum of the function value of the first branch function of the second loss function and the function value of the second branch function of the second loss function as the function value of the second loss function, and calculates the second loss function L 2 can be expressed as follows:
[0098]
number
[0099] S307: Determine a function value of a third loss function, the third loss function being a function value of the second sample image y kis used to indicate the difference between the sample avatar and the third predicted avatar in
[0100] In one possible implementation, the process of determining the function value of the third loss function can be implemented as follows:
[0101] Call the second discrimination device to obtain the second sample image y k , and a third predicted avatar, respectively; Second sample image y k and a fifth discrimination result of the third predicted avatar, a function value of a first branch function of the third loss function is determined based on the fourth discrimination result of the second sample image y k is used to indicate whether the third predicted avatar is a real image; and the fifth discrimination result is used to indicate whether the third predicted avatar is a real image; A function value of a second branch function of the third loss function is determined, the second branch function of the third loss function being k and indicating a difference between a third perceptual feature of the sample avatar and a fourth perceptual feature of the third predicted avatar; The sum of the function value of the first branch function of the third loss function and the function value of the second branch function of the third loss function is determined as the function value of the third loss function.
[0102] 10 shows a schematic diagram of the process of determining the third loss function provided by an exemplary embodiment of the present application. As shown in FIG. 10, the server realizes the determination of the function value of the third loss function by calling the second discrimination device (general discrimination device), the identity recognition network, and the feature perception network. The image processing model generates a second sample image y k For an embodiment of processing the third predicted image [Equation 3] to obtain the third predicted image [Equation 3], reference may be made to step S304, which will not be described in detail here.
[0103] After obtaining the third predicted image [Equation 3], the server, on the other hand, uses a second discrimination device to obtain a second sample image y k, and the third predicted image [Equation 3], i.e., the second sample image y k , and the third predicted image [Equation 3] is judged to be an actual image, and the second sample image y k Based on the fourth discrimination result of the third predicted image [Equation 3] and the fifth discrimination result of the third predicted image [Equation 3], a function value of a first branch function of the third loss function is determined. The first branch function can be expressed as the following [Equation 22].
[0104]
number
[0105] Here, L GAN3 represents the first branch function of the third loss function of the second generative adversarial network GAN' (including the second generative network (G2) and the second discriminator (D2)). [Equation 23] represents the third predicted image [Equation 3] generated by the second generative network (G) and the second sample image y k The formula 24 represents the difference between the discrimination result of the third predicted image [Formula 3] by the second discrimination device and the second sample image y k The E(x) function is used to calculate the expected value of x, and D general (x) is used to represent the discrimination of x by using the second discrimination device, and I other is the second sample image y k where Enc(x) is used to denote the encoding of x using an encoder, and Equation 25 is used to denote the decoding of x according to y using a second decoder. As can be deduced from this, Equation 26 is used to denote the decoding of x according to y using a second discriminator. k [Equation 27] represents the third predicted image [Equation 3], and [Equation 28] represents the second discrimination device being employed to discriminate the third predicted image [Equation 3].
[0106]
number
number
number
number
number
number
[0107] On the other hand, the server uses a feature perception network (LPIPS network) to generate a second sample image y k , and the third predicted image [Equation 3] are subjected to feature recognition to obtain the second sample image y k and a fourth perceptual feature corresponding to the third predicted image [Equation 3], and perform feature comparison between the third perceptual feature and the fourth perceptual feature to obtain a second feature comparison result. The second feature comparison result is k and the third predicted image [Equation 3]. After obtaining the second feature comparison result, the server determines a second branch function of the third loss function according to the second feature comparison result, and the second branch function can be expressed as follows [Equation 29]:
[0108]
number
[0109] Here, L LPIPS2 represents the second branch function of the third loss function of the feature perception network (LPIPS network), and LPIPS(x) represents feature perception for x by the feature perception network (LPIPS network). As is clear from the first branch function of the third loss function, I other is the second sample image, i.e. y k[Equation 30] is the third predicted image, i.e., [Equation 3]. Based on this, [Equation 31] represents feature perception for the third predicted image, i.e., [Equation 3], by the LPIPS network. [Equation 32] represents feature perception for the second sample image, i.e., y k It represents feature perception.
[0110]
number
number
number
[0111] After obtaining the function value of the first branch function of the third loss function and the function value of the second branch function of the third loss function, the server determines the sum of the function value of the first branch function of the third loss function and the function value of the second branch function of the third loss function as the function value of the third loss function, and calculates the third loss function L 3 can be expressed as follows: [Equation 33]
[0112]
number
[0113] S308: Determine a function value of a target loss function based on the function value of the first loss function, the function value of the second loss function, and the function value of the third loss function.
[0114] Optionally, in the model training process, a function value of a target loss function of the image processing model can be determined based on the function value of the first loss function and the function value of the second loss function.
[0115] Taking training based on the third loss function value as an example, the target loss function can be expressed as follows:
[0116]
number
[0117] Alternatively, the target loss function value may be a sum of the first loss function value, the second loss function value, and the third loss function value, or the target loss function value may be a weighted sum of the first loss function value, the second loss function value, and the third loss function value, and the weights between any two of the three may be the same or different. In general, the target loss function can be expressed as follows:
[0118]
number
[0119] Here, a represents a weight value corresponding to the first loss function, b represents a weight value corresponding to the second loss function, and c represents a weight value corresponding to the third loss function, and any two of the three weight values may be the same or different, where a+b+c=1.
[0120] In one possible implementation, the steps of calculating the first loss function, the second loss function, and the third loss function may be performed simultaneously.
[0121] S309: Train an image processing model according to the function value of the target loss function.
[0122] In one embodiment, the server reduces the loss value of the total loss function by adjusting the parameters of the image processing model (e.g., the number of convolution layers, the number of upsampling layers, the number of downsampling layers, dilation, etc.). In general, the server back-propagates the error to the first generating network and the second generating network (encoder and decoder) according to the function value of the target loss function, and updates the parameter values of the first generating network and the second generating network using the gradient descent method. Optionally, in the model training process, the parts to be updated in parameters include the encoder, the specific decoder, the general decoder, the specific discriminator, and the general discriminator, and the LPIPS network, the identity identification network, and the attribute identification network are not involved in the parameter update.
[0123] In one possible implementation, the server adjusts the parameters of the image processing model by 1 +L 2 Adjust the parameters of the first generating network by L 3 The parameters of the second generator network can be adjusted by
[0124] During the training process of the image processing model, each time a parameter update is performed, the server can determine whether the updated image processing model reaches the training completion condition. If the updated image processing model does not reach the training completion condition, it repeats based on the above steps S302 to S309 and continues to train the image processing model until the image processing model reaches the training completion condition. If the updated image processing model reaches the training completion condition, the training of the image processing model is terminated.
[0125] Optionally, the image processing model in the embodiment of the present application can further include a second generating network, a first discriminating device, a second discriminating device, a feature perception network, an identity identification network and an attribute identification network in addition to the first generating network. In the model training process, the parameters in the first generating network, the second generating network, the first discriminating device and the second discriminating device are all updated, and after the training of the image processing model is completed, the first generating network is kept, that is, the image processing model that reaches the training completion condition includes the first generating network after the training is completed.
[0126] Alternatively, the image processing model of the present application includes a first generation network, and during the model training process, other network structures other than the image processing model are called to assist in the training of the image processing model, and after the training completion condition is reached, the image processing model after completion of training is obtained and applied.
[0127] When the first type of attributes are non-identifying attributes and the second type of attributes are identifying attributes, the input image is processed by the image processing model after completion of training, and the obtained predicted image retains the non-identifying attributes of the avatar in the input image, but may have the identifying attributes of the target avatar. In general, the characteristic attributes of the avatar in the predicted image, such as facial expression, posture, lighting, etc., are consistent with the input image, and the facial features, such as the five senses, skin, hair, glasses, etc., in the predicted image are consistent with the target avatar.
[0128] In general, the image processing model after training can be applied to the avatar replacement scene of the video. Take the process of the server performing avatar replacement processing on the test video as an example, Figure 11 shows a processing flowchart of the test video provided by one exemplary embodiment of the present application. As shown in Figure 11, the server obtains a test video, the test video includes R frames of test images, each frame of the test image includes one calibration avatar, R is a positive integer, Call the first generating network of the image processing model after training to process the test images of R frames respectively to obtain the test images of R frames and the corresponding predicted images, where the predicted images of R frames include a target avatar of a target person, and a first kind of attribute of the avatar in the predicted images of R frames is consistent with a first kind of attribute of the calibrated avatar in the corresponding test images; In the test video, image completion is performed on the test image of the R frame in which the calibration avatar is deleted. The predicted images of R frames are respectively fused with the corresponding test images in the test video after image interpolation to obtain the target video.
[0129] Meanwhile, the server performs image frame extraction (image extraction) on the test video to obtain a test image set. The test image of each frame includes a test avatar. The server calibrates the test avatar in the test image set, and the calibration method for the test avatar can refer to the embodiment in S3012, but will not be described in detail here. After the calibration is completed, the server calls the avatar replacement model to process the test image of each frame respectively, and obtains a predicted image corresponding to the test image of each frame. The avatar in the predicted image corresponds to a first type of attribute of the original avatar and a second type of attribute of the avatar of the target person included in the first sample image when it can include a training image processing model.
[0130] Optionally, after obtaining the predicted avatar corresponding to the test image of each frame, the avatar in the test image of each frame can be replaced later, and Fig. 12 shows a flowchart of processing the test image with the trained image processing model provided by an exemplary embodiment of the present application. In Fig. 12, a specific embodiment of the server employing the trained image processing model to process the test image of each frame respectively to obtain the predicted image corresponding to the test image of each frame may refer to S302 or S303, which will not be described in detail here.
[0131] On the other hand, for example, if the process of performing avatar conversion on an image in a test video is processed by a server, the server needs to delete the calibration avatar in the test image in the test video, and perform image completion (Inpainting) processing on the test image of the R frame in the test video after deleting the calibration avatar. The purpose is to make the restored image look relatively natural by completing the image itself (e.g., background) to be restored or the missing area (i.e., avatar conversion area) of the image to be restored according to the image library information after deleting the original avatar in the test image.
[0132] Fig. 13 shows a processing flow chart of the test video provided by one exemplary embodiment of the present application. As shown in Fig. 13, the server fuses the predicted image of each frame with the test image after the corresponding image complementation to obtain the target video.
[0133] Optionally, the server may perform color calibration (eg, skin color adjustment) on the fused video to make the captured target video more realistic.
[0134] In one embodiment, a user can upload a video of himself singing or dancing, and the server can use the trained image processing model to replace the avatar in the video uploaded by the user with a star avatar / anime avatar, etc. (the avatar of the target person), to obtain the target video after avatar conversion, which can further improve the entertainment value of the video. In addition, the user can also use the trained image processing model to perform "avatar conversion" live broadcasting (i.e., convert the avatar of the live broadcasting user into the avatar of the target person in real time), which can further increase the entertainment value of the live broadcasting.
[0135] In another embodiment, since mobile payment can be made through "face recognition", the image processing model after completion of training, which has relatively high requirements for the accuracy of the face identification model, can be used to generate training data (attack data) to train the face identification model (i.e., to train the ability of the face identification model to identify the veracity of the predicted avatar), thereby further improving the reliability and security of mobile payment.
[0136] In the embodiment of the present application, the image processing method provided in the embodiment of the present application uses a first sample set including a first sample image of a target avatar of the same target person and a second sample set including a second sample image of the sample avatar to train an image processing model including a first generation network, and in the training process, the image processing model is trained according to the avatar difference between the first sample image and a first predicted image obtained by processing it with the first generation network, and the difference of a first type of attribute between the second sample image and a second predicted image obtained by processing it with the first generation network, so that the image processing model after training can replace the avatar in the input image with the target avatar and retain the first type of attribute of the avatar in the input image, thereby enabling the avatar replacement based on the image processing model obtained by the present application to retain the specified attribute of the replaced avatar at the same time, thereby making the authenticity of the replaced avatar relatively high, and improving the accuracy of image processing and the replacement effect.
[0137] The method of the embodiment of the present application has been discussed in detail above, and in order to better implement the above means of the embodiment of the present application, the apparatus of the embodiment of the present application is presented below as well.
[0138] Figure 14 shows a structural schematic diagram of an image processing device provided by one exemplary embodiment of the present application. The image processing device can be installed in the server 102 shown in Figure 1. The image processing device shown in Figure 14 can be used to perform some or all of the functions in the method embodiments described by Figures 2 and 3 above.
[0139] Here, the image processing device is Invoke the first generative network in the image processing model and compute the first sample image x in the first sample set. i a first predicted image acquisition module 1410 for processing the first predicted image [Equation 1] to obtain a first predicted image [Equation 1], the first predicted image [Equation 1] including a first predicted avatar, the first sample set including N first sample images, each of the first sample images including a target avatar corresponding to a same target person, N being a positive integer, i being a positive integer, and i≦N; Invoke the first generator network and generate a second sample image y k a second predicted image obtaining module 1420 for processing the second predicted image [Equation 2] to obtain a second predicted image [Equation 2], the second predicted image [Equation 2] including a second predicted avatar, the second sample set including M second sample images, each second sample image including a sample avatar, M being a positive integer, k being a positive integer, and k≦M; Above 1st sample image x i the difference between the target avatar and the first predicted avatar in the second sample image y kand a model training module 1430 used to train the image processing model based on a difference between a first type of attributes of the sample avatar in the input image and the first type of attributes of the second predicted avatar, wherein the image processing model is used to replace an avatar in an input image with the target avatar and retain the first type of attributes of the avatar in the input image.
[0140] In one embodiment, the model training module 1430 includes: A first determination sub-module for determining a function value of a first loss function, the first loss function being a function value of the first sample image x i a first determination submodule for indicating a difference between the target avatar and the first predicted avatar in a second determination sub-module for determining a function value of a second loss function, the second loss function being a function value of the second sample image y k a second determination submodule for indicating a difference between the first type of attributes of the sample avatar and the first type of attributes of the second predicted avatar; A third determination submodule is used for determining a function value of a target loss function of the image processing model according to the function value of the first loss function and the function value of the second loss function; and a model training sub-module, which is used to train the image processing model according to the function value of the target loss function.
[0141] In one embodiment, the first determination submodule comprises: The first discrimination device is called, and the first sample image x i , and the first predicted avatar; Above 1st sample image x i and determining a function value of a first branch function of the first loss function based on a first discrimination result of the first sample image xi is used to indicate whether the first predicted avatar is a real image, and the second discrimination result is used to indicate whether the first predicted avatar is a real image; and determining a function value of a second branch function of the first loss function, the second branch function being a function value of the first sample image x i and indicating a difference between a first perceptual feature of the target avatar and a second perceptual feature of the first predicted avatar; determining a sum of a function value of a first branch function of the first loss function and a function value of a second branch function of the first loss function as a function value of the first loss function.
[0142] In one embodiment, the second determination submodule comprises: calling a first discrimination device to discriminate the second predicted avatar; determining a function value of a first branch function of the second loss function based on a third discrimination result of the second predicted avatar, the third discrimination result being used to indicate whether the second predicted avatar is a real image; and Above second sample image y k performing an attribute comparison between the first type of attribute of the sample avatar and the first type of attribute of the second predicted avatar to obtain an attribute comparison result; determining a second split function of the second loss function based on the attribute comparison; and determining, as a function value of the second loss function, a sum of a function value of a first branch function of the second loss function and a function value of a second branch function of the second loss function.
[0143] In one embodiment, the apparatus comprises: Call the second generator network and generate the second sample image y ka third predicted image obtaining module used for processing the third predicted image [equation 3] to obtain a third predicted image [equation 3], where the third predicted image [equation 3] includes a third predicted avatar, and the second generating network and the first generating network have the same feature extraction unit; The model training module 1430 includes: a fourth determination sub-module for determining a function value of a third loss function, the third loss function being a function value of the second sample image y k a fourth determination submodule for indicating a difference between the sample avatar and the third predicted avatar in The third determination sub-module is used for determining a function value of the target loss function according to the function value of the first loss function, the function value of the second loss function, and the function value of the third loss function.
[0144] In one embodiment, the fourth determination submodule: The second discrimination device is called, and the second sample image y k , and the third predicted avatar; and Above second sample image y k and determining a function value of a first branch function of the third loss function based on a fourth discrimination result of the second sample image y k is used to indicate whether the fifth discrimination result is used to indicate whether the third predicted avatar is a real image; and determining a function value of a second branch function of the third loss function, the second branch function of the third loss function being a function value of the second sample image y k and indicating a difference between a third perceptual feature of the sample avatar and a fourth perceptual feature of the third predicted avatar in the sample avatar. determining, as a function value of the third loss function, a sum of a function value of a first branch function of the third loss function and a function value of a second branch function of the third loss function.
[0145] In one embodiment, the first generating network includes an encoder and a first decoder; The first predicted image acquisition module 1410 includes: Call the encoder to generate the first sample image x i to obtain a first feature vector; calling the first decoder to decode the first feature vector to obtain a first generated image and the first region division information, the first region division information being used to indicate an avatar region in the first generated image; and extracting the first predicted image [Equation 1] from the first generated image based on the first region division information.
[0146] In one embodiment, the first generating network includes an encoder and a first decoder; The second predicted image acquisition module 1420 is Call the encoder to generate the second sample image y k to obtain a second feature vector; calling the first decoder to decode the second feature vector to obtain a second generated image and the second region division information, the second region division information being used to indicate an avatar region in the second generated image; and and extracting the second predicted image [Equation 2] from the second derived image based on the second region division information.
[0147] In one embodiment, the encoder includes P feature extraction networks and one feature aggregation layer, where P is a positive integer, each feature extraction network includes one downsampling layer, and the scale parameters of the P downsampling layers are different; The P downsampling layers are used to extract feature information of the image input to the encoder under the P scale parameters; The feature aggregation layer is used to perform an aggregation process on the feature information under the P scale parameters to obtain a feature vector corresponding to the image input to the encoder.
[0148] In one embodiment, the first decoder includes a first feature transformation network, Q first image reconstruction networks, and a first convolution network, where Q is a positive integer, and each of the first image reconstruction networks includes a first residual network and a first upsampling layer; The first feature transformation network is used to transform the feature vector input to the first decoder into a feature map; The Q first image reconstruction networks are used to perform a first feature restoration process on the feature map to obtain a fusion feature image; The first convolutional network is used for performing a convolution process on the fusion feature image, and outputting a generated image corresponding to the feature vector input to the first decoder.
[0149] In one embodiment, the second generating network includes an encoder, a second decoder, and an identity identification network; Call the second generator network and generate the second sample image y k The above step of processing to obtain the third predicted image [Equation 3] is Call the encoder to generate the second sample image y k to obtain a second feature vector; The second sample image y is obtained by calling the identification network. k Identify and use the second sample image y k obtaining an identity of the Invoke the second decoder and generate the second sample image y kdecoding the second feature vector according to the identification information to obtain a third generated image and the third region division information, the third region division information being used to indicate an avatar region in the third generated image; and extracting the third predicted image [Equation 3] from the third derived image based on the third region division information.
[0150] In one embodiment, the second decoder includes a second feature transformation network, Q second image reconstruction networks, and a second convolution network, where Q is a positive integer, and each of the second image reconstruction networks includes a second residual network, a second upsampling layer, and a self-adaptation module; The self-adaptation module, in the decoding process of the second decoder, k and obtaining a third feature vector corresponding to the identification information based on the identification information, and the third feature vector is used to instruct the second decoder to decode the feature vector input to the second decoder.
[0151] In one embodiment, the apparatus comprises: a test video acquisition module used to acquire the test video after the model training module 1430 completes training for the image processing model, the test video including R frames of test images, each frame of the test image including one calibration avatar, where R is a positive integer; a fourth predicted image acquisition module, which is used to call the first generating network of the image processing model after training is completed to process the test images of R frames respectively, and obtain predicted images respectively corresponding to the test images of R frames, where the predicted images of R frames include the target avatar of the target person, and the first kind of attributes of the avatar in the predicted images of R frames are consistent with the first kind of attributes of the calibration avatar in the corresponding test images; an image completion module is used for performing image completion on the test image of the R frame in the test video from which the calibration avatar is removed; a target video acquisition module used for fusing the predicted images of R frames with the corresponding test images in the test video after image interpolation to obtain a target video.
[0152] In one embodiment, the first type of attributes refers to non-identifying attributes.
[0153] According to one embodiment of the present application, some steps of the image processing method shown in FIG. 2 and FIG. 3 can be performed by each module or sub-module in the image processing device shown in FIG. 14. Each or all of the modules or sub-modules in the image processing device shown in FIG. 14 can be configured by merging with one or several other structures, or some of the modules can be further functionally decomposed into several smaller structures, thereby realizing the same operation and not affecting the realization of the technical effect of the embodiment of the present application. The above modules are divided based on logic functions, and in the practical application, the function of one module can be realized by several modules, or the function of several modules can be realized by one module. In other embodiments of the present application, the image processing device can also include other modules, and in the practical application, the functions can be realized with the assistance of other modules and can be realized jointly by several modules.
[0154] According to another embodiment of the present application, a computer program (including program code) capable of executing each step of the corresponding method shown in Fig. 2 or Fig. 3 can be operated in a general-purpose computing device, such as a computer, including a processing element such as a central processing unit (CPU), a random access storage medium (RAM), a read-only storage medium (ROM), and a storage element, to configure the image processing device shown in Fig. 14 and realize the image processing method of the embodiment of the present application. The computer program can be stored in a computer-readable recording medium, for example, and can be loaded into the computing device by the computer-readable recording medium and can operate in the computing device.
[0155] Based on the technical idea of the same invention, the principle by which the image processing device provided in the embodiments of the present application solves the problem and the beneficial effects thereof are similar to the principle by which the image processing method in the method embodiments of the present application solves the problem and the beneficial effects thereof, so reference may be made to the principle by which the method is implemented and the beneficial effects thereof, but for the sake of brevity, they will not be described in detail here.
[0156] Referring to FIG. 15, FIG. 15 shows a structural schematic diagram of an image processing device provided by one exemplary embodiment of the present application. The image processing device may be the server 102 shown in FIG. 1, and the image processing device includes at least a processor 1501, a communication interface 1502, and a memory 1503. Here, the processor 1501, the communication interface 1502, and the memory 1503 can be connected by a bus or other methods, and in the embodiment of the present application, they are connected by a bus as an example. Here, the processor 1501 (also called a central processing unit (CPU)) is the calculation core and control core of the image processing device, which can analyze various commands in the image processing device and process various data of the image processing device, for example, the CPU can be used to analyze the power on / off command sent by the user to the image processing device and control the image processing device to perform the power on / off operation, and for example, the CPU can transmit various interaction data between the internal structures of the image processing device. The communication interface 1502 may selectively include a standard wired interface, a wireless interface (e.g., WI-FI, a mobile communication interface, etc.), and may be controlled by the processor 1501 and used to receive and transmit data. The communication interface 1502 may further be used for internal data transmission and interaction of the image processing device. The memory 1503 is a storage device in the image processing device and may be used to store programs and data. As can be understood, the memory 1503 here may not only include the built-in memory of the image processing device, but may also include an expansion memory supported by the image processing device. The memory 1503 provides a storage area in which an operating system of the image processing device is stored, including but not limited to an Android system, an iOS system, a Windows Phone system, etc., and the present application is not limited in this respect.
[0157] In the embodiment of the present application, the processor 1501 realizes the means disclosed in the present application by operating the executable program code in the memory 1503. The operations performed by the processor 1501 may refer to the introductions in the above method embodiments, but will not be described in detail here.
[0158] Based on the technical idea of the same invention, the principle by which the image processing device provided in the embodiments of the present application solves the problem and the beneficial effects thereof are similar to the principle by which the image processing method in the method embodiments of the present application solves the problem and the beneficial effects thereof, and reference can be made to the principle by which the method is implemented and the beneficial effects thereof, so for the sake of brevity, they will not be described in detail here.
[0159] An embodiment of the present application further provides a computer readable storage medium, in which a computer program is stored, the computer program being suitable to be loaded by a processor and to execute the image processing method of the above method embodiment.
[0160] An embodiment of the present application further provides a computer program product or a computer program, the computer program product or the computer program comprising computer instructions stored in a computer readable storage medium, a processor of a computing device reading the computer instructions from the computer readable storage medium, and the processor executing the computer instructions to cause the computing device to perform the image processing method.
[0161] It is necessary to explain that, for the sake of simplicity, the above method embodiments are expressed as a combination of a series of operations, but those skilled in the art can understand that the present application is not limited to the order of operations described, because some steps can be performed in other orders or simultaneously according to the present application. Next, those skilled in the art can also understand that the embodiments described in the specification belong to preferred embodiments, and the related operations and modules are not necessarily essential to the present application.
[0162] The steps in the methods of the embodiments of the present application may be rearranged, merged, or reduced according to practical needs.
[0163] The modules in the device of the embodiment of the present application can be merged, divided and reduced according to actual needs.
[0164] As can be understood by those skilled in the art, all or some of the steps in the various methods of the above embodiments can be completed by issuing instructions to related hardware through a program. The program can be stored in a computer-readable storage medium, which can include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0165] The above is merely one relatively preferred embodiment of the present application, and of course, it should not be used to limit the scope of the claims of the present application. As can be understood by those skilled in the art, the implementation of all or part of the process of the above embodiment and the equivalent modifications made based on the claims of the present application are also within the scope of coverage of the present application. [Explanation of symbols]
[0166] 101 Terminal Equipment 102 Server 1410 1st predicted image acquisition module 1420 Second predicted image acquisition module 1430 Model Training Module 1501 Processor 1502 Communication Interface 1503 Memory
Claims
1. 1. An image processing method, comprising: Invoke a first generator network in the image processing model and generate a first sample image x in the first sample set. i to obtain a first predicted image [Equation 1], the first predicted image [Equation 1] including a first predicted avatar, the first sample set including N first sample images, each of the first sample images including a target avatar corresponding to a same target person, N being a positive integer, i being a positive integer, and i≦N; Invoke the first generator network and generate a second sample image y k to obtain a second predicted image [Equation 2], the second predicted image [Equation 2] including a second predicted avatar, the second sample set including M second sample images, each second sample image including a sample avatar, M being a positive integer, k being a positive integer, and k≦M; The first sample image x i and the difference between the target avatar and the first predicted avatar in the second sample image y k training the image processing model based on a difference between a first type of attributes of the sample avatar in the input image and the first type of attributes of the second predicted avatar, the image processing model being used to replace an avatar in an input image with the target avatar and retain the first type of attributes of the avatar in the input image; The first generator network includes an encoder and a decoder; The encoder extracts image features to encode the first sample image x i or the second sample image y k and obtain the corresponding feature vector, the decoder obtains a generated image and region segmentation information used to indicate an avatar region in the generated image based on the feature vector; an image processing method comprising: extracting the first predicted image [Equation 1] or the second predicted image [Equation 2] from the generated image based on the region division information. [0010] [0025]
2. The first sample image x i the difference between the target avatar and the first predicted avatar in the second sample image y k training the image processing model based on a difference between a first type of attribute of the sample avatar and the first type of attribute of the second predicted avatar in determining a function value of a first loss function, the first loss function being a function value of the first sample image x i and indicating a difference between the target avatar and the first predicted avatar in determining a function value of a second loss function, the second loss function being calculated based on the second sample image y k and indicating a difference between a first type of attribute of the sample avatar and the first type of attribute of the second predicted avatar in the determining a function value of a target loss function of the image processing model based on a function value of the first loss function and a function value of the second loss function; and training the image processing model according to function values of the target loss function.
3. The step of determining a function value of a first loss function comprises: A first discrimination device is called to obtain the first sample image x i , and the first predicted avatar, respectively; The first sample image x i and a second discrimination result of the first predicted avatar, the first discrimination result being determined based on the first sample image x i the first discrimination result is used to indicate whether the first predicted avatar is a real image, and the second discrimination result is used to indicate whether the first predicted avatar is a real image; Determining a function value of a second branch function of the first loss function, the second branch function of the first loss function being determined based on the first sample image x i and indicating a difference between a first perceptual feature of the target avatar and a second perceptual feature of the first predicted avatar in the step of determining a sum of a function value of a first branch function of the first loss function and a function value of a second branch function of the first loss function as the function value of the first loss function.
4. The step of determining a function value of a second loss function comprises: calling a first discrimination device to discriminate the second predicted avatar; determining a function value of a first branch function of the second loss function based on a third discrimination result of the second predicted avatar, the third discrimination result being used to indicate whether the second predicted avatar is a real image; The second sample image y k performing an attribute comparison between the first type of attribute of the sample avatar and the first type of attribute of the second predicted avatar to obtain an attribute comparison result; determining a function value of a second branch function of the second loss function based on the attribute comparison result; determining a sum of a function value of a first branch function of the second loss function and a function value of a second branch function of the second loss function as the function value of the second loss function.
5. Before constructing a target loss function of the image processing model according to the first loss function and the second loss function, the method further comprises: Invoke a second generator network and generate the second sample image y k to obtain a third predicted image [Equation 3], the third predicted image [Equation 3] including a third predicted avatar, the second generating network and the first generating network having the same feature extraction unit; determining a function value of a third loss function, the third loss function being calculated based on the second sample image y k and indicating a difference between the sample avatar and the third predicted avatar in The step of determining a function value of a target loss function of the image processing model based on a function value of the first loss function and a function value of the second loss function includes:
3. The method of claim 2, comprising determining a function value of the target loss function based on a function value of the first loss function, a function value of the second loss function, and a function value of the third loss function. [0030]
6. The step of determining a function value of a third loss function comprises: A second discrimination device is called to obtain the second sample image y k , and the third predicted avatar, respectively; The second sample image y k determining a function value of a first branch function of the third loss function based on a fourth discrimination result of the second sample image y k the fifth discrimination result is used to indicate whether the third predicted avatar is a real image; and the fifth discrimination result is used to indicate whether the third predicted avatar is a real image. Determining a function value of a second branch function of the third loss function, the second branch function of the third loss function being determined based on the second sample image y k and indicating a difference between a third perceptual feature of the sample avatar and a fourth perceptual feature of the third predicted avatar in determining a sum of a function value of a first branch function of the third loss function and a function value of a second branch function of the third loss function as a function value of the third loss function.
7. the first generator network includes an encoder and a first decoder; Invoke a first generator network in the image processing model and generate a first sample image x in the first sample set. i to obtain a first predicted image [Equation 1], Invoke the encoder to generate the first sample image x i to obtain a first feature vector; calling the first decoder to decode the first feature vector to obtain a first generated image and first region segmentation information, the first region segmentation information being used to indicate an avatar region in the first generated image; The method of claim 1 , further comprising: extracting the first predicted image [Equation 1] from the first generated image based on the first region division information.
8. the first generator network includes an encoder and a first decoder; Invoke the first generator network and generate a second sample image y k to obtain a second predicted image [Equation 2], Invoke the encoder to generate the second sample image y k to obtain a second feature vector; calling the first decoder to decode the second feature vector to obtain a second generated image and second region segmentation information, the second region segmentation information being used to indicate an avatar region in the second generated image; The method of claim 1 , further comprising: extracting the second predicted image [Equation 2] from the second generated image based on the second region division information.
9. The encoder includes P feature extraction networks and one feature aggregation layer, where P is a positive integer, each feature extraction network includes one downsampling layer, and the scale parameters of the P downsampling layers are different; The P downsampling layers are used to extract feature information of the image input to the encoder under the P scale parameters; The method according to claim 7 or 8, wherein the feature aggregation layer is used to perform an aggregation process on the feature information under the P scale parameters to obtain a feature vector corresponding to an image input to the encoder.
10. The first decoder includes a first feature transformation network, Q first image reconstruction networks, and a first convolution network, where Q is a positive integer, and each of the first image reconstruction networks includes a first residual network and a first upsampling layer; The first feature transformation network is used to transform the feature vector input to the first decoder into a feature map; The Q first image reconstruction networks are used to perform a first feature restoration process on the feature map to obtain a fusion feature image; The method according to claim 7 or 8, wherein the first convolutional network is used to perform a convolution process on the fusion feature image and output a generated image corresponding to a feature vector input to the first decoder.
11. the second generating network includes an encoder, a second decoder, and an identity recognition network; Invoke a second generator network and generate the second sample image y k to obtain a third predicted image [Equation 3], Invoke the encoder to generate the second sample image y k to obtain a second feature vector; The identity network is then called to retrieve the second sample image y k and extracting the second sample image y k obtaining an identity of the Invoke the second decoder and generate the second sample image y k decoding the second feature vector according to the identification information to obtain a third generated image and third region division information, the third region division information being used to indicate an avatar region in the third generated image; The method according to claim 5 , further comprising: extracting the third predicted image [Equation 3] from the third generated image based on the third region division information.
12. The second decoder includes a second feature transformation network, Q second image reconstruction networks, and a second convolution network, where Q is a positive integer, and each of the second image reconstruction networks includes a second residual network, a second upsampling layer, and a self-adaptation module; The self-adaptation module, in the decoding process of the second decoder, k and obtaining a third feature vector corresponding to the identification information based on the identification information of the first decoder, the third feature vector being used to instruct the second decoder to decode the feature vector input to the second decoder.
13. After completing training of the image processing model, the method further comprises: acquiring a test video, the test video including R frames of test images, each frame of the test image including one calibration avatar, where R is a positive integer; calling a first generation network of the image processing model after training is completed to process the test images of R frames respectively to obtain predicted images respectively corresponding to the test images of R frames, the predicted images of R frames including the target avatar of the target person, and the first type of attributes of the avatar in the predicted images of R frames being consistent with the first type of attributes of the calibration avatar in the corresponding test images; performing image interpolation on the test image of an R frame in the test video from which the calibration avatar has been removed; The method of claim 1 , further comprising: fusing each of the predicted images of R frames with a corresponding test image in a test video after image interpolation to obtain a target video.
14. The method according to any one of claims 1 to 13, wherein the first type of attributes refers to non-identifying attributes.
15. An image processing device, comprising: Invoke a first generator network in the image processing model and generate a first sample image x in the first sample set. i a first predicted image acquisition module for processing the above to obtain a first predicted image [Equation 1], the first predicted image [Equation 1] including a first predicted avatar, the first sample set including N first sample images, each of the first sample images including a target avatar corresponding to a same target person, N being a positive integer, i being a positive integer, and i≦N; Invoke the first generator network and generate a second sample image y k a second predicted image acquisition module for processing the second predicted image to obtain a second predicted image, the second predicted image including a second predicted avatar, the second sample set including M second sample images, each of the second sample images including a sample avatar, M being a positive integer, k being a positive integer, and k≦M; The first sample image x i the difference between the target avatar and the first predicted avatar in the second sample image y k a model training module used to train the image processing model based on a difference between a first type of attribute of the sample avatar in the input image and the first type of attribute of the second predicted avatar, the image processing model being used to replace an avatar in an input image with the target avatar and retain the first type of attribute of the avatar in the input image; The first generator network includes an encoder and a decoder; The encoder extracts image features to encode the first sample image x i or the second sample image y k and obtain the corresponding feature vector, the decoder obtains a generated image and region segmentation information used to indicate an avatar region in the generated image based on the feature vector; An image processing device that extracts the first predicted image [Equation 1] or the second predicted image [Equation 2] from the generated image based on the region division information.
16. An image processing device, comprising: a storage device; and a processor; The storage device stores a computer program, The image processing device, wherein the processor is used to realize the image processing method according to any one of claims 1 to 14 by loading and executing the computer program.
17. A computer program for causing a computer to execute the image processing method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Data processing method and device for face image generation and medium
CN110084193A
Method and device for generating face image
CN111523413A
A method and device for processing persona data
CN112017140A
Learning system, analysis system, method for learning, method for analysis, program, and storage medium
JP2021043839A