Method for training neural network model and action migration

By combining cross-reconstruction and self-reconstruction training of neural network models with transformation prediction and decoding networks, the problems of poor transfer effect and poor versatility in motion transfer are solved, achieving efficient and accurate motion or facial expression transfer, which is applicable to animation or film production with various control rules and character types.

CN120832931APending Publication Date: 2025-10-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410468701.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-17
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing technologies struggle to ensure that the transferred actions are both realistic and natural during the motion transfer process. Furthermore, when dealing with differences in actions between different individuals, they rely on a pre-established database of character expression correspondences and training methods tailored to specific control rules, resulting in poor versatility and high costs for manual annotation.

Method used

A neural network model is used to train a transformation prediction network through cross-reconstruction and self-reconstruction processes. Combined with a transformation decoding network, the transfer of actions or expressions is realized to generate target images, reducing the dependence on a pre-established database of character expression correspondences and adapting to different control rules.

Benefits of technology

It achieves high-accuracy motion or facial expression transfer under unsupervised learning, reduces manual annotation costs, simplifies video or animation production, and is applicable to various control rules and character types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832931A_ABST
    Figure CN120832931A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a method for training a neural network model, a method for motion migration, information processing equipment, a computer program product and a storage medium. The method for training the neural network model comprises the steps that a first image sample, a second image sample and a third image sample are determined, the first image sample comprises a first role, and the second image sample and the third image sample comprise a second role; the actions of the second characters in the second image sample and the third image sample are different; and training the transformation prediction network based on the first image sample, the second image sample and the third image sample. According to the method disclosed by the invention, the trained transformation prediction network can accurately realize transformation prediction so as to further realize action migration based on the trained transformation prediction network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of artificial intelligence, and more particularly, to a method, apparatus, computer program product and storage medium for training a neural network model, and a method, apparatus, computer program product and storage medium for action transfer. BACKGROUND

[0002] Artificial intelligence (AI) is the theory, method, technology and application system that use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology of computer science, which aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0003] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model or basic model, which can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0004] Based on artificial intelligence technology, action transfer can bring great changes to the fields of movies, animation, games, etc. For example, in the field of movie production, artificial intelligence technology can be used to transfer the actions of real actors to movie characters, so that movie characters can better present various actions close to real actors, thus creating more shocking and more enjoyable scenes. In the field of game development, artificial intelligence technology can be used to transfer the actions of real people to game characters, thus simplifying the game development process and improving the immersion of players. In the field of animation production, artificial intelligence technology can be used to transfer the actions of real people to animation characters, thus making the actions of animation characters more smooth and simplifying the animation development process.

[0005] In the case of more and more complex action transfer scenarios, how to ensure that the transferred actions are both real and natural, and how to handle the action differences between different individuals are among the key research directions in the field of action transfer. SUMMARY

[0006] To accurately complete the motion transfer task by using artificial intelligence technology, the present disclosure provides a method for training a neural network model, comprising: determining a first image sample, a second image sample and a third image sample, wherein the first image sample contains a first character, the second image sample and the third image sample contain a second character, and the actions of the second character in the second image sample and the third image sample are different; training a transformation prediction network through a first training process, wherein the first training process trains the transformation prediction network based on the first image sample, the second image sample and the third image sample, wherein the first training process comprises: predicting a first transformation from the first image sample to the second image sample by using the transformation prediction network; generating a first target image based on the first image sample and the first transformation by using an image generation network, wherein the character in the first target image is consistent with the character in the first image sample, and has an action consistent with the character in the second image sample; predicting a second transformation from the third image sample to the first target image by using the transformation prediction network; generating a second target image based on the third image sample and the second transformation by using the image generation network, wherein the character in the second target image is consistent with the character in the third image sample, and has an action consistent with the character in the first target image; and calculating a value of a first loss function based on the difference between the second image sample and the second target image to train the transformation prediction network.

[0007] According to an embodiment of the present disclosure, a real character is randomly taken as the first character, a virtual character is randomly taken as the second character, or a virtual character is taken as the first character and a real character is taken as the second character, and the real character and the virtual character are character characters or animal characters.

[0008] According to an embodiment of the present disclosure, the method further comprises: training the transformation prediction network through a second training process, wherein the second training process trains the transformation prediction network based on the second image sample and the third image sample, wherein the second training process comprises: predicting a third transformation from the third image sample to the second image sample by using the transformation prediction network; generating a third target image based on the third image sample and the third transformation by using the image generation network, wherein the character in the third target image is consistent with the character in the third image sample, and has an action consistent with the character in the second image sample; and calculating a value of a second loss function based on the difference between the second image sample and the third target image to train the transformation prediction network.

[0009] According to an embodiment of the present disclosure, the transform prediction network can be trained based on the first training process and the second training process simultaneously; or the transform prediction network can be trained based on the second training process first and then trained based on the first training process.

[0010] According to an embodiment of the present disclosure, the second role is a specific virtual role, the first role is selected from a plurality of real roles, and the first training process or the second training process is performed respectively for the plurality of real roles.

[0011] According to an embodiment of the present disclosure, the first role is selected from a plurality of real roles, the second role is a specific virtual role, and the method further comprises: obtaining a fourth image sample containing the second role; obtaining a control parameter corresponding to a change from a standard image of the specific role to the fourth image sample; predicting a fourth transform between the standard image of the specific role and the fourth image sample by using the transform prediction network; decoding the fourth transform by using a transform decoding network to obtain a predicted control parameter; and calculating a value of a fourth loss function based on a difference between the control parameter and the predicted control parameter to train the transform decoding network.

[0012] Embodiments of the present disclosure also provide a method for action transfer, comprising: obtaining a first image containing a first role and a second image containing a second role; predicting a predicted transform between the first image and the second image by using a transform prediction network; and generating a third image based on the first image and the predicted transform, wherein a role in the third image is consistent with a role in the first image and has a consistent action with a role in the second image.

[0013] According to an embodiment of the present disclosure, generating a third image based on the first image and the predicted transform comprises: generating the third image based on the first image and the predicted transform by using an image generation network; or decoding the predicted transform by using a transform decoding network to obtain a predicted control parameter, and generating the third image based on the first image and the predicted control parameter.

[0014] An embodiment of the present disclosure further provides an apparatus for training a neural network model, comprising: a sample determination module, configured to: determine a first image sample, a second image sample, and a third image sample, wherein the first image sample contains a first character, the second image sample and the third image sample contain a second character, and the actions of the second character in the second image sample and the third image sample are different; a model training module, configured to: train a transformation prediction network through a first training process, wherein the first training process trains the transformation prediction network based on the first image sample, the second image sample, and the third image sample, wherein the first training process includes: using the transformation prediction network to predict a first transformation from the first image sample to the second image sample exchange; using an image generation network to generate a first target image based on the first image sample and the first transformation, wherein the character in the first target image is consistent with the character in the first image sample and has the same action as the character in the second image sample; using the transformation prediction network to predict a second transformation from the third image sample to the first target image; using the image generation network to generate a second target image based on the third image sample and the second transformation, wherein the character in the second target image is consistent with the character in the third image sample and has the same action as the character in the first target image; and calculating the value of a first loss function based on the difference between the second image sample and the second target image to train the transformation prediction network.

[0015] An embodiment of the present disclosure also provides a device for action migration, including: an image acquisition module, configured to acquire a first image containing a first character and a second image containing a second character; a transformation prediction module, configured to predict a predicted transformation from the first image to the second image using a transformation prediction network; an image generation module, configured to generate a third image based on the first image and the predicted transformation, wherein the character in the third image is consistent with the character in the first image and has an action consistent with the character in the second image.

[0016] An embodiment of the present disclosure further provides an information processing device, including: a memory and a processor, wherein the processor is coupled to the memory and configured to provide the above method.

[0017] An embodiment of the present disclosure further provides a computer program product, which includes computer software code. When the computer software code is executed by a processor, it provides the above method.

[0018] The embodiments of the present disclosure also provide a computer readable storage medium, which stores computer executable instructions. The instructions, when executed by a processor, provide the method described above.

[0019] The method of the present disclosure can enable the trained transformation prediction network to accurately implement transformation prediction, so as to further implement action migration based on the trained transformation prediction network. The method of the present disclosure can fully utilize the neural network model to implement action (including posture, expression, gesture) migration to generate the required target image, without further post-processing or pre-establishing a perfect role expression corresponding relationship database. The method of the present disclosure can also obtain good training effect for the transformation prediction network in the case of unsupervised learning, reducing the cost of manual annotation. The training method (and model architecture) of the present disclosure has universality, and does not need to design a training process for each control rule. The neural network model (including the transformation prediction network and the transformation decoding network) for action migration obtained by training can obtain control parameters for controlling the role, for subsequent video or animation production, thereby simplifying the video or animation production process. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings needed in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some exemplary embodiments of the present disclosure, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0021] Herein, in the drawings:

[0022] Figure 1 a schematic diagram of an application scenario according to an embodiment of the present disclosure is shown;

[0023] Figure 2 is an example schematic diagram showing a scenario of action migration and training based on a neural network model for action migration according to an embodiment of the present disclosure;

[0024] Figure 3 is a schematic diagram showing a training process of a transformation prediction network according to an embodiment of the present disclosure;

[0025] Figure 4 is a schematic diagram showing a training process of a transformation decoding network according to an embodiment of the present disclosure;

[0026] Figures 5A-5B is a schematic diagram showing an action migration process according to an embodiment of the present disclosure;

[0027] Figures 6A-6C is a schematic flowchart of a method of training a transformation prediction network according to an embodiment of the present disclosure;

[0028] Figure 6D FIG. 14 is a schematic flowchart illustrating a method of training a transformation-decoding network according to an embodiment of the disclosure;

[0029] Figure 7A FIG. 15 is a schematic flowchart illustrating a method of action transfer according to an embodiment of the disclosure;

[0030] Figures 7B-7C FIG. 16 is an effect schematic diagram illustrating a method of action transfer according to an embodiment of the disclosure;

[0031] Figure 8 FIG. 17 is a constituent schematic diagram of an apparatus for training a neural network model according to an embodiment of the disclosure;

[0032] Figure 9 FIG. 18 is a constituent schematic diagram of an apparatus for action transfer according to an embodiment of the disclosure; and

[0033] Figure 10 FIG. 19 is an architecture of a computing device according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0034] In order to make the objectives, technical solutions and advantages of the disclosure more obvious, the following will describe the example embodiments according to the disclosure in detail with reference to the drawings. Obviously, the described embodiments are only part of the embodiments of the disclosure, not all the embodiments of the disclosure, and it should be understood that the disclosure is not limited to the example embodiments described herein.

[0035] In addition, in the specification and drawings, substantially identical or similar steps and elements are denoted by the same or similar reference numerals, and repeated descriptions thereof will be omitted.

[0036] In addition, in the specification and drawings, elements are described in singular or plural form according to the embodiments. However, the selection of singular and plural forms for the proposed case is only for the convenience of explanation and is not intended to limit the disclosure thereto. Therefore, the singular form can include the plural form, and the plural form can also include the singular form, unless the context clearly indicates otherwise.

[0037] In the specification and drawings, substantially identical or similar steps or elements are denoted by the same or similar reference numerals, and repeated descriptions thereof will be omitted. Meanwhile, in the description of the disclosure, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance or order.

[0038] For the convenience of describing the disclosure, the following introduces the concepts related to the disclosure.

[0039] Action transfer technology is a technique that transfers an action or action pattern from one entity or environment to another. In the field of artificial intelligence, action transfer technology uses deep learning and artificial intelligence to learn and understand the characteristics of human actions from a large amount of training data. By applying these characteristics to other characters or virtual figures, the transfer and reproduction of actions are achieved. This technology not only reduces costs, but also improves production efficiency and the realism of actions. For example, in the process of computer animation, the pose of a human body can be transferred to a cartoon model through action transfer, providing great convenience for animation creation. In addition, action transfer technology plays an important role in virtual fitting, character animation, film and game creation, etc. Through transfer technology, virtual characters can be given lively actions more easily, improving the visual appeal and interactivity of the work.

[0040] Expression transfer refers to transferring the pose and expression of a face object in a driving image to a face object in a source image to generate a driving result image. The driving result image has the identity information of the face object in the source image and the pose and expression information of the face object in the driving image. Expression transfer can be realized based on deep learning and computer vision technology, and the core is to build a model that can learn and transfer facial expressions. First, the model extracts features from the source image containing the face of the target person. Then, the model uses the expression features in the driving image or video (containing the expression to be transferred) to transfer the expression features in the driving image to the target person's face in the source image. Finally, the model generates a new image in which the target person's facial expression has been replaced with the expression in the driving image. This new image not only retains the identity information of the target person in the source image, but also has the expression features of the driving image, thus realizing expression transfer.

[0041] The various neural networks (or neural network models) that can be used in the embodiments of the present disclosure below can all be artificial intelligence models, especially artificial intelligence-based neural network models. Generally, artificial intelligence-based neural network models are implemented as acyclic graphs, with neurons arranged in different layers. Generally, a neural network model includes an input layer and an output layer separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating output in the output layer. Network nodes (i.e., neurons) are fully connected to nodes in adjacent layers via edges, and there are no edges between nodes within each layer. Data received at the nodes of the input layer of the neural network is propagated to the nodes of the output layer via any of the hidden layers, activation layers, pooling layers, convolutional layers, etc. The input and output of the neural network model can take various forms, which are not limited by the present disclosure.

[0042] In summary, the present disclosure relates to technologies such as artificial intelligence, action transfer (including expression transfer), etc. Embodiments of the present disclosure will be further described below in conjunction with the accompanying drawings.

[0043] Firstly, refer to Figure 1 The application scenarios of the method and the corresponding apparatus, etc. according to embodiments of the present disclosure are described. Figure 1 A schematic diagram of an application scenario 100 according to embodiments of the present disclosure is shown, in which a server 110 and a plurality of terminals 120 are shown schematically.

[0044] The neural network model for action transfer of embodiments of the present disclosure can be integrated in an apparatus for action transfer and located in various electronic devices, for example, Figure 1 any electronic device in the server 110 and the plurality of terminals 120. For example, the neural network model for action transfer can be integrated in the terminal 120. The terminal 120 includes but is not limited to a mobile phone, a computer, a smart voice interactive device, a smart home appliance, and a vehicle terminal. For another example, the neural network model for action transfer can also be integrated in the server 110. The server 110 can be a stand-alone physical server, a server cluster composed of multiple physical servers or a distributed system, and can also be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server can be connected directly or indirectly through wired or wireless communication, which is not limited in the present disclosure.

[0045] It can be understood that the apparatus for action transfer applying embodiments of the present disclosure can be a terminal, a server, or a system composed of a terminal and a server. The method for action transfer applying embodiments of the present disclosure can be executed on a terminal, a server, or a terminal and a server jointly.

[0046] The neural network model for action transfer provided by embodiments of the present disclosure can be used to perform various types of action (including expression or posture) transfer tasks. For example, the action (or expression / posture) of a real role can be transferred to a 3D virtual role, the action (or expression / posture) of a 3D virtual role can be transferred to a 2D virtual role, etc. The role in the present disclosure can refer to a person (for example, the method of the present disclosure is used for the transfer of expressions of a person) or an animal (for example, the method of the present disclosure is used for the transfer of actions of an animal).

[0047] The neural network model for target detection provided by the embodiments of the present disclosure can also relate to artificial intelligence cloud services in the field of cloud technology. The cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and network in a wide area network or a local area network to realize data calculation, storage, processing, and sharing. The cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, and the like applied based on a cloud computing business model, can form a resource pool, and is used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background service of a technical network system needs a large amount of computing and storage resources, such as a video website, a picture website, and more portals. With the high development and application of the Internet industry, in the future, every item can have its own identification mark and needs to be transmitted to the background system for logical processing. Different levels of data will be processed separately, and various industry data need strong system support, which can only be realized through cloud computing.

[0048] The artificial intelligence cloud service is also commonly referred to as AIaaS (AI as a Service). This is a mainstream service mode of an artificial intelligence platform, specifically, the AIaaS platform splits several common AI services and provides independent or packaged services on the cloud. This service mode is similar to opening an AI theme mall: all developers can access one or more artificial intelligence services provided by the platform through an application programming interface (API), and some experienced developers can also use the AI framework and AI infrastructure provided by the platform to deploy and maintain exclusive cloud artificial intelligence services.

[0049] Figure 2 is an example schematic diagram showing an example schematic diagram 200 of a scenario of action migration and training based on a neural network model for action migration according to an embodiment of the present disclosure.

[0050] In the training phase, the server 110 can train the neural network model for action migration based on training samples (for example, image samples containing a role to be migrated). After training is completed, the server can deploy the trained neural network model for action migration to one or more servers (or on a cloud service) to provide artificial intelligence services related to action migration.

[0051] It is worth noting that all the images used in the present disclosure are legal, moral and private in accordance with the legal regulations. Specifically, the source of all the images is legal, and the explicit permission of the user has been obtained during the collection process. In addition, all the images used in the present disclosure comply with the privacy protection principle, and the images have been strictly screened and cleaned, and will not be disclosed to any third party without explicit authorization.

[0052] In the stage of action migration based on the neural network model for action migration, it is assumed that the user terminal 120 for action migration has installed a client or application (i.e., various action migration (e.g., character expression migration, character action migration, animal action migration, virtual image action migration, etc.) applications) interacting with the server 110 for action migration. The user terminal 120 can send an action migration request to the server 110 corresponding to the application through the network to request the neural network for action migration deployed on the server 110 to perform action migration. For example, after the server 110 receives the action migration request, the trained neural network model for action migration responds to the request to perform action migration processing, and feeds back the action migration result (e.g., provides the image or video after action migration) to the user terminal 120. The user terminal 120 can receive the action migration result. After that, the user terminal 120 can perform further analysis or processing based on the action migration result.

[0053] It is worth noting that, Figure 2 The training sample data shown in the above-mentioned can also be updated in real time. For example, the user can score the result of action migration. For example, if the user thinks that the rationality and accuracy of the action migration result are both high, the user can give a higher score to the action migration result, and the server 110 can take the action migration result as a positive sample of the neural network model for action migration. If the user gives a lower score to the action migration result, the server 110 can take the action migration result as a negative sample.

[0054] Figure 2 The training sample set shown in the above-mentioned can also be set in advance. For example, referring to Figure 2 The server can obtain the training sample from the database, and then generate the training sample set of the neural network model for action migration. Of course, the present disclosure is not limited thereto.

[0055] At present, for the action (or expression / pose) migration task, a method without a neural network model or a method based on a neural network model is usually used to realize it.

[0056] For example, in the case of transferring the expression of a real role to a 3D virtual role, a method without a neural network model usually needs to obtain a large number of expression-matched image pairs of a real role and a 3D virtual role in advance to establish a role expression correspondence database, and then in the case of transferring the expression of multiple images of a real role to a virtual role, the required image pairs can be sequentially retrieved from the role expression correspondence database to obtain the matched multiple images of the virtual role. This method has high dependence on the pre-established role expression correspondence database, needs to spend a lot of manpower to establish a perfect role expression correspondence database, and has poor universality (can only be used for a specific real role and a specific 3D virtual role).

[0057] A method based on a neural network model is usually designed for a specific 3D virtual role under a specific control rule. The neural network model is trained based on images of a real role and images of a specific 3D virtual role. The trained neural network model can predict the control parameters (which can be used in an animation production pipeline) of the specific 3D virtual role under the specific control rule based on the images of the real role, thereby transferring the expression of the real role to the specific 3D virtual role. However, since different platforms have different 3D virtual roles and different control rules (the control rules include predetermined rules for control degrees of freedom, physical meanings of control parameters, etc., for example, for a first control rule, there can be 10 control parameters for the mouth, respectively representing the deformation of different regions on the mouth and the changes of muscles in different regions around the mouth, where 1 represents the maximum deformation and 0 represents the minimum deformation; for a second control rule, there can be 50 control parameters for the mouth, respectively representing the deformation of different regions on the mouth and the changes of muscles in different regions around the mouth, which have a different region division manner from the first control rule, where 100 represents the maximum deformation and 0 represents the minimum deformation), a neural network model training method (and a neural network model architecture) needs to be specially designed to obtain the control parameters according to the characteristics, and therefore the training method is not applicable to different control rules or 3D virtual roles (for example, some training methods require that the control parameters must be derivable, different training methods have different requirements for input data, such as some training methods need images of a specific actor as model input). Further, if the expression of a real role is to be transferred to other 3D virtual roles under other control rules, control rule (i.e., rigging rule in the field of animation or film production) migration is also needed, and post-processing operations such as transfer of different images and matching of facial accessories (such as eyes and teeth) are also needed.

[0058] To solve the above problems, the present application proposes: using a neural network model to realize action (or expression) migration based on a source image (for example, a virtual character image) and a driving image (for example, a real character image) to generate a required target image (for example, a new virtual character image with the image of a virtual character in the source image and the action (or expression) of a real character in the driving image). The method of the present disclosure has high migration accuracy and can obtain the required new target image without manual adjustment of the virtual character processed by the neural network model.

[0059] The neural network model for action migration of the present disclosure can include two parts: a transformation prediction network and a transformation decoding network. The transformation prediction network is used to predict the transformation between the source image and the driving image, and the transformation decoding network is used to decode the transformation to obtain control parameters for controlling the action (including expression and posture) of the character (the control parameters can be used as binding parameters in the field of animation or film production pipeline). Based on the source image and the obtained control parameters, the required target image can be obtained by using simulation software to realize animation or film production.

[0060] To improve the prediction accuracy of the transformation prediction network, the present disclosure combines cross-reconstruction and self-reconstruction processes to train the transformation prediction network. That is, two image samples of the same category are respectively used as the source image and the driving image (for example, both the source image and the driving image are virtual character images; or both the source image and the driving image are real character images), and two image samples of different categories are respectively used as the source image and the driving image (for example, the source image is a virtual character image and the driving image is a real character image; or the source image is a real character image and the driving image is a virtual character image) to train the transformation prediction network. The cross-reconstruction and self-reconstruction processes of the present disclosure can also obtain good training results in the case of unsupervised learning, reducing the cost of manual annotation.

[0061] Then, after the training of the transformation prediction network is completed, the corresponding transformation decoding network can be trained for different control rules. The trained transformation decoding network can convert the transformation predicted by the transformation prediction network into control parameters that meet the control rule. Based on the control parameters, the required target image can be obtained, thereby realizing action (or expression / posture) migration.

[0062] The training method (and model architecture) of the present disclosure has universality and does not need to design a training process for each control rule. The neural network model for action migration (including the transformation prediction network and the transformation decoding network) obtained by training can obtain control parameters for controlling the character, which can be used for subsequent video or animation production, thereby simplifying the video or animation production process.

[0063] The following takes the example of migrating the expression of a real role to a 3D virtual role, combined with Figures 3-5B The process of training and action migration based on the neural network model for action migration is described. Figure 5A and Figure 5B The neural network model used in the foregoing can be trained through the process shown in Figures 3-4 . Figures 3-5B Corresponding reference signs and symbols in the foregoing can have corresponding relationships. In the training process, each real role image is from the following common data sets: multi-view affective audio-visual dataset (MEAD) and speaker recognition dataset (VoxCeleb).

[0064] Figure 3 The training process of the transformation prediction network according to an embodiment of the present disclosure is shown.

[0065] The training of the transformation prediction network includes two processes of self-reconstruction and cross-reconstruction, wherein in the self-reconstruction process, two image samples of the same category are respectively taken as a source image and a driving image to generate a target image; in the cross-reconstruction process, two image samples of different categories are respectively taken as a source image and a driving image to generate a target image. By calculating the value of the loss function based on the difference between the driving image and the target image, the transformation prediction network can be trained.

[0066] For the example of Figure 3 , in the self-reconstruction process, two 3D virtual role images A S and A D randomly selected are respectively taken as a source image and a driving image to generate a target image A R ; in the cross-reconstruction process, a real role image H S and a 3D virtual role image A D are respectively taken as a source image and a driving image to generate a target image H R ; further, a 3D virtual role image A S and a target image H R may also be respectively taken as a source image and a driving image to generate a target image A' R .

[0067] It should be understood that Figure 3 the process of training the transformation prediction network by taking two 3D virtual role images (A S and A D ) of the same role and a real role image H S as inputs of the transformation prediction network is also introduced, two real role images (H S and H D ) of the same role can also be taken as inputs of the transformation prediction network to train the transformation prediction network.) and a 3D virtual character image A S As the input of the transformation prediction network, the transformation prediction network is trained (the process is similar to Figure 3 Similar, no further description here).

[0068] It should be noted that for Figure 3 In the training process shown, during multiple training iterations, all 3D virtual character images are images of different actions of a specific character, while all real character images are images of different actions of multiple characters.

[0069] Specifically, for Figure 3 For example, image A S and image A D As source image and driving image respectively, to utilize the transformation prediction network Predicting from image A S To image A D Transformation between Right now, Next, we use the image generation network G I , based on image A S and transformation Generate image A R ,Right now,

[0070] Then, the image H can be S and image A D As source image and driving image respectively, to utilize the transformation prediction network Predict from image H S To image A D Transformation between Right now, Next, we use the image generation network G I , based on image H S and transformation Generate image H R ,Right now,

[0071] Furthermore, image A can be S and image H R As source image and driving image respectively, to utilize the transformation prediction network Predicting from image A S To image H R Transformation between Right now, Next, we use the image generation network G I , based on image A S and transformation Generate image A′R i.e.

[0072] By the above processing, it can be seen that the images A D , A R , H R and A' R all have consistent actions (including expressions and postures), wherein the images A D , A R and A' R are 3D virtual character images.

[0073] Therefore, the value of the loss function can be calculated by the following formula (1) to train the transformation prediction network.

[0074]

[0075] wherein the values of and can be calculated based on formula (2) and formula (3) respectively.

[0076]

[0077] wherein |||1 represents the minimum absolute value deviation, i.e., the L1 norm, here the minimum absolute value deviation between the features calculated after extracting the features from the images; represents the perceptual loss; represents the adversarial loss; λ p and λ g are predetermined coefficients.

[0078]

[0079] wherein the values of and can be calculated based on formula (4) and formula (5) respectively.

[0080]

[0081]

[0082] wherein |||2 represents the minimum square error, i.e., the L2 norm.

[0083] Similarly, for the process of taking the real character images (H S and H D ) of two same characters and a 3D virtual character image A S as the input of the transformation prediction network to train the transformation prediction network, the loss function can also be calculated in a similar manner

[0084] In the process of training the transformation prediction network, the total loss function may be represented by equation (6).

[0085]

[0086] wherein some batches of training samples are based on the loss function In the process of training the transformation prediction network, some batches of training samples are based on the loss function In the process of training the transformation prediction network.

[0087] In the process of training the transformation prediction network , the parameters of the image generation network G I may be optimized accordingly. Alternatively, for the image generation network G I with better performance, in the process of training the transformation prediction network , the parameters of the image generation network G I may be fixed to simplify the training process.

[0088] Finally, the trained transformation prediction network can accurately predict the transformation from the source image to the driving image, regardless of whether the source image and the driving image are real character (which can be different real characters) images or 3D virtual character (specific 3D virtual character) images.

[0089] The training process of the transformation decoding network (D R ) according to the embodiments of the present disclosure is described below with reference to Figure 4 .

[0090] As shown in FIG. 1, in order to obtain the control parameters (i.e., the binding parameters in the field of animation or film production) that can be used in the animation or film production pipeline by migrating the motion of the real character to the specific 3D virtual character, the corresponding control parameters P D = (e A , p A ) when the image A A changes from the standard image A0 (the standard image A0 is usually an image of the front face of the character without special expressions) of the specific 3D virtual character to the image A A may be obtained, wherein e A is the control parameter for controlling the expression of the character, p A is the control parameter for controlling the pose of the character, and e A and p D are known parameters.

[0091] Then, the image A0 and the image AD respectively as source and drive images to predict the transformation from image A0 to image H using the trained transformation prediction network D i.e. Next, the transformation R is decoded using the transformation decoding network D to obtain the predicted control parameters P D corresponding to the transformation of the standard image A0 to the 3D virtual character image A A ′ = (e′ A , p′ A ). The difference between the control parameters P A and the predicted control parameters P A ′ can be used to compute the value of the loss function to train the transformation decoding network. For example, the value of the loss function can be computed by equation (7).

[0092]

[0093] where λ pose is a predetermined coefficient.

[0094] Finally, the trained transformation decoding network can decode the transformation predicted by the transformation prediction network to obtain the control parameters that can be directly used in the animation or movie production pipeline to control the 3D virtual character to achieve fine animation or movie production.

[0095] Figure 5A is a schematic diagram showing a motion transfer process according to an embodiment of the present disclosure.

[0096] In the motion transfer process (i.e., the inference process of the neural network model for motion transfer) shown in Figure 5B , the standard image A0 of a specific 3D virtual character and the real character image H D are respectively taken as source and drive images to predict the transformation from image A0 to image H using the trained transformation prediction network D i.e. Next, the image generation network G I is used to generate the image D based on the transformation from image A0 to image H i.e. After this process, the 3D virtual character image A after motion (including expression) transfer can be obtained.​​

[0097] Figure 5B is a schematic diagram showing an action migration process according to another embodiment of the present disclosure.

[0098] In the action migration process (i.e., the inference process of the neural network model for action migration) shown, Figure 5B the standard image A0 of a specific 3D virtual character and the real character image H D are taken as the source image and the driving image respectively, so as to utilize the trained transformation prediction network to predict the transformation from the image A0 to the image H D . That is, Then, the transformation is decoded by utilizing the transformation decoding network D R , and the corresponding control parameter P predicted when the standard image A0 is changed to the real character image H D can be obtained, i.e. H = (e H , p H ). Then, based on the control parameter P H and the standard image A0, the computer can obtain the 3D virtual character image after action (including expression) migration by utilizing the simulation software.

[0099] Through experimental verification based on the common dataset multi-view emotion audio-visual dataset (MEAD) and the speaker recognition dataset (VoxCeleb), it is shown that the neural network model for action migration trained by combining the cross-reconstruction and self-reconstruction processes can accurately migrate the expressions and actions of the characters in each driving image (not limited to the training dataset) to the 3D virtual character targeted in the training process, without further post-processing or pre-establishing a perfect expression database of the characters. The cross-reconstruction and self-reconstruction processes of the present disclosure can also obtain good training results in the case of unsupervised learning, reducing the cost of manual labeling.

[0100] The training method (and model architecture) of the present disclosure is universal, and does not need to design a training process for each control rule (including the predetermined rules for the physical meaning of the control freedom and the control parameter). The neural network model for action migration (including the transformation prediction network and the transformation decoding network) trained can obtain the control parameter for controlling the character, for subsequent video or animation production, thereby simplifying the video or animation production process.

[0101] Figure 6Ais a schematic flowchart showing a method 600 of training a transformation prediction network according to an embodiment of the present disclosure.

[0102] In step S610, a first image sample, a second image sample and a third image sample are determined, wherein the first image sample contains a first character, the second image sample and the third image sample contain a second character, and the actions of the second character in the second image sample and the third image sample are different.

[0103] According to an embodiment of the present disclosure, the transformation prediction network is used to predict the transformation between different images for further realizing action (including behavior, expression, posture) transfer based on the predicted transformation.

[0104] The first character and the second character can be different types of characters. For example, the first character can be a real character, and the second character can be a virtual character (for example, the first character is a real character, and the second character is a 3D virtual character or a 2D virtual character); or the first character and the second character can be different types of virtual characters (for example, the first character is a 3D virtual character, and the second character is a 2D virtual character), etc. Alternatively, the real character and the virtual character can be a human character or an animal character.

[0105] According to an embodiment of the present disclosure, in order to improve the transformation prediction accuracy of the trained transformation prediction network, the types of the first character and the second character can also be randomly exchanged. For example, in the case where the transformation prediction network is used to predict the transformation between real character images and virtual character images, the real character can be randomly taken as the first character, and the virtual character can be randomly taken as the second character, or the virtual character can be randomly taken as the first character, and the real character can be randomly taken as the second character.

[0106] In step S620, the transformation prediction network is trained through a first training process, wherein the first training process trains the transformation prediction network based on the first image sample, the second image sample and the third image sample.

[0107] It should be understood that this training process can not need to input the label corresponding to each image sample, and can also obtain good training effect for the transformation prediction network in the case of unsupervised learning, reducing the cost of manual annotation.

[0108] For the first training process of step S620, the specific training process can include steps S621-S624 as shown in Figure 6B

[0109] ​In step S621, a first transform between the first image sample and the second image sample is predicted by using the transform prediction network (i.e., the first image sample is processed by the first transform to obtain the second image sample).

[0110] In step S622, a first target image is generated based on the first image sample and the first transform by using the image generation network, where a character in the first target image is consistent with a character in the first image sample and has a same action as a character in the second image sample.

[0111] The processing of steps S621 and S622 is equivalent to taking the first image sample as a source image and the second image sample as a driving image to generate the first target image with action (including behavior, expression, and posture) migration.

[0112] In step S623, a second transform between the third image sample and the first target image is predicted by using the transform prediction network (i.e., the third image sample is processed by the second transform to obtain the first target image).

[0113] In step S624, a second target image is generated based on the third image sample and the second transform by using the image generation network, where a character in the second target image is consistent with a character in the third image sample and has a same action as a character in the first target image.

[0114] The processing of steps S623 and S624 is equivalent to taking the third image sample as a source image and the first target image as a driving image to generate the second target image with action (including behavior, expression, and posture) migration.

[0115] In step S625, a value of a first loss function is calculated based on a difference between the second image sample and the second target image to train the transform prediction network.

[0116] According to an embodiment of the present disclosure, during the training of the transform prediction network by the first training process, the parameters of the image generation network can be updated and optimized at the same time. When the performance of the image generation network is good enough, the parameters of the image generation network can also be fixed to simplify the training process.

[0117] It should be understood that, after the processing of steps S621-S624, the second image sample and the second target image actually both contain a same second character, and the actions of the second character in the second image sample and the second target image are the same.

[0118] According to an embodiment of the present disclosure, the value of the first loss function can be determined based on the minimum absolute value deviation (i.e., L1 norm) between the second image sample and the second target image. Optionally, the value of the first loss function can also be determined additionally based on at least one of the perceptual loss and the adversarial loss between the second image sample and the second target image (for example, by summing or weighted summing (not limited to this calculation manner, other calculation manners can also be used) the minimum absolute value deviation and the perceptual loss between the second image sample and the second target image; by summing or weighted summing the minimum absolute value deviation, the perceptual loss, and the adversarial loss between the second image sample and the second target image, etc.).

[0119] Figure 6B The first training process shown is actually a process of training based on image samples of different types of roles (i.e., cross-reconstruction process). Optionally, the method 600 can further include Figure 6C The second training process 630 shown is a process of training based on image samples of the same type of role (i.e., self-reconstruction process).

[0120] In step S631, a third transformation from the third image sample to the second image sample is predicted by using the transformation prediction network (i.e., the second image sample can be obtained after the third image sample is processed by the third transformation).

[0121] In step S632, a third target image is generated based on the third image sample and the third transformation by using the image generation network, wherein the role in the third target image is consistent with the role in the third image sample, and has the same action as the role in the second image sample.

[0122] The processing procedures of steps S631 and S632 are equivalent to taking the third image sample as a source image and the second image sample as a driving image to generate the third target image after the action (including behavior, expression, and posture) migration.

[0123] In step S633, the value of a second loss function is calculated based on the difference between the second image sample and the third target image to train the transformation prediction network.

[0124] According to an embodiment of the present disclosure, during the training of the transformation prediction network by the second training process, the parameters of the image generation network can be updated and optimized at the same time. When the performance of the image generation network is good enough, the parameters of the image generation network can also be fixed to simplify the training process.

[0125] According to an embodiment of the present disclosure, the value of the second loss function can be determined based on a minimum absolute value deviation (i.e., L1 norm) between the second image sample and the third target image. Optionally, the value of the second loss function can also be determined additionally based on at least one of a perceptual loss and an adversarial loss between the second image sample and the third target image (e.g., by summing or weighted summing (not limited to this calculation manner, other calculation manners can also be used) the minimum absolute value deviation and the perceptual loss between the second image sample and the third target image; by summing or weighted summing the minimum absolute value deviation, the perceptual loss, and the adversarial loss between the second image sample and the third target image, etc.).

[0126] According to an embodiment of the present disclosure, the transform prediction network can be trained based on the first training process and the second training process simultaneously (i.e., training considering both self-reconstruction process and cross-reconstruction process); or the transform prediction network can be trained based on the second training process first and then based on the first training process (i.e., training considering self-reconstruction process first and then cross-reconstruction process, or training considering both self-reconstruction process and cross-reconstruction process). In the case of training the transform prediction network based on the first training process and the second training process simultaneously, a fourth loss function value can be determined based on the value of the first loss function and the value of the third loss function to train the transform prediction network (e.g., by summing or weighted summing (not limited to this calculation manner, other calculation manners can also be used) the value of the first loss function and the value of the third loss function).

[0127] According to an embodiment of the present disclosure, the second role can be a specific virtual role, and the first role can be selected from a plurality of real roles, and the first training process or the second training process is performed respectively for the plurality of real roles. In this way, the trained transform prediction network can be used to predict the transformation between various different real roles to the specific virtual role, so as to further realize the action (including behavior, expression, posture) migration from various different real roles to the specific virtual role based on the predicted transformation.

[0128] According to an embodiment of the present disclosure, a third loss function value can also be calculated based on the difference between the second transformation and the third transformation (e.g., the third loss function value can be determined based on the minimum square error (i.e., L2 norm) between the second transformation and the third transformation); and the transform prediction network can be trained based on the third loss function value, or the transform prediction network can be trained based on the third loss function value and the first loss function value.

[0129] Figure 6D is a schematic flowchart illustrating a method of training a transformation prediction network. Figures 6A-6C Based on the method of training the transformation prediction network shown in FIG. 6, a method 640 of further training a transformation decoding network is shown in a schematic flowchart. The transformation decoding network is used to decode the transformation predicted by the transformation prediction network to obtain control parameters for controlling a specific virtual role for subsequent video or animation production, thereby simplifying the video or animation production process.

[0130] In step S641, a fourth image sample containing the second role is obtained, wherein the first role is selected from a plurality of real roles, and the second role is a specific virtual role. The specific virtual role can be consistent with the specific virtual role selected when training the transformation prediction network.

[0131] In step S642, the control parameters corresponding to the change from the standard image of the specific role to the fourth image sample are obtained.

[0132] It should be understood that the control parameters here are known parameters, which are equivalent to sample labels, and the training process of the transformation prediction network is supervised learning. The standard image of the specific role is a predetermined standard image. For convenience of processing, the standard image of the specific role is usually an image in which the specific role is facing forward, without special gestures or expressions.

[0133] In step S643, the transformation prediction network is used to predict a fourth transformation between the standard image of the specific role and the fourth image sample (i.e., the standard image of the specific role can obtain the fourth image sample after being processed by the fourth transformation).

[0134] In step S644, the transformation decoding network is used to decode the fourth transformation to obtain predicted control parameters.

[0135] In step S645, the value of a fourth loss function is calculated based on the difference between the control parameters and the predicted control parameters to train the transformation decoding network.

[0136] According to an embodiment of the present disclosure, the value of the first loss function can be determined based on the minimum absolute value deviation (L1 norm) between the control parameters and the predicted control parameters.

[0137] Figure 7A is a schematic flowchart illustrating a method 700 of action migration according to an embodiment of the present disclosure.

[0138] In step S710, a first image containing a first role and a second image containing a second role are obtained.

[0139] According to an embodiment of the present disclosure, the first role and the second role can be different types of roles. For example, the first role can be a real role and the second role can be a virtual role (e.g., the first role is a real role and the second role is a 3D virtual role or a 2D virtual role); or the first role and the second role can be different types of virtual roles (e.g., the first role is a 3D virtual role and the second role is a 2D virtual role), etc. Optionally, the real role and the virtual role can be a human role or an animal role. That is, the action migration method 700 can be used to migrate the action of a human role (e.g., used to migrate the expression of a real human to a virtual role, migrate the expression of a 3D virtual human to a 2D virtual human, etc.) or an animal role (e.g., used to migrate the action of a real animal to a corresponding virtual animal, migrate the expression of a 3D virtual cartoon image to a 2D virtual cartoon image, etc.).

[0140] In step S720, a transformation prediction network is used to predict a prediction transformation from the first image to the second image.

[0141] In step S730, a third image is generated based on the first image and the prediction transformation, wherein the role in the third image is consistent with the role in the first image and has the action consistent with the role in the second image.

[0142] The processing procedures of steps S720 and S730 are equivalent to taking the first image as a source image and the second image as a driving image to generate the third image after action (including behavior, expression, and posture) migration.

[0143] According to an embodiment of the present disclosure, an image generation network can be used to generate the third image based on the first image and the prediction transformation (e.g., the process as shown in Figure 5A , which can be implemented completely based on a neural network model.

[0144] According to an embodiment of the present disclosure, the prediction transformation can also be decoded by a transformation decoding network to obtain a prediction control parameter, and the third image can be generated based on the first image and the prediction control parameter (e.g., the process as shown in Figure 5B , which can be implemented based on simulation software (e.g., animation production software, movie production software). The advantage of this embodiment over the generation of the third image by the image generation network is that the prediction control parameter can be directly used in the animation or movie production pipeline to enable animators or filmmakers to achieve more delicate animation or movie production.

[0145] It should be understood that the actions disclosed herein are all broadly defined and may include at least one of behavior, posture, and expression. Similarly, the predicted control parameters may also include at least one of the following: a parameter for controlling the behavior of the second character, a parameter for controlling the expression of the second character, and a parameter for controlling the posture of the second character.

[0146] It should be understood that according to the embodiments of the present disclosure, the driving image (i.e., the second image) can be used as input to achieve the migration of the action of the character in the driving image to the character in the first image; or the driving video can be used as input to achieve the migration of the action of the character in the driving video to the character in the first image (for example, the driving video can be decomposed into multiple frames of images to serve as the second image respectively to execute method 700, and the generated multiple frames of third images can be synthesized into the desired target video).

[0147] Figure 7A The transform prediction network, the image generation network, and the transform decoding network used in the embodiment can be Figures 6A-6D The method shown is trained.

[0148] When the real character image is used as the driving image and the expression in the real character image is transferred to the virtual character, Figure 7B 、 Figure 7C and Table 1 shows the utilization Figure 7A The action transfer method shown in the figure is experimentally tested on the shared datasets Multi-view Emotion Audiovisual Dataset (MEAD) and Speaker Recognition Dataset (VoxCeleb). Figure 7B It shows Figure 7A A comparison of the effects of the motion transfer method shown here and the face tracking method developed by Apple (hereinafter referred to as Apple), wherein the Apple method is based on the control rules (binding rules) of the augmented reality development component (ARKit); Figure 7C It shows Figure 7A The comparison diagram of the effect of the motion transfer method and the 3D face reconstruction method (the method is referred to as DECA), wherein the DECA method is based on the FLAME control rule (binding rule); Table 1 shows the scores of a large number of users on the generated images after motion transfer (the full score is 5 points).

[0149] Table 1

[0150]

[0151] pass Figure 7B 、 Figure 7CAs can be seen from Table 1, the action migration method of the present disclosure has a better action migration effect than the ARKit-based method and the DECA-based method, and can more accurately migrate the expression of the real character to the virtual character. Moreover, the training method (and model architecture) of the present disclosure has universality, and does not need to design a training process for each control rule (including a predetermined rule for the physical meaning of the control degree of freedom and the control parameter), further making the action migration method of the present disclosure applicable to characters under various control rules.

[0152] Figure 8 FIG. 8 is a constituent schematic diagram illustrating an apparatus 800 for training a neural network model according to an embodiment of the present disclosure.

[0153] According to an embodiment of the present disclosure, the apparatus 800 for training a neural network model can include a sample determination module 810 and a model training module 820.

[0154] The sample determination module 810 can be configured to determine a first image sample, a second image sample, and a third image sample, wherein the first image sample contains a first character, the second image sample and the third image sample contain a second character, and the actions of the second character in the second image sample and the third image sample are different.

[0155] The model training module 820 can be configured to train the transformation prediction network through a first training process, wherein the first training process trains the transformation prediction network based on the first image sample, the second image sample, and the third image sample, wherein the first training process includes: predicting a first transformation from the first image sample to the second image sample by using the transformation prediction network; generating a first target image based on the first image sample and the first transformation by using an image generation network, wherein the character in the first target image is consistent with the character in the first image sample, and has an action consistent with the character in the second image sample; predicting a second transformation from the third image sample to the first target image by using the transformation prediction network; generating a second target image based on the third image sample and the second transformation by using the image generation network, wherein the character in the second target image is consistent with the character in the third image sample, and has an action consistent with the character in the first target image; and calculating a value of a first loss function based on the difference between the second image sample and the second target image to train the transformation prediction network.

[0156] According to an embodiment of the present disclosure, the model training module 820 can be further configured to train the transform prediction network through a second training process, wherein the second training process is based on the second image sample and the third image sample to train the transform prediction network, and wherein the second training process comprises: predicting a third transform between the third image sample and the second image sample by using the transform prediction network; generating a third target image based on the third image sample and the third transform by using the image generation network, wherein the character in the third target image is consistent with the character in the third image sample and has the same action as the character in the second image sample; and calculating a value of a second loss function based on a difference between the second image sample and the third target image to train the transform prediction network.

[0157] According to an embodiment of the present disclosure, the model training module 820 can be further configured to obtain a fourth image sample containing the second character, wherein the first character is selected from a plurality of real characters, and the second character is a specific virtual character; obtain a control parameter corresponding to a change from a standard image of the specific character to the fourth image sample; predict a fourth transform between the standard image of the specific character and the fourth image sample by using the transform prediction network; decode the fourth transform by using the transform decoding network to obtain a predicted control parameter; and calculate a value of a fourth loss function based on a difference between the control parameter and the predicted control parameter to train the transform decoding network.

[0158] It should be understood that, Figure 8 The apparatus 800 for training a neural network model shown can implement various methods of training a neural network model as described with respect to Figures 6A-6C The sample determination module 810 can be used to implement the processing procedure of step S610, and the model training module 820 can be used to implement the processing procedures of steps S621-S625, steps S631-S633, and steps S641-S645, which are not described herein again. By Figure 8 The apparatus 800 for training a neural network model shown can train a neural network model to obtain a transform prediction network in Figure 7A Optionally, the apparatus 800 for training a neural network model shown can train a neural network model to further obtain a trained image generation network and / or a trained transform decoding network. Figure 8

[0159] Figure 9 is a constituent schematic diagram showing an apparatus 900 for action migration according to an embodiment of the present disclosure.

[0160] ​According to an embodiment of the present disclosure, the apparatus 900 for motion migration may include: an image acquisition module 910 , a transformation prediction module 920 , and an image generation module 930 .

[0161] The image acquisition module 910 may be configured to acquire a first image containing a first character and a second image containing a second character. The transformation prediction module 920 may be configured to use a transformation prediction network to predict a predicted transformation from the first image to the second image. The image generation module 930 may be configured to generate a third image based on the first image and the predicted transformation, wherein the character in the third image is consistent with the character in the first image and has the same action as the character in the second image.

[0162] It should be understood that Figure 9 The device 900 for action migration shown in FIG. Figure 7A The various methods for motion migration described above, including the image acquisition module 910 , the transformation prediction module 920 , and the image generation module 930 , can be used to implement the processing of steps S710 , S720 , and S730 , respectively, and are not described in detail here. Figure 9 The neural network model used in the embodiment can be obtained by Figures 6A-6C The device 900 for action migration can be located at Figure 1 The server 110 shown may also be located at Figure 1 The device 900 for action transfer can take a driving image (i.e., the second image) as input to transfer the action of the character in the driving image to the character in the first image; or can take a driving video as input to transfer the action of the character in the driving video to the character in the first image.

[0163] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0164] For example, the method or apparatus according to the embodiment of the present disclosure may also be implemented by Figure 10 The architecture of the computing device 3000 shown in FIG. Figure 10As shown, computing device 3000 can include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port connected to a network 3050, an input / output component 3060, a hard disk 3070, and the like. The storage devices in computing device 3000, such as ROM 3030 or hard disk 3070, can store various data or files used by the processes and / or communications of the methods provided by the present disclosure, as well as the program instructions to be executed by the CPUs. Computing device 3000 can also include a user interface 3080. Of course, Figure 10 The illustrated architecture is exemplary only, and in implementing different devices, components shown can be omitted, components shown can be duplicated, and / or different components can be utilized. Figure 10 One or more components of the computing device are shown.

[0165] According to yet another aspect of the present disclosure, a computer-readable storage medium is also provided. The computer-readable storage medium has computer-readable instructions stored thereon. When the computer-readable instructions are run by a processor, the method according to the embodiments of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiments of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct

[0166] Embodiments of the present disclosure also provide an information processing device, comprising: a memory and a processor coupled to the memory and configured to execute a method according to embodiments of the present disclosure.

[0167] Embodiments of the present disclosure also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method according to an embodiment of the present disclosure.

[0168] In summary, the embodiments of the present disclosure provide methods, apparatuses, information processing devices, computer program products, and storage media for training neural network models, as well as methods, apparatuses, computer program products, and storage media for action migration.

[0169] The method for training a neural network model disclosed herein includes: determining a first image sample, a second image sample, and a third image sample, wherein the first image sample contains a first character, the second image sample and the third image sample contain a second character, and the actions of the second character in the second image sample and the third image sample are different; training a transformation prediction network through a first training process, wherein the first training process trains the transformation prediction network based on the first image sample, the second image sample, and the third image sample, wherein the first training process includes: using the transformation prediction network to predict a first transformation from the first image sample to the second image sample; using an image generation network based on the first image sample, the second image sample, and the third image sample; an image sample and the first transformation, generating a first target image, wherein the character in the first target image is consistent with the character in the first image sample and has the same action as the character in the second image sample; using the transformation prediction network to predict a second transformation from the third image sample to the first target image; using the image generation network to generate a second target image based on the third image sample and the second transformation, wherein the character in the second target image is consistent with the character in the third image sample and has the same action as the character in the first target image; and calculating the value of a first loss function based on the difference between the second image sample and the second target image to train the transformation prediction network.

[0170] The method of the present disclosure can enable the trained transformation prediction network to accurately implement transformation prediction, so as to further implement action migration based on the trained transformation prediction network. The method of the present disclosure can fully utilize the neural network model to implement action (including behavior, expression, and posture) migration to generate a required target image without further post-processing or pre-establishment of a perfect role expression corresponding relationship database. The method of the present disclosure can also obtain good training effect for the transformation prediction network in the case of unsupervised learning, reducing the cost of manual annotation. The training method (and model architecture) of the present disclosure has universality and does not need to design a training process for each control rule. The neural network model (including the transformation prediction network and the transformation decoding network) for action migration obtained by training can obtain control parameters for controlling a role, for subsequent video or animation production, thereby simplifying the video or animation production process.

[0171] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectural, functional, and operational scenarios of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a portion of code that includes at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the accompanying drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and they can also be executed in reverse order, depending on the involved functions. It should also be noted that each block in the block diagrams and / or flowcharts, and a combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0172] The present disclosure uses certain terms to describe embodiments of the present disclosure. As used in the description of the disclosure and the appended claims, the terms "first / second embodiments", "an embodiment", and / or "some embodiments" mean that a certain feature, structure, or characteristic described is a part of at least one embodiment. It is also noted that the features, structures, or characteristics of one or more embodiments of the present disclosure can be combined in any suitable manner.

[0173] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that contains the functions of the module or unit.

[0174] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0175] The above is a description of the present disclosure and should not be considered limiting. Although several exemplary embodiments of the present disclosure are described, those skilled in the art will readily understand that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the present disclosure. Accordingly, all such modifications are intended to be included within the scope of the present disclosure defined by the claims. It should be understood that the above is a description of the present disclosure and should not be considered limiting to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the claims. The present disclosure is defined by the claims and their equivalents.

Claims

1. A method for training a neural network model, comprising: determining a first image sample, a second image sample, and a third image sample, wherein the first image sample contains a first character, the second image sample and the third image sample contain a second character, and actions of the second character in the second image sample and the third image sample are different; training a transformation prediction network by a first training process, wherein the first training process trains the transformation prediction network based on the first image sample, the second image sample, and the third image sample, wherein the first training process comprises: predicting, with the transformation prediction network, a first transformation from the first image sample to the second image sample; generating, with an image generation network, a first target image based on the first image sample and the first transformation, wherein a character in the first target image is consistent with a character in the first image sample and has an action consistent with a character in the second image sample; predicting, with the transformation prediction network, a second transformation from the third image sample to the first target image; generating, with the image generation network, a second target image based on the third image sample and the second transformation, wherein a character in the second target image is consistent with a character in the third image sample and has an action consistent with a character in the first target image; and computing a value of a first loss function based on a difference between the second image sample and the second target image to train the transformation prediction network.

2. The method of claim 1, wherein, Computing the value of the first loss function based on the difference between the second image sample and the second target image comprises: determining the value of the first loss function based on a least absolute deviation between the second image sample and the second target image.

3. The method of claim 2, wherein, Computing the value of the first loss function based on the difference between the second image sample and the second target image further comprises: determining the value of the first loss function based on at least one of a perceptual loss and an adversarial loss between the second image sample and the second target image.

4. The method of claim 1, further comprising: training the transformation prediction network by a second training process, wherein the second training process trains the transformation prediction network based on the second image sample and the third image sample, wherein the second training process comprises: predicting, with the transformation prediction network, a third transformation from the third image sample to the second image sample; generating, with the image generation network, a third target image based on the third image sample and the third transformation, wherein a character in the third target image is consistent with a character in the third image sample and has an action consistent with a character in the second image sample; and computing a value of a second loss function based on a difference between the second image sample and the third target image to train the transformation prediction network.

5. The method of claim 4, wherein, Computing the value of the second loss function based on the difference between the second image sample and the third target image comprises: determining the value of the second loss function based on at least one of a perceptual loss and an adversarial loss between the second image sample and the third target image.

6. The method of claim 5, wherein, determining the value of the second loss function based on at least one of a perceptual loss and an adversarial loss between the second image sample and the third target image. determining the value of the second loss function based on at least one of a perceptual loss and an adversarial loss between the second image sample and the third target image.

7. The method of claim 4, wherein, training the transform prediction network based on the first training process and the second training process simultaneously; or training the transform prediction network based on the second training process first and then training the transform prediction network based on the first training process.

8. The method of claim 7, wherein, In the case of training the transform prediction network based on the first training process and the second training process simultaneously, determining a value of a fourth loss function based on the value of the first loss function and the value of the third loss function to train the transform prediction network.

9. The method of claim 7, wherein, the second role is a specific virtual role, the first role is selected from a plurality of real roles, and the first training process or the second training process is performed respectively for the plurality of real roles.

10. The method of claim 4, further comprising: calculating a value of a third loss function based on a difference between the second transform and the third transform; and training the transform prediction network based on the value of the third loss function.

11. The method of claim 10, wherein, calculating a value of a third loss function based on a difference between the second transform and the third transform includes: determining the value of the third loss function based on a least square error between the second transform and the third transform.

12. The method of claim 1, wherein, the real role and the virtual role are a character role or an animal role.

13. The method of claim 1, wherein, the first role is selected from a plurality of real roles, and the second role is a specific virtual role, the method further comprising: obtaining a fourth image sample containing the second role; obtaining a control parameter corresponding to a change from a standard image of the specific role to the fourth image sample; predicting a fourth transform between the standard image of the specific role and the fourth image sample by using the transform prediction network; decoding the fourth transform by using a transform decoding network to obtain a predicted control parameter; calculating a value of a fourth loss function based on a difference between the control parameter and the predicted control parameter to train the transform decoding network.

14. The method of claim 13, wherein, calculating a value of a fourth loss function based on a difference between the control parameter and the predicted control parameter includes: determining the value of the first loss function based on a least absolute value deviation between the control parameter and the predicted control parameter.

15. A method for action transfer, comprising: obtaining a first image containing a first role and a second image containing a second role; predicting a prediction transform from the first image to the second image using a transform prediction network; generating a third image based on the first image and the prediction transform, wherein a character in the third image is consistent with a character in the first image and has an action consistent with a character in the second image.

16. The method of claim 15, wherein, Generating a third image based on the first image and the prediction transform comprises: generating the third image based on the first image and the prediction transform using an image generation network; or decoding the prediction transform using a transform decoding network to obtain a prediction control parameter, and generating the third image based on the first image and the prediction control parameter.

17. The method of claim 16, wherein, In the case of generating the third image based on the first image and the prediction control parameter, the prediction control parameter comprises at least one of: a parameter for controlling a behavior of the second character, a parameter for controlling an expression of the second character, and a parameter for controlling a pose of the second character.

18. An information processing apparatus comprising: a memory and a processor coupled to the memory and configured to perform the method of any one of claims 1-17.

19. A computer program product comprising computer software code which, when executed using a processor, is adapted to implement the method of any one of claims 1-17.

20. A computer readable storage medium having stored thereon computer executable instructions adapted to implement the method of any one of claims 1-17 when executed using a processor.