A Human Motion Style Transfer Method, System and Storable Medium
By disassembling and integrating semantic information and style information, the generation of adversarial networks is carried out for adversarial training, which solves the problems of single style transfer effect and low content preservation in the existing technology, and realizes the generation and content retention of fine-grained stylized actions.
Patent Information
- Application Number
- CN202311019036.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-08-14
AI Technical Summary
The existing sports style transfer method has a single style transfer effect and low content preservation, which ignores the semantic prior knowledge related to the content, resulting in the generated action style lacking the style characteristics related to the content.
By obtaining semantic information of input content actions and style actions, using a fully connected neural network and a convolutional neural network to disassemble data, combining the generated adversarial network for data integration and adversarial training, generating fine-grained stylized actions, fusing semantic latent code and style latent code, and adjusting data distribution using an adaptive instance normalization layer.
The movement after style transfer is realized with a fine-grained style expression with semantic guidance, while retaining the content information of the action to the greatest extent. The experimental results show that it is better than the existing methods.
Smart Images

Figure CN117036550B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer animation and virtual reality technologies, and particularly to a human motion style transfer method, system, and storage medium. Background Art
[0002] Motion style transfer is of great benefit to the virtual reality and computer animation industries. Animators usually seek to create stylized motions to represent the character or emotion of a character, thereby making the character more vivid. However, traditional motion capture methods rely on a huge amount of manual labor and specific equipment. To generate realistic animations, a large number of studies have proposed motion style transfer methods driven by the captured motion style. The main problem is how to effectively transfer the style while maximizing the preservation of the content of the motion.
[0003] In recent years, due to the rise of deep learning, some methods use neural networks to separately separate the content and style of a motion and then combine them to achieve style transfer. There are various ways of this combination. A relatively effective one is to change the style of the motion by changing the data distribution of the content latent code, that is, to use adaptive instance normalization (AdaIN) weighted by the style to adjust the data distribution of the content latent code, thereby achieving motion style transfer. And there are also many ways of encoding motion, such as one-dimensional temporal convolutional neural network and graph convolutional neural network. The main difference between the two is that the graph convolutional neural network can preserve the temporal and spatial information of the motion, but correspondingly its network cost is also very high. However, although these data-driven motion style transfers have made progress for many years, there are still some problems in practice, mainly manifested in that the style transfer effect of the existing methods is still single in style and the content preservation degree is not high. The reasons for these problems are mainly that these methods can only obtain a single style from the style motion, resulting in the generated motion style lacking style features related to the content of the content motion, thus showing a poor style transfer effect. Moreover, they ignore the use of semantic prior knowledge related to the content.
[0004] In view of this, it is an urgent problem for those skilled in the art to propose a human motion style transfer method, system, and storage medium to solve the difficulties existing in the prior art. Summary of the Invention
[0005] In view of this, the present invention provides a human motion style transfer method, system, and storage medium, which can effectively transfer the style and maximize the preservation of the content of the motion.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A human motion style transfer method includes the following steps:
[0008] Data acquisition step: Obtain the input content action and the style action;
[0009] Data disassembling step: Vectorize the semantic information of the input content action and input it into a fully connected neural network to obtain a semantic latent code, input the input style action into a convolutional neural network to obtain a style latent code, and input the input content action into the encoding network of the generation module of the generative adversarial network to obtain an intermediate content latent code;
[0010] Data integration step: Fuse the semantic latent code and the style latent code together to obtain a semantic-aware style latent code containing specific semantic content information, and input the semantic-aware style latent code and the content latent code into an Adaptive Instance Normalization (AdaIN) layer to change the data distribution of the content latent code and obtain a stylized content latent code;
[0011] Data decoding step: Input the stylized content latent code into the decoding network of the generation module of the generative adversarial network to decode the action;
[0012] Adversarial generation step: The adversarial module of the generative adversarial network discriminates the style of the generated action to perform adversarial training on the network, and obtains a stylized action with fine-grained style guided by semantics.
[0013] For the above method, optionally, in the data disassembling step, the semantic information of the input content action is vectorized through embedding and then input into a fully connected neural network to obtain a semantic latent code containing content features.
[0014] For the above method, optionally, in the data disassembling step, the semantic information refers to the action content label in the input content action.
[0015] For the above method, optionally, in the data disassembling step, the generation module of the generative adversarial network is composed of an encoder and a decoder, and both the encoder and the decoder are composed of convolutional neural networks.
[0016] For the above method, optionally, in the data integration step, the formula for fusing the semantic latent code and the style latent code is as follows:
[0017]
[0018] Among them, sig() represents the sigmoid activation function; represents the semantic latent code; z s represents the style latent code; z cs represents the obtained semantic-aware style latent code.
[0019] In the above method, optionally, in the data integration step, the data distribution is the mean and variance of the data, and the method implemented by the AdaIN layer is as follows:
[0020]
[0021] where z c represents the content latent code; z cs represents the semantic-aware style latent code; μ and σ respectively represent the mean and variance of the data.
[0022] In the above method, optionally, in the adversarial generation step, the adversarial module of the generative adversarial network consists of a discriminator, which is used to discriminate the style category of the generated action
[0023] In the above method, optionally, in the adversarial generation step, the training output action of the generative adversarial network includes the style of the input style action and the content of the input content action, and has a fine-grained style guided by the input semantic information.
[0024] A human motion style automatic migration system that executes a human action style migration method according to any one of the above, including a data acquisition module, a data disassembling module, a data integration module, a data decoding module, and an adversarial generation module;
[0025] Data acquisition module: Acquire the input content action and style action;
[0026] Data disassembling module: Vectorize the semantic information of the input content action and input it into a fully connected neural network to obtain the semantic latent code, input the input style action into a convolutional neural network to obtain the style latent code, and input the input content action into the encoding network of the generation module of the generative adversarial network to obtain the intermediate content latent code;
[0027] Data integration module: Fuse the semantic latent code and the style latent code together to obtain a semantic-aware style latent code containing specific semantic content information, input the semantic-aware style latent code and the content latent code into the adaptive instance normalization AdaIN layer, change the data distribution of the content latent code, and obtain the stylized content latent code;
[0028] Data decoding module: Input the stylized content latent code into the decoding network of the generation module of the generative adversarial network to decode the action;
[0029] Adversarial generation module: The adversarial module of the generative adversarial network discriminates the style of the generated action to perform adversarial training on the network, and obtains a stylized action with a fine-grained style guided by semantics.
[0030] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the human motion style transfer method of any one of the above are implemented.
[0031] As can be seen from the above technical solutions, the present invention provides a human motion style transfer method, system and storage medium. Compared with the prior art, it has the following beneficial effects:
[0032] (1) By inputting the style motion, content motion and semantic information of the content motion, a fine-grained stylized motion with semantic guidance after style transfer is finally obtained;
[0033] (2) Due to the addition of the semantic information of the motion content, the obtained stylized motion can greatly retain the content information of the motion;
[0034] (3) Extensive experimental results and comprehensive evaluations confirm that our motion style transfer method can well transfer the style of the motion while retaining the content of the motion. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0036] Figure 1 It is a flowchart of a human motion style transfer method provided by the present invention;
[0037] Figure 2 It is a schematic diagram of a stylized motion provided by the present invention;
[0038] Figure 3 It is a system structure diagram of a human motion style transfer provided by the present invention. Detailed Embodiments
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0040] Referring to Figure 1 as shown, the present invention discloses a human motion style transfer method, including the following steps:
[0041] Data acquisition step: Obtain the input content action and the style action;
[0042] Data decomposition step: Vectorize the semantic information of the input content action and input it into a fully connected neural network to obtain a semantic latent code. Input the input style action into a convolutional neural network to obtain a style latent code. Input the input content action into the encoding network of the generation module of the generative adversarial network to obtain an intermediate content latent code;
[0043] Data integration step: Fuse the semantic latent code and the style latent code together to obtain a semantic-aware style latent code containing specific semantic content information. Input the semantic-aware style latent code and the content latent code into the Adaptive Instance Normalization (AdaIN) layer to change the data distribution of the content latent code and obtain a stylized content latent code;
[0044] Data decoding step: Input the stylized content latent code into the decoding network of the generation module of the generative adversarial network to decode the action;
[0045] Adversarial generation step: The adversarial module of the generative adversarial network discriminates the style of the generated action to perform adversarial training on the network, and obtains a stylized action with fine-grained style guided by semantics.
[0046] Furthermore, in the data decomposition step, the semantic information of the input content action is vectorized through embedding and then input into a fully connected neural network to obtain a semantic latent code containing content features.
[0047] Furthermore, in the data decomposition step, the semantic information refers to the action content label in the input content action.
[0048] Furthermore, in the data decomposition step, the generation module of the generative adversarial network consists of an encoder and a decoder, and both the encoder and the decoder are composed of convolutional neural networks.
[0049] Furthermore, in the data integration step, the formula for fusing the semantic latent code and the style latent code is as follows:
[0050]
[0051] Among them, sig() represents the sigmoid activation function; represents the semantic latent code; z s represents the style latent code; z cs represents the obtained semantic-aware style latent code.
[0052] Furthermore, in the data integration step, the data distribution is the mean and variance of the data, and the method implemented by the AdaIN layer is as follows:
[0053]
[0054] Among them, z c represents the potential code of the content; z cs represents the potential code of the semantic perception style; μ and σ respectively represent the mean and variance of the data.
[0055] Furthermore, in the adversarial generation step, the adversarial module of the generative adversarial network consists of a discriminator, which is used to discriminate the style category of the generated action
[0056] Furthermore, in the adversarial generation step, the training output action of the generative adversarial network contains the style of the input style action and the content of the input content action, and has a fine-grained style guided by the input semantic information.
[0057] Specifically, referring to Figure 2 as shown, the stylized action is finally reconstructed, the content potential code after style transfer is decoded, and the stylized content potential code is input into the decoding network of the generation module of the generative adversarial network to decode the action. This decoding network is composed of a convolutional neural network. Subsequently, the adversarial module of the generative adversarial network discriminates the style of the generated action to perform adversarial training on the network. The adversarial module consists of a discriminator, which is used to discriminate the style category of the generated action. Finally, since the training output action of the generative adversarial network contains the style of the input style action and the content of the input content action, the style performance of the stylized action has a fine-grained style performance guided by the input semantic information
[0058] Referring to Figure 3 as shown, a human motion style automatic transfer system executes a human action style transfer method according to any one of the above, and includes a data acquisition module, a data disassembling module, a data integration module, a data decoding module, and an adversarial generation module;
[0059] Data acquisition module: acquire the input content action and the style action;
[0060] Data disassembling module: vectorize the semantic information of the input content action and input it into a fully connected neural network to obtain the semantic potential code, input the input style action into a convolutional neural network to obtain the style potential code, and input the input content action into the encoding network of the generation module of the generative adversarial network to obtain the intermediate content potential code;
[0061] Data integration module: fuse the semantic potential code and the style potential code together to obtain a semantic perception style potential code containing specific semantic content information, input the semantic perception style potential code and the content potential code into the adaptive instance normalization AdaIN layer to change the data distribution of the content potential code, and obtain the stylized content potential code;
[0062] Data decoding module: Input the stylized content latent code into the decoding network of the generation module of the generative adversarial network to decode the action;
[0063] Adversarial generation module: The adversarial module of the generative adversarial network discriminates the style of the generated action to perform adversarial training on the network, and obtains a stylized action with fine-grained style guided by semantics.
[0064] Specifically, the present invention inputs the semantic information of the content action, where the semantic information refers to the category information of the action. After vectorizing it, it is input into a multi-layer perceptron network to construct a semantic latent code; the input style action is input into an action style encoder composed of a convolutional neural network to obtain a single style latent code; the obtained semantic latent code and the single style latent code are fused through a dot product and residual structure, so that the single style latent code can fuse specific action semantic features to express more fine-grained style features. This feature makes the action style transfer result obtained by our method have stronger style expressiveness; the obtained semantic-aware style latent code is input into an adaptive instance normalization (AdaIN) layer, thereby changing the data distribution of the content latent code obtained by encoding the input content action into the convolutional neural network. This data distribution includes the mean and variance of the data; the obtained stylized content latent code is input into the decoding network to generate a stylized action with fine-grained style guided by semantics. Through the above scheme, the present invention can finally obtain a stylized action with fine-grained style guided by semantics after style transfer by inputting the style action, content action, and semantic information of the content action. At the same time, due to the addition of the action content semantics, the obtained stylized action can greatly retain the content information of the action. Extensive experimental results and comprehensive evaluations confirm that the action style transfer method of the present application can well transfer the style of the action while retaining the content of the action.
[0065] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of any one of the above human action style transfer methods.
[0066] In one embodiment, this method is compared with other high-performance human shape estimation methods in a quantitative manner. The performance of the state-of-the-art methods: Deep-Motion-Editing, Diverse-Style-Stylization, and Motion-Puzzle. The comparison data and results are shown in Table 1.
[0067] Table 1 evaluates the action quality, content preservation degree, and style performance degree respectively
[0068] Method FMD CRA SRA Deep-Motion-Editing 103.25 57.38 37.38 Diverse-Style-Stylization 30.52 66.07 25.98 Motion-Puzzle 49.18 35.22 29.14 Ours 7.04 84.64 55.63
[0069] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0070] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for human motion style transfer, characterized in that, It includes the following steps: Data acquisition step: Acquire the input content action and the style action; Data disassembling step: Vectorize the semantic information of the input content action and input it into a fully-connected neural network to obtain a semantic latent code, input the input style action into a convolutional neural network to obtain a style latent code, and input the input content action into the encoding network of the generation module of the generative adversarial network to obtain an intermediate content latent code; Data integration step: Fuse the semantic latent code and the style latent code together to obtain a semantic-aware style latent code containing specific semantic content information, input the semantic-aware style latent code and the content latent code into an adaptive instance normalization (AdaIN) layer to change the data distribution of the content latent code and obtain a stylized content latent code; Data decoding step: Input the stylized content latent code into the decoding network of the generation module of the generative adversarial network to decode the action; Adversarial generation step: The adversarial module of the generative adversarial network discriminates the style of the generated action to perform adversarial training on the network, and obtains a stylized action with fine-grained style guided by semantics.
2. A method for human action style transfer according to claim 1, wherein In the data disassembling step, the semantic information of the input content action is vectorized through embedding and then input into a fully-connected neural network to obtain a semantic latent code containing content features.
3. A method for human action style transfer according to claim 1, wherein In the data disassembling step, the semantic information refers to the action content label in the input content action.
4. A method for human action style transfer according to claim 1, wherein In the data disassembling step, the generation module of the generative adversarial network is composed of an encoder and a decoder, and both the encoder and the decoder are composed of convolutional neural networks.
5. A method for human action style transfer according to claim 1, wherein In the data integration step, the formula for fusing the semantic latent code and the style latent code is as follows: Among them, sig() represents the sigmoid activation function; represents the semantic latent code; z s represents the style latent code; z cs represents the obtained semantic-aware style latent code.
6. A method for human action style transfer according to claim 1, wherein In the data integration step, the data distribution is the mean and variance of the data, and the method implemented by the AdaIN layer is as follows: Among them, z c represents the potential code of the content; z cs represents the potential code of the semantic perception style; μ and σ represent the mean and variance of the data respectively.
7. A method for human action style transfer according to claim 1, wherein In the adversarial generation step, the adversarial module of the generative adversarial network is composed of a discriminator, which is used to discriminate the style category of the generated action.
8. A method for human action style transfer according to claim 1, wherein In the adversarial generation step, the generated output action of the generative adversarial network training contains the style of the input style action and preserves the content of the input content action, and has a fine-grained style guided by the input semantic information.
9. An automatic human motion style transfer system, characterized in that, Performing a method for human action style transfer according to any one of claims 1-8, including a data acquisition module, a data disassembling module, a data integration module, a data decoding module, and an adversarial generation module connected in sequence; Data acquisition module: Acquire the input content action and the style action; Data disassembling module: Vectorize the semantic information of the input content action and input it into a fully connected neural network to obtain a semantic latent code. Input the input style action into a convolutional neural network to obtain a style latent code. Input the input content action into the encoding network of the generation module of the generative adversarial network to obtain an intermediate content latent code; Data integration module: Fuse the semantic latent code and the style latent code together to obtain a semantic-aware style latent code containing specific semantic content information. Input the semantic-aware style latent code and the content latent code into an adaptive instance normalization (AdaIN) layer to change the data distribution of the content latent code and obtain a stylized content latent code; Data decoding module: Input the stylized content latent code into the decoding network of the generation module of the generative adversarial network to decode the action; Adversarial generation module: The adversarial module of the generative adversarial network discriminates the style of the generated action to perform adversarial training on the network, and obtains a stylized action with fine-grained style guided by semantics.
10. A computer-readable storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a processor, the steps of the human action style transfer method described in any one of claims 1-8 are implemented.
Citation Information
Patent Citations
Video-animation style migration method based on deep adversarial network
CN112164130A
Style migration method for automatically generating stylized video
CN112884636A