An Unsupervised Action Style Transfer Method Based on Reversible Flow Networks
The reversible flow network-based unsupervised action style transfer method addresses data requirement challenges and maintains content integrity in action style transfer, ensuring high-quality results without paired data.
Patent Information
- Application Number
- CN202111415623.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The existing deep learning-based action style transfer method has high requirements for training data, and the generated actions are prone to losing some frames, and the content integrity and consistency cannot be guaranteed.
Unsupervised action style transfer method based on reversible flow network is adopted, action style transfer is carried out through projection flow network and adaptive instance normalization layer, and training is performed using loss functions, including content loss, style loss, action reconstruction loss and joint torsion loss, reducing training data requirements and maintaining content integrity.
Without the need for paired training data and labels, stylized actions that maintain content integrity and consistency can be learned and generated to maintain the style transfer effect.
Smart Images

Figure CN113947525B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and more specifically, to an unsupervised action style transfer method based on a reversible flow network. Background Art
[0002] The style characteristics of actions often convey information about a person's mood, personality, or identity. Games and animated movies often need to consider various different action styles in order to pursue realistic and expressive characters. However, traditional motion capture technology is costly, and it is not practical to rely on equipment to capture all styles of motion. Style transfer based on existing actions is a more economically feasible approach.
[0003] Action style transfer refers to extracting style information from an action segment S and applying it to the content of an action segment C. Here, style refers to the motion attributes that convey the mood and personality of a character, usually reflected in the amplitude of the action and the degree of limb bending. Content is the type of an action, the pace, and the motion trajectory.
[0004] Early style transfer work used manually extracted features, such as extracting the spectral intensity representation of any action style and using the spectral differences of different styles for style transfer. However, since style is an abstract attribute and cannot be described by an exact mathematical definition. Therefore, in contrast, data-driven methods infer style features from a large number of samples, which have more advantages than manually crafted features to characterize style and content. However, most deep learning-based training methods are supervised, and these methods rely on paired samples for training, that is, actions with exactly the same content but only different styles. The acquisition of such training data has certain difficulties. In addition, these methods usually require that the data must have style annotation information.
[0005] Recently, some training methods that do not require paired action data and do not require data annotation information have emerged. However, this method still has some problems. The generated actions will lose some frames and are prone to changing the original pace frequency, that is, it cannot well guarantee the content integrity and consistency. Summary of the Invention
[0006] In order to solve the problems of high requirements for training data and inability to guarantee the content integrity and consistency of generated actions existing in the existing deep learning-based action style transfer methods, the present invention provides an unsupervised action style transfer method based on a reversible flow network, which can reduce the requirements for training data, retain the integrity of the content, and not reduce the effect of style transfer.
[0007] To achieve the above object of the present invention, the following technical solutions are adopted:
[0008] An unsupervised action style transfer method based on a reversible flow network, the method comprising the following steps:
[0009] S1: Obtain action data in BVH format and perform preprocessing, and use the preprocessed action data as training samples; the action data includes content actions and style actions;
[0010] S2: Construct an action style transfer model, the action style transfer model including a projection flow network for encoding action hidden features and an adaptive instance normalization layer for hidden feature style transfer;
[0011] S3: Train the action style transfer model using the training samples, and adjust the parameters of the action style transfer model using a loss function in each round of training;
[0012] S4: Input the action data to be transferred into the trained action style transfer model to achieve style transfer between any two actions.
[0013] Preferably, in step S1, the preprocessing of the action data is as follows:
[0014] S101: Convert the action data in BVH format into quaternion form x and coordinate form x pos , and perform segment cutting with a length of 32 frames;
[0015] S102: Normalize the data in quaternion form x and coordinate form x pos respectively.
[0016] Furthermore, in step S2, the projection flow network is composed of 16 stacked flow structures, and each flow includes three parts: an activation layer, a reversible 1×1 convolutional layer, and an affine coupling layer;
[0017] The activation layer makes each channel have zero mean and unit variance;
[0018] The reversible 1×1 convolutional layer is used to permute the channel dimension of the feature map;
[0019] The affine coupling layer uses affine coupling to transform the data of each channel.
[0020] Still further, the affine coupling layer uses affine coupling to transform the data, specifically as follows:
[0021] First, use the split() function to divide the data of each channel into two halves x a , x b along the channel dimension; then use the affine parameters ψ and δ to perform an affine transformation on one of the halves x b to obtain z b; Finally, use the concat() function to concatenate x along the channel dimension a , z b Two pieces of data;
[0022] Among them, the affine parameters ψ, δ are calculated using a Transformer network based on half of the data for each channel x a and the coordinate form x of the action data pos for calculation.
[0023] Preferably, in step S2, the adaptive instance normalization layer is used for style transfer between hidden features; the specific calculation of the adaptive instance normalization is as follows:
[0024]
[0025] Among them, z c represents the hidden feature of the content action, and z s represents the hidden feature of the style action. μ represents the mean, and σ represents the standard deviation.
[0026] Preferably, the action style transfer model processes the input action data as follows:
[0027] D1: Input the content action and the style action into the projection flow network, and respectively obtain the first hidden features corresponding to the content action and the style action through the forward propagation process;
[0028] D2: Input the first hidden features corresponding to the content action and the style action into the adaptive instance normalization layer respectively, map the first hidden feature of the style action to two affine parameters ψ, δ, and perform adaptive instance normalization on the first hidden feature of the content action with this, and finally output the second hidden feature of the stylized action;
[0029] D3: Input the second hidden feature of the stylized action into the projection flow network and perform backpropagation to reconstruct the finally style - transferred action.
[0030] Furthermore, in step S3, the loss function includes a content loss function, a style loss function, an action reconstruction loss function, and a joint torsion loss function;
[0031] The content loss function:
[0032] L c =‖G(G -1 (t)) - t‖2
[0033] In the formula, G represents the forward propagation process of the action style transfer model; G -1 represents the backward propagation process of the action style transfer model; t represents the output of the adaptive instance normalization;
[0034] The described style loss function:
[0035] L s = ‖μ(G(G -1 (t))) - μ(G(s))‖² + ‖σ(G(G -1 (t))) - σ(G(s))‖²
[0036] Wherein, μ represents the average value, σ represents the standard deviation, and s represents the input style action;
[0037] The described action reconstruction loss:
[0038] L r = ‖G -1 (G(c|c)) - c‖₁
[0039] Wherein, c represents the content action;
[0040] The described joint torsion loss:
[0041]
[0042] Wherein, q is the quaternion - form action data obtained by denormalizing the reverse output x act of the action style transfer model; euler y represents the value of the y - axis after converting the quaternion to the Euler angle; α is the maximum angle of joint torsion.
[0043] Furthermore, before step D1, first perform quaternion spherical interpolation on the content action and the style action to increase the time dimension. Specifically:
[0044] Given two quaternions q1 and q2, the interpolation formula between them is as follows:
[0045]
[0046] Wherein, θ represents the angle between q1 and q2; t represents the interpolation position.
[0047] Furthermore, in step D3, input the second hidden feature of the stylized action into the projection flow network and perform backpropagation, specifically as follows:
[0048] First, perform inverse affine coupling layer operation, inverse invertible 1×1 convolution layer operation, and inverse activation layer operation in sequence; then, for the x act obtained by the inverse activation layer operation, perform downsampling using the interpolation frequency:
[0049] x sample (i) = x act (3i)
[0050] Among them, x(i) represents the i-th frame of action x;
[0051] Finally, affine parameters ψ and δ are used for denormalization:
[0052] q = x act ⊙δ(x) + ψ(x)
[0053] The finally obtained q is the action after style transfer.
[0054] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0055] The beneficial effects of the present invention are as follows:
[0056] Compared with other action style transfer methods, the present invention can learn and generate stylized actions that maintain content integrity and consistency under the condition of reducing the requirements for training data, that is, without using paired training data and without any annotation of training data, and has the integrity of retaining content and does not reduce the effect of style transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a flowchart of the unsupervised action style transfer method described in Embodiment 1.
[0058] Figure 2 is a framework diagram of the action style transfer in Embodiment 1.
[0059] Figure 3 is a schematic diagram of the principle of the Transformer to generate affine parameters of the affine coupling layer. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0061] Embodiment 1
[0062] As Figure 1 shown, an unsupervised action style transfer method based on a reversible flow network, the method includes the following steps:
[0063] S1: Obtain action data in BVH format and perform preprocessing, and use the preprocessed action data as training samples; the action data includes content actions and style actions.
[0064] In a specific embodiment, in step S1, the action data is preprocessed as follows:
[0065] S101: Convert the action data in the original Euler angle representation form into quaternion form \(x\) and coordinate form \(x\). pos , and then perform segment cutting. Specifically, take a sampling window of 32 frames in length to sample the entire action segment, moving 8 frames each time, that is, there are overlapping segments of 8 frames.
[0066] S102: After cutting all action segments, normalize the data of quaternion form \(x\) and coordinate form \(x\) respectively. pos
[0067]
[0068] In the formula, \(\mu\) represents the mean value, and \(\sigma\) represents the standard deviation.
[0069] S2: Construct an action style transfer model. The action style transfer model includes a projection flow network for encoding action hidden features and an adaptive instance normalization layer for hidden feature style transfer.
[0070] In a specific embodiment, in step S2, in order to extract deep first hidden features, the projection flow network is composed of 16 stacked flow structures, and each flow includes three parts: an activation layer, a reversible 1×1 convolutional layer, and an affine coupling layer.
[0071] The activation layer makes each channel have zero mean and unit variance, which is convenient for subsequent calculations; the activation layer uses a scaling parameter \(w\) and an offset parameter \(b\) to calculate the action data as follows:
[0072] x act = w⊙x i + b (2)
[0073] In the formula, \(x\) act represents the data output by the activation layer; \(x\) i represents the action data of each frame, and \(w\) and \(b\) are initialized with the mean and variance of the initial batch of training data, and then learned and modified during the training process to achieve optimization.
[0074] The reversible 1×1 convolutional layer is used to permute the channel dimension of the feature map. The purpose is to make each dimension affect all other dimensions during the subsequent affine coupling layer transformation. The specific calculation of the reversible 1×1 convolutional layer is as follows:
[0075] x conv = Wx act (3)
[0076] In the formula, \(W\) represents a randomly initialized weight matrix, the size of which is \(c×c\), and \(c\) is the size of the feature dimension. The affine coupling layer uses affine coupling to transform the data of each channel, and the specific calculation is as follows:
[0077] x a , x b = Split(x conv ) (4)
[0078] z b = (x b + ψ) ⊙ δ (5)
[0079] z = Concat(x a , z b ) (6)
[0080] First, use the split() function to divide the data of each channel into two halves x a , x b ; then, use the affine parameters ψ and δ to perform an affine transformation on one of the halves x b to obtain z b ; finally, use the concat() function to concatenate x a , z b for the two data along the channel dimension.
[0081] Among them, the affine parameters ψ, δ are calculated using the Transformer network according to half of the data of each channel x a and the coordinate form x of the action data pos as shown in Figure 3 . Since the Transformer network constructs long-range dependencies based on the attention mechanism, it avoids the problem of error accumulation in sequence tasks and can better utilize the global sequence information. In addition, the purpose of adding the coordinate form x pos of the action data is to enhance the input information, thereby further strengthening the differences between different actions and helping the training of style features.
[0082] Specifically, connect x a and x pos along the feature channel to obtain the input data x with a feature dimension of d model . First, calculate the position embedding vector PE as follows:
[0083]
[0084]
[0085] In the formula, pos represents the position subscript of each hidden feature, and sin and cos are alternately encoded for even and odd positions. The purpose is to give the encoding a certain value range but be independent of the action length.
[0086] After that, add the position embedding vector to the original input data x to obtain xpe = x + PE
[0087] Next, calculate x pe for the multi - head self - attention of Q , use three different random matrices W K , W V pe k , and multiply them by x respectively to obtain three vectors Q, K, and V. According to these three vectors, the method of calculating self - attention using scaled dot - product attention is as follows:
[0088]
[0089] where d k is the dimension of these three vectors.
[0090] According to the above expression, the multi - head self - attention can be further calculated. Specifically, divide x pe into h heads, and for each head, calculate its self - attention as follows:
[0091] head i = Attention(QW i Q , KW i K , VW i V ) (10)
[0092] Here, W is used to perform a linear transformation on the three vectors Q, K, and V.
[0093] After that, concatenate the self - attention results of h times to obtain the multi - head self - attention result as follows:
[0094] MultiHead(Q, K, V) = Concat(head1, ……, head h )W O (11)
[0095] where W O is the parameter used to perform a linear transformation on the concatenated result. After that, the encoding result of the Transformer can be obtained:
[0096] x encoder = FFN(LayerNorm(x pe + MultiHead)) (12)
[0097] In the formula, LayerNorm represents normalizing the data; FFN is a feed - forward neural network layer composed of a linear transformation and an activation layer.
[0098] Then, a multi - layer perceptron (MLP) is used as the decoder to map the dimension of x encoder to twice the dimension of x b , and then it is split into affine parameters, and the expression is as follows:
[0099] ψ,δ=SplitCross(MLP(x encoder )) (13)
[0100] In the formula, SplitCross means splitting the result of the linear layer into two affine parameters. Specifically, the odd - numbered bits are taken as the scaling parameter ψ, and the even - numbered bits are taken as the offset parameter δ.
[0101] In another specific embodiment, in step S2, the adaptive instance normalization layer is used for style transfer between hidden features; the specific calculation of the adaptive instance normalization is as follows:
[0102]
[0103] where z c represents the first hidden feature of the content action, z s represents the first hidden feature of the style action, μ represents the mean, and σ represents the standard deviation.
[0104] S3: Design a loss function to perform unsupervised training on the action style transfer. Specifically, use training samples to train the action style transfer model, and in each round of training, use the loss function to adjust the parameters of the action style transfer model.
[0105] In a specific embodiment, each time during training, the action data obtained in S1 is input, which includes the content action c and the style action s. The action style transfer model processes the input action data as follows:
[0106] D1: Input the content action and the style action into the projection flow network, and respectively obtain the first hidden features corresponding to the content action and the style action through the forward propagation process;
[0107] D2: Respectively input the first hidden features corresponding to the content action and the style action into the adaptive instance normalization layer, map the first hidden feature of the style action into two affine parameters ψ,δ, and perform adaptive instance normalization on the first hidden feature of the content action with this, and finally output the second hidden feature of the stylized action;
[0108] D3: Input the second hidden feature of the stylized action into the projection flow network and perform backpropagation to reconstruct the finally style - transferred action.
[0109] In each round of training, the loss function is used to adjust the network parameters, specifically including content loss, style loss, action reconstruction loss, and joint torsion loss. Assume that the forward and backward propagation processes of the network model are G and G -1 , and the specific calculation methods of each loss are as follows:
[0110] In a specific embodiment, in step S3, the loss function includes a content loss function, a style loss function, an action reconstruction loss function, and a joint torsion loss function. Assume that the forward propagation process of the action style transfer model is G, and the backward propagation process of the action style transfer model is G -1 . The specific calculation methods of each loss are as follows:
[0111] The content loss function:
[0112] L c =‖G(G -1 (t)) - t‖2 (15)
[0113] In the formula, t represents the output of adaptive instance normalization; the content loss function is mainly used to ensure that the forward and backward propagation processes of the projection flow network do not change the action features.
[0114] The style loss function:
[0115] L s =‖μ(G(G -1 (t))) - μ(G(s))‖2 + ‖σ(G(G -1 (t))) - σ(G(s))‖2 (16)
[0116] In the formula, μ represents the mean value, σ represents the standard deviation, and s represents the input style action; the style loss function is mainly used to ensure that the mean value and standard deviation of the second hidden feature of the stylized action are consistent with the affine parameters generated by the style input action.
[0117] The action reconstruction loss:
[0118] L r =‖G -1 (G(c|c)) - c‖1 (17)
[0119] In the formula, c represents the content action; the purpose of the action reconstruction loss is that when the style action and the content action are the same action, the generated result should be as close as possible to the input.
[0120] The joint torsion loss:
[0121]
[0122] Where q is the reverse output x of the action style transfer model act The action data in quaternion form obtained after denormalization; euler y represents the value of the y-axis after converting the quaternion to Euler angles; α is the maximum angle of joint twist, which is set to 100° here.
[0123] S4: Using the trained action style transfer model in S3, the style transfer between any two actions can be performed. Input the action data to be transferred into the trained action style transfer model to achieve the style transfer between any two actions.
[0124] Embodiment 2
[0125] Based on the unsupervised action style transfer method based on the reversible flow network described in Embodiment 1, in the unsupervised action style transfer method described in this embodiment, before step D1, a content action c and a style action s are given.
[0126] First, perform quaternion spherical interpolation on the content action and the style action to increase the time dimension, thereby enhancing the data. Specifically:
[0127] Given two quaternions q1 and q2, the interpolation formula between them is as follows:[[]]
[0128]
[0129] Where θ represents the angle between q1 and q2; t represents the interpolation position.
[0130] For an action, insert two frames between every two frames. For example, for a 32-frame action, it will be expanded to 94 frames after interpolation. Then input the two actions into the projection flow network in the first step respectively, and use the forward propagation process of the projection flow network to obtain the first hidden feature z of the content action c c and the first hidden feature z of the style action S s .
[0131] Input the first hidden feature z of the obtained content action c c and the first hidden feature z of the style action s s into the adaptive instance normalization layer to obtain the second hidden feature of the stylized action:
[0132] z stylized = AdaIN(z c , z s ) (20)
[0133] Input the second hidden feature z of the stylized action stylizedInput into the projection flow network and perform backpropagation to reconstruct the final action after style transfer.
[0134] Specifically, first is the inverse operation of the affine coupling layer:
[0135] z a ,z b = Split(z stylized ) (21)
[0136] ψ,δ = Transformer(z a ,x pos ) (22)
[0137] x b = z b / δ - ψ (23)
[0138] x coupling = Concat(z a ,x b ) (24)
[0139] Then are the inverse operations of the invertible 1×1 convolutional layer and the activation layer:
[0140] x′ conv = W -1 x coupling (25)
[0141] x′ act = (x′ conv - b) / w (26)
[0142] where W -1 represents the inverse operation of the weight matrix W.
[0143] After that, downsampling is performed on x′ act obtained from the inverse operation of the activation layer using the interpolation frequency (interpolation at an interval of two frames):
[0144] x sample (i) = x′ act (3i) (27)
[0145] where x(i) represents the i-th frame of the action x;
[0146] Finally, denormalization is performed using the affine parameters ψ,δ:
[0147] q = x act ⊙σ(x)+ψ(x)
[0148] The finally obtained q is the action after style transfer.
[0149] Example 3
[0150] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following method steps are implemented:
[0151] S1: Obtain action data in BVH format and perform preprocessing, and use the preprocessed action data as training samples; the action data includes content actions and style actions;
[0152] S2: Construct an action style transfer model, where the action style transfer model includes a projection flow network for encoding action hidden features and an adaptive instance normalization layer for hidden feature style transfer;
[0153] S3: Train the action style transfer model using the training samples, and adjust the parameters of the action style transfer model using a loss function in each round of training;
[0154] S4: Input the action data to be transferred into the trained action style transfer model to achieve style transfer between any two actions.
[0155] Embodiment 4
[0156] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the following method steps are implemented:
[0157] S1: Obtain action data in BVH format and perform preprocessing, and use the preprocessed action data as training samples; the action data includes content actions and style actions;
[0158] S2: Construct an action style transfer model, where the action style transfer model includes a projection flow network for encoding action hidden features and an adaptive instance normalization layer for hidden feature style transfer;
[0159] S3: Train the action style transfer model using the training samples, and adjust the parameters of the action style transfer model using a loss function in each round of training;
[0160] S4: Input the action data to be transferred into the trained action style transfer model to achieve style transfer between any two actions.
[0161] The various embodiments of the present invention can be combined arbitrarily to achieve different technical effects.
[0162] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk).
[0163] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by hardware instructed by a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disk that can store program codes.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An unsupervised action style transfer method based on a reversible flow network, characterized in that: The method described above includes the following steps: S1: Obtain motion data in BVH format and perform preprocessing, and use the preprocessed motion data as training samples; the motion data includes content motions and style motions; S2: Construct a motion style transfer model, and the motion style transfer model includes a projection flow network for encoding motion hidden features and an adaptive instance normalization layer for style transfer of hidden features; The motion style transfer model processes the input motion data as follows: D1: Input the content motion and the style motion into the projection flow network, and respectively obtain the first hidden features corresponding to the content motion and the style motion through the forward propagation process; D2: Input the first hidden features corresponding to the content action and the style action into the adaptive instance normalization layer respectively, map the first hidden feature of the style action to two affine parameters , and perform adaptive instance normalization on the first hidden feature of the content action with these parameters, and finally output the second hidden feature of the stylized action; D3: Input the second hidden feature of the stylized motion into the projection flow network and perform backpropagation to reconstruct the finally style-transferred motion; S3: Use the training samples to train the motion style transfer model, and adjust the parameters of the motion style transfer model using the loss function in each round of training; S4: Input the motion data to be transferred into the trained motion style transfer model to achieve style transfer between any two motions.
2. The unsupervised action style transfer method based on a reversible flow network according to claim 1, characterized in that: In step S1, preprocess the motion data as follows: S101: Convert the BVH format motion data into quaternion form and coordinate form , and perform segment cutting with a length of 32 frames; S102: Normalize the data in quaternion form and coordinate form respectively.
3. The unsupervised action style transfer method based on a reversible flow network according to claim 1, wherein: In step S2, the projection flow network is composed of 16 stacked flow structures, and each flow includes three parts: an activation layer, a reversible 1×1 convolutional layer, and an affine coupling layer; The activation layer makes each channel have zero mean and unit variance; The reversible 1×1 convolutional layer is used to permute the channel dimension of the feature map; The affine coupling layer uses affine coupling to transform the data of each channel.
4. The unsupervised action style transfer method based on a reversible flow network according to claim 3, characterized in that: The affine coupling layer uses affine coupling to transform the data as follows: First, use the split( ) function to divide each channel data into two halves along the channel dimension ; Then use the affine parameters and to perform an affine transformation on one of the halves to obtain ; Finally, use the concat( ) function to concatenate the two data along the channel dimension; Among them, the affine parameters are calculated using a Transformer network based on half of the data for each channel and the coordinate form of the action data to obtain the result 5. The unsupervised action style transfer method based on a reversible flow network according to claim 1, wherein: In step S2, the adaptive instance normalization layer is used for style transfer between hidden features; the specific calculation of the adaptive instance normalization is as follows: Among them, represents the hidden feature of the content action, represents the hidden feature of the style action, represents the mean value, represents the standard deviation.
6. The unsupervised action style transfer method based on a reversible flow network according to claim 1, characterized in that: In step S3, the loss function includes a content loss function, a style loss function, a motion reconstruction loss function, and a joint torsion loss function; The content loss function: In the formula, represents the forward propagation process of the action style transfer model; represents the backward propagation process of the action style transfer model; represents the output of the adaptive instance normalization; The style loss function: where 𝜇 represents the mean value and 𝜎 represents the standard deviation, represents the input style action; The motion reconstruction loss function: In the formula, represents the content action; The joint torsion loss function: In the formula, is the reverse output of the action style transfer model The action data in quaternion form obtained after denormalization; represents the value of the y-axis after converting the quaternion to Euler angles; is the maximum angle of joint twist.
7. The unsupervised action style transfer method based on a reversible flow network according to claim 6, characterized in that: Before step D1, first perform quaternion spherical interpolation on the content motion and the style motion to increase the time dimension. Specifically: Given two quaternions , , the interpolation formula between them is as follows: In the formula, ; represents the interpolation position.
8. The unsupervised action style transfer method based on a reversible flow network according to claim 7, characterized in that: In step D3, input the second hidden feature of the stylized motion into the projection flow network and perform backpropagation as follows: First, perform the inverse operation of the affine coupling layer, the inverse operation of the reversible 1×1 convolutional layer, and the inverse operation of the activation layer in sequence; After performing inverse operations on the activation layer, the Perform downsampling using the interpolation frame rate: Among them, represents an action of the i nth frame; Finally, affine parameters are used again for denormalization: The finally obtained q is the action after style transfer.
9. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Clothing image artistry generation method based on deep learning style migration
CN110490791A
Generative stream model-based human motion style migration method and system
CN113012036A