Sign language action transfer method and device based on motion redirection
By decoupling skeletal movements and structural information through a recurrent generative adversarial network, the problems of low finger redirection accuracy and training data in sign language synthesis are solved, and high-precision sign language movement migration is achieved. It is applicable to a variety of skeletal structures, and the movements are natural, coherent, and accurately positioned.
Patent Information
- Application Number
- CN202310803770.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-06-30
AI Technical Summary
Existing motion redirection technology in sign language synthesis has problems such as low finger redirection accuracy, severe deformation of sign language movements, and difficulty in obtaining paired training data. In particular, when the hierarchical structures of the source and target skeletons are inconsistent or there is a large difference in body proportions, the redirection error is large.
A cyclic generative adversarial network is used for unsupervised training to construct encoder, decoder and discriminator models. Skeletal motion and structural information are decoupled through motion encoder, static encoder and latent encoder. Combined with attention mechanism and multiple loss functions, high-precision sign language animation data is generated.
The accuracy and applicability of sign language movement transfer are improved, the finger redirection accuracy is high, the movement is natural and coherent, it is applicable to different skeleton structures, does not require a consistent hierarchical structure, and reduces the training data requirements and reasoning complexity.
Smart Images

Figure CN116844231B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of skeleton animation motion redirection, and in particular to a sign language action migration method and device based on motion redirection. Background Art
[0002] Sign language is a visual language that uses body movements, finger gestures, facial expressions, and lip movements to express meaning and communicate. It is the primary means for hearing-impaired people to communicate with hearing people in daily life. Sign language synthesis refers to the technology of translating natural language (such as Chinese) into sign language movements. Sign language synthesis includes video real-person sign language synthesis and virtual digital human sign language synthesis. Among them, digital human sign language synthesis has attracted much attention due to its high efficiency and diverse presentation forms. Digital human sign language synthesis mainly displays sign language movements through skeletal animation. Since the use of motion capture technology to construct skeletal animation data for sign language movements is labor-intensive and costly, in order to use the same skeletal animation data on different digital humans, a motion redirection technology is urgently needed to migrate a set of standard digital human skeletal animation data to different digital human skeletons.
[0003] Traditional motion redirection is typically implemented using an inverse kinematics (IK) algorithm. IK is first applied to each frame to satisfy constraints, and then the resulting motion is smoothed by assembling multiple layers of B-spline curves. To respond to changing effector positions while preserving the details of the original motion, the IK algorithm must also account for changes in joint angles. The IK algorithm requires a significant amount of time to construct the constraint matrix, which is then calculated through iterative reasoning to achieve the redirection result. This traditional algorithm requires the source and target skeletons to have consistent hierarchical structures, and the redirected positional error can be significant when the body proportions of the source and target skeletons differ significantly.
[0004] Existing motion redirection widely uses deep learning technology. This technology removes hand joints during redirection, redirecting only the relevant movements of the torso, arms, and legs, making it more effective for limb movements. However, sign language movements are crucial to finger gestures. Because finger redirection and the influence of other bones on the current skeleton during the redirection process are not considered, direct application leads to severe finger deformation and low finger position redirection accuracy. Furthermore, existing deep learning methods require paired data for training, which is generally difficult to obtain in motion redirection. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention proposes a sign language action migration method and device based on motion redirection, which uses a cyclic generative adversarial network for unsupervised training, solving the problem of difficulty in obtaining paired training data.
[0006] In order to achieve the above-mentioned object, the present invention provides, on one hand, a method for transferring sign language actions based on motion redirection, comprising:
[0007] Constructing an encoder model, wherein the encoder model is configured as a motion encoder, a static encoder, and a latent encoder;
[0008] The motion encoder is configured to: input the original skeleton sign language animation data and output the encoded skeleton motion information;
[0009] The static encoder is configured as follows: input is skeleton space static data, and output is encoded skeleton structure information;
[0010] The latent encoder is coupled to the motion encoder and the static encoder, and is configured to: decouple the skeleton action information from the skeleton structure information to extract the abstract sign language action;
[0011] Constructing a decoder model, the decoder model being coupled to the encoder model and configured to: redirect the sign language abstract action and skeleton structure information to generate skeleton-reconstructed sign language animation data;
[0012] A discriminator model is constructed, and the discriminator model is coupled to the encoder model and the decoder model, and is configured to: input the skeleton original sign language animation data, the skeleton reconstructed sign language animation data, and the skeleton space static data, and output identification results including a first identification result regarding the skeleton original sign language animation data and the skeleton space static data, and a second identification result regarding the skeleton reconstructed sign language animation data and the skeleton space static data.
[0013] Optionally, the method further includes constructing a target loss function, wherein the target loss function includes a shallow loss function, and the shallow loss function is used to constrain skeleton motion information of both the source skeleton and the target skeleton;
[0014] The shallow loss function is determined according to the skeleton action information generated by inputting the original sign language animation data of the source skeleton into the motion encoder and the skeleton action information generated by inputting the original sign language animation data of the redirected target skeleton into the motion encoder.
[0015] Optionally, the shallow loss function is expressed as:
[0016]
[0017] in, The original sign language animation data Q representing the source skeleton A A Input the skeleton motion information generated by the motion encoder, The original sign language animation data Q representing the target skeleton B B Input the skeleton motion information generated by the motion encoder, l ltc Represents a shallow loss function.
[0018] Optionally, the target loss function further includes a reconstruction loss function, and the reconstruction loss function is used to constrain the reconstruction information of both the source skeleton and the target skeleton;
[0019] The reconstruction loss function is determined according to the original sign language animation data and the reconstructed sign language animation data of the source skeleton, and the original sign language animation data and the reconstructed sign language animation data of the target skeleton.
[0020] Optionally, the reconstruction loss function is expressed as:
[0021]
[0022] Among them, the Q A 、 Respectively represent the original sign language animation data and the reconstructed sign language animation data of the source skeleton A,
[0023] Q B 、 They represent the original sign language animation data and the reconstructed sign language animation data of the target skeleton B, respectively. rec represents the reconstruction loss function.
[0024] Optionally, the target loss function further includes an adversarial loss function, and the adversarial loss function is used to constrain adversarial information between the source skeleton and the target skeleton;
[0025] The adversarial loss function is determined based on the first identification result and the second identification result of the source skeleton and the first identification result and the second identification result of the target skeleton.
[0026] Optionally, the adversarial loss function is expressed as:
[0027]
[0028] in, Represents the skeleton reconstruction sign language animation data of the source skeleton A With the skeleton space static data S A The second identification result, C A (Q A ,S A ) represents the original sign language animation data Q of the source skeleton A A With the skeleton space static data S A The first identification result; Represents the target skeleton B's skeleton reconstruction sign language animation data With the skeleton space static data S B The second identification result, C B (Q B ,S B) represents the original sign language animation data Q of the target skeleton B B With the skeleton space static data S B The first identification result, l adv represents the adversarial loss function.
[0029] Optionally, the target loss function further includes an end loss function, and the end loss function is used to constrain the movement speed of each skeletal joint at the end of the skeleton;
[0030] The terminal loss function is determined according to the movement speed of each skeletal joint at the end of the source skeleton and the movement speed of each skeletal joint at the end of the target skeleton.
[0031] Optionally, the terminal loss function is expressed as:
[0032]
[0033] in, Indicates the speed of each bone joint at the end of source skeleton A, Indicates the velocity of each bone joint at the end of the target skeleton B; h A Indicates the height of the source skeleton A, h B Indicates the skeleton height of the target skeleton B, l ee represents the terminal loss function.
[0034] Another aspect of the present invention provides a sign language motion transfer device based on motion redirection, which adopts the above-mentioned sign language motion transfer method based on motion redirection, and at least includes:
[0035] An encoder module, configured to construct an encoder model, wherein the encoder model is configured as a motion encoder, a static encoder, and a latent encoder;
[0036] The motion encoder is configured to: input the original skeleton sign language animation data and output the encoded skeleton motion information;
[0037] The static encoder is configured as follows: input is skeleton space static data, and output is encoded skeleton structure information;
[0038] The latent encoder is coupled to the motion encoder and the static encoder, and is configured to: decouple the skeleton action information from the skeleton structure information to extract the abstract sign language action;
[0039] A decoder module is used to construct a decoder model, the decoder model is coupled to the encoder model, and is configured to: redirect the sign language abstract action and skeleton structure information to generate skeleton-reconstructed sign language animation data;
[0040] The discriminator module is used to construct a discriminator model, which is coupled to the encoder model and the decoder model and configured to: input the skeleton original sign language animation data, the skeleton reconstructed sign language animation data, and the skeleton space static data, and output identification results including a first identification result regarding the skeleton original sign language animation data and the skeleton space static data, and a second identification result regarding the skeleton reconstructed sign language animation data and the skeleton space static data.
[0041] From the above scheme, it can be seen that the advantages of the present invention are:
[0042] The motion redirection-based sign language action migration method provided by the present invention uses a cyclic generative adversarial network for unsupervised training. It specifically constructs an encoder model, a decoder model, and a discriminator model, and configures the encoder model as a motion encoder, a static encoder, and a latent encoder; wherein the motion encoder is configured as follows: the input is skeleton original sign language animation data, and the output is encoded skeleton action information; the static encoder is configured as follows: the input is skeleton space static data, and the output is encoded skeleton structure information; the latent encoder is configured as follows: the skeleton action information is decoupled from the skeleton structure information to extract abstract sign language actions; the decoder model is configured as follows: the sign language abstract actions and skeleton structure information are redirected to generate skeleton reconstructed sign language animation data; the discriminator model is configured as follows: the input is the skeleton original sign language animation data, the skeleton reconstructed sign language animation data, and the skeleton space static data, and the output identification results include a first identification result of the skeleton original sign language animation data and the skeleton space static data, and a second identification result of the skeleton reconstructed sign language animation data and the skeleton space static data, thereby solving the problem of difficulty in obtaining paired training data. In addition, compared with traditional methods, the present invention does not require the hierarchical structure of the source skeleton and the target skeleton to be consistent. As long as the skeleton is a human-shaped structure with a head, hands, and feet, the scope of application of motion migration is wider. At the same time, it has significant advantages when the body proportions of the source skeleton and the target skeleton are significantly different. Not only are the gestures after migration natural and coherent, but the position accuracy is also high, and the sign language expression is reliable and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 The overall structure of the sign language action redirection model is shown;
[0044] Figure 2 A schematic diagram of the process of a sign language action transfer method based on motion redirection is shown;
[0045] Figure 3 The attention layer structure diagram is shown;
[0046] Figure 4 A sign language animation dataset is shown;
[0047] Figure 5 Different types of skeleton size structures are shown;
[0048] Figure 6 A visual comparison of the present invention and other models is shown;
[0049] Figure 7a-7c Shows the right hand velocity curve when transferring the original motion to skeletons of different sizes. DETAILED DESCRIPTION
[0050] In order to make the above features and effects of the present invention more clearly understood, embodiments are given below and described in detail with reference to the accompanying drawings.
[0051] This paper proposes a sign language action transfer method based on motion redirection. The transfer model is a recurrent generative adversarial network composed of an encoder, a decoder, and a discriminator. The overall network structure is as follows: Figure 1 As shown, Figure 1 The overall structure of the sign language action redirection model is shown. Figure 2 A flowchart of a sign language action transfer method based on motion redirection is shown.
[0052] Specifically, a sign language motion transfer method based on motion redirection includes:
[0053] S1. Construct an encoder model, where the encoder model is configured as a motion encoder, a static encoder, and a latent encoder.
[0054] Among them, the motion encoder is configured as follows: the input is the original skeleton sign language animation data, and the output is the encoded skeleton action information. Through the motion encoder, the purpose of aligning the original action with the redirected action is achieved.
[0055] Specifically, motion encoder Encoder (Q) It consists of several convolutional layers, whose input is the original sign language animation data Q, and the convolution kernel size is k×k. The original sign language animation data undergoes multiple two-dimensional convolutions and uses the Leakey ReLU activation function to obtain the skeleton action information L (Q) .
[0056] L (Q) =conv(k×k)(Q)
[0057]
[0058] Among them, the static encoder is configured as follows: the input is the static data of the skeleton space, including the initial spatial coordinates of the bone nodes, etc., and the output is the encoded skeleton structure information. Through the static encoder, the recognition of different skeleton structures and proportions is completed.
[0059] Specifically, static encoder Encoder (S) The convolution operation is analogous to the motion encoder Encoder (Q) , whose input is the skeleton space static data S. The skeleton space static data S undergoes multiple one-dimensional convolutions and performs activation operations to generate skeleton structure information l (s) .
[0060] L (s) =conv(k)(S)
[0061]
[0062] The latent encoder is coupled to the motion encoder and the static encoder, and is configured to decouple the skeleton action information from the skeleton structure information to extract abstract sign language actions.
[0063] Specifically, the latent encoder Latent converts the skeleton action information l (q) and skeleton structure information L (S) Then, the convolutional network is used to generate the action result L that is independent of the skeleton structure and extract the abstract sign language action. The specific formula is as follows:
[0064] L = conv(k×k)(L (Q) +broadcast(L (S) W E +b E ))
[0065] Among them, W E represents the weight matrix, b E Represents the bias matrix, and broadcast means broadcasting the vector into a matrix.
[0066] S2. Construct a decoder model, wherein the decoder model is coupled to the encoder model and configured to redirect the sign language abstract action and skeleton structure information to generate skeleton-reconstructed sign language animation data.
[0067] Specifically, the decoder decodes the action result L and the skeleton structure information L (s) Re-decode into reconstructed sign language animation data to achieve sign language movement redirection.
[0068] First, a convolution operation similar to the encoder is performed to generate the matrix U.
[0069] U=conv(k×k)(conv(k×k)(L)+broadcast(L (S) W D +b D ))
[0070] Among them, WD represents the weight matrix, b D represents the deviation matrix;
[0071] Then, U is sent to the attention layer to calculate the attention score and reconstruct the attention layer structure. Figure 3 As shown, the model uses three different weights W Q 、W K and W V Get three matrices Query, Key, and Value.
[0072] Query=W Q U,Key=W K U,Value=W V U
[0073] Then, the matrix Query is transposed and multiplied by the matrix Key, and the softmax function is used to calculate the attention score map Score. The specific formula is as follows, where Score ji Indicates the degree of attention that the j-th bone node pays to the i-th bone node during the redirection process.
[0074]
[0075]
[0076] Finally, the Leakey Relu activation function is used to form the reconstructed sign language animation data Q with the same size as the original sign language animation data Q (rec) , the specific calculation formula is as follows, where δ is the weight coefficient.
[0077] Q (rec) =δValueScore+U
[0078]
[0079] S3. Construct a discriminator model, which is coupled to the encoder model and the decoder model and configured as follows: input is the skeleton original sign language animation data, the skeleton reconstructed sign language animation data, and the skeleton space static data; output identification results include a first identification result about the skeleton original sign language animation data and the skeleton space static data, and a second identification result about the skeleton reconstructed sign language animation data and the skeleton space static data.
[0080] Specifically, the structure of the discriminator is similar to that of the encoder. Its input is the original sign language animation data Q or the reconstructed sign language animation data Q. (rec), skeleton space static data S, the output is the identification result C(Q,S), which includes the first identification result about the skeleton original sign language animation data and the skeleton space static data, and the second identification result about the skeleton reconstructed sign language animation data and the skeleton space static data. The main difference between the discriminator and the encoder is the addition of a sigmoid activation function at the exit.
[0081] O=conv(k×k)(Q)+broadcast(conv(k)(S)W C +b C )
[0082]
[0083] Among them, W C represents the weight matrix, b C represents the deviation matrix.
[0084] S4. Construct the target loss function. The constructed target loss function includes the shallow loss function, reconstruction loss function, adversarial loss function, and terminal loss function.
[0085] Specifically, let's take the redirection process of migrating a sign language action from source skeleton A to target skeleton B as an example. Here, the number of skeletal nodes in source skeleton A is N, the total number of frames of the sign language animation is T, and the number of skeletal nodes in target skeleton B is M. Obviously, the total number of frames of the sign language animation after migrating from source skeleton A to target skeleton B is also T. The original sign language animation data of source skeleton A can be expressed as:
[0086]
[0087] in It is composed of the posture quaternions of each skeleton node and the position coordinates of the root node in the k-th frame of the sign language animation. The specific form is expressed as:
[0088]
[0089] in Indicates the quaternion corresponding to the i-th bone node, Represents the position coordinates of the root node. Similarly, the skeleton space static data of the source skeleton A can be expressed as:
[0090]
[0091] The original sign language animation data migrated to the target skeleton B is represented as
[0092]
[0093] First, in theory, a good model may result in the original sign language animation data Q of the source skeleton A being different due to different skeleton parameters. A and the original sign language animation data Q of the target skeleton B B The difference is large, but the action matrix obtained by mapping it to the common latent space is the skeleton action information of the source skeleton and the target skeleton's skeleton motion information should be the same, so in order to align the latent space vectors, the present invention introduces the latent loss function l ltc , the shallow loss function is used to constrain the skeleton motion information of both the source skeleton and the target skeleton.
[0094] Specifically, the shallow loss function is determined based on the skeleton motion information generated by inputting the original sign language animation data of the source skeleton into the motion encoder and the skeleton motion information generated by inputting the original sign language animation data of the redirected target skeleton into the motion encoder. The shallow loss function is expressed as:
[0095]
[0096] in, The original sign language animation data Q representing the source skeleton A A Input the skeleton motion information generated by the motion encoder, The original sign language animation data Q representing the target skeleton B B Input the skeleton motion information generated by the motion encoder, l ltc Represents a shallow loss function.
[0097] In addition, the present invention draws on the structure of the cyclic generative adversarial network to realize the training of the model by reconstructing the original sign language animation data Q during the encoding and decoding process. Similarly, the model also converts the encoded skeleton action information L (Q) Input to the decoder to generate the reconstructed sign language animation data Q corresponding to the original sign language animation data Q (rec) In theory, the reconstructed sign language animation data Q (rec) It should be close to the original sign language animation data Q, so the reconstruction loss l is introduced rec The reconstruction loss function is used to constrain the reconstruction information of both the source skeleton and the target skeleton.
[0098] Specifically, the reconstruction loss function is determined based on the original sign language animation data and the reconstructed sign language animation data of the source skeleton, and the original sign language animation data and the reconstructed sign language animation data of the target skeleton. The reconstruction loss function is expressed as:
[0099]
[0100] Among them, Q A 、 They represent the original sign language animation data and the reconstructed sign language animation data of the source skeleton A, Q B 、 They represent the original sign language animation data and the reconstructed sign language animation data of the target skeleton B, respectively. rec represents the reconstruction loss function.
[0101] In addition, to ensure the reconstruction of sign language animation data Q (rec) It is a real sign language animation data. This invention introduces the adversarial loss l adv To determine the correctness of the reconstruction result. The adversarial loss function is used to constrain the adversarial information between the source skeleton and the target skeleton.
[0102] Specifically, the adversarial loss function is determined based on the first identification result and the second identification result of the source skeleton and the first identification result and the second identification result of the target skeleton. The adversarial loss function is expressed as:
[0103]
[0104] in, Represents the skeleton reconstruction sign language animation data of the source skeleton A With the skeleton space static data S A The second identification result, C A (Q A ,S A ) represents the original sign language animation data Q of the source skeleton A A With the skeleton space static data S A The first identification result; Represents the target skeleton B's skeleton reconstruction sign language animation data With the skeleton space static data S B The second identification result, C B (Q B ,S B ) represents the original sign language animation data Q of the target skeleton B B With the skeleton space static data S B The first identification result, l adv represents the adversarial loss function.
[0105] In addition, in order to ensure that the movement speed of the skeleton end remains unchanged, the present invention introduces the end loss l ee Constrain the movement speed of each skeletal joint at the end of the skeleton (including the end of the limb and the end of the hand) to ensure that the redirection result conforms to the normal movement law. The end loss function is used to constrain the movement speed of each skeletal joint at the end of the skeleton.
[0106] Specifically, the terminal loss function is determined based on the movement speed of each skeletal joint at the end of the source skeleton and the movement speed of each skeletal joint at the end of the target skeleton. The terminal loss function is expressed as:
[0107]
[0108] in, Represents the velocity of the i-th end bone joint of the source skeleton A, represents the velocity of the i-th end bone joint of the target skeleton B; h A Indicates the height of the source skeleton A (corresponding to the end of the source skeleton limb or hand), h B Indicates the skeleton height of the target skeleton B (corresponding to the end of the target skeleton limb or hand), l ee represents the terminal loss function.
[0109] Therefore, the final objective loss function l=l is formed ltc +αl rec +βl adv +γl ee , α, β, γ represent weight coefficients.
[0110] Next, the effectiveness of the model constructed by the present invention is verified by conducting experiments on a sign language animation dataset.
[0111] (1) Dataset
[0112] First, the sign language animation data used is from the National General Sign Language Dictionary. It was collected by a sign language teacher using motion capture equipment. The data set consists of 6707 sign language movements. Each sign language movement corresponds to a bvh format animation file with a frame rate of 60. Each frame of animation consists of 53 human bone nodes. The start and end of the sign language movement are represented by the natural raising and lowering of the arms. The specific form is as follows: Figure 4 shown.
[0113] In order to verify the model's ability to transfer different skeleton sign language movements, the experiment replaced the target skeleton with high-bones, normal-bones and low-bones skeletons respectively while ensuring that the source skeleton sign language movements were the same. Figure 5 A comparison chart of different types of skeleton sizes is shown, where the measurement range of skeleton height is the height from the heel to the head, and the measurement range of hand size is the length from the wrist to the middle finger when the hand is naturally hanging.
[0114] (2) Experimental results
[0115] When quantifying the results, the experiment of the present invention adopts the method of calculating the MSE error of the skeleton nodes to represent the error between the redirection result and the actual result. The experimental results are shown in Table 1.
[0116] Table 1 MSE error of performing redirection task on different skeleton sizes
[0117]
[0118] The experiment of this invention visually compares the proposed model with the existing SOTA model. To reflect the universality of the redirection results, the experiment shows the visualization results of each model performing the redirection task when the target skeleton is Normal-Bones. Figure 5 As shown in the figure, the leftmost image is the input original sign language animation data, and the remaining four columns of images represent the redirection results of different models, that is, the reconstructed sign language animation data. The dark skeleton represents the target result.
[0119] When performing sign language actions, the speed of the sign language actions will also affect the interpretation of the sign language meaning. Therefore, when performing the redirection task, whether the sign language speed of the original action can be normally transferred to the target skeleton is also an important indicator for judging the effectiveness of the model. The experiment of the present invention shows the curves of the hand speed and expected speed after redirection of each task in Table 1. The experiment decomposes the speed into three directions of X, Y, and Z, and after actual observation, the speed curves of the left and right hands for the same redirection task are not much different, and considering that the sign language actions are mainly with the right hand, only the hand speed curve of the right hand is shown. The specific experimental results are as follows. Figure 7a-7c As shown, Figure 7a 、 7b 7c and 7d respectively show the results of redirecting to High-Bones, redirecting to Low-Bones, and redirecting to Normal-Bones.
[0120] (3) Result analysis
[0121] As shown in Table 1, the average error of the proposed model is much smaller than that of NKN, PMnet, and SAD, indicating that the proposed model has the best performance among all models for the redirection task of skeletons of different sizes.
[0122] Depend on Figure 6 It can be seen that Figure 6The visualization comparison results of the model of the present invention and other models are shown in the figure. The redirection result of the NKN model is higher than the target hand position, and the redirection result of the PMnet model is lower than the target hand position. When the palms are close together in the second and third frames of the original sign language animation, the palms redirected by the PMnet model cannot be closed. In addition, the redirection result of the PMnet model also has obvious errors in the standing posture. The redirection result of the SAD model is better than the NKN model and the PMnet model. The movement trajectory of its hands is closer to the real trajectory when performing sign language actions, but it also shows a more obvious hand position offset, and there are also more obvious errors in the standing posture. The redirection result of the model proposed in the present invention is closer to the target result when performing sign language actions and standing postures than the other three models. There is no phenomenon of palm separation when the palms are closed, and the finger shape is clearly visible without obvious deformation.
[0123] Figure 7a-7c The right-hand velocity curves are shown when migrating the original motion to skeletons of different sizes. The NKN and PMnet models exhibit significant speed fluctuations, with their maximum abnormal speeds significantly different from the target speed. The SAD model, while lacking significantly different abnormal speed points compared to the NKN and PMnet models, exhibits speed fluctuations at certain times, and its curve shape differs significantly from the target curve. The velocity curve of the proposed model is relatively smooth, lacking speed fluctuations, and its trajectory closely matches the target curve.
[0124] In summary, the method for transferring sign language actions based on motion redirection provided by the present invention uses a cyclic generative adversarial network for unsupervised training, which solves the problem of difficulty in obtaining paired training data. In addition, the present invention includes the hand skeleton in the redirection scope during redirection and introduces an attention mechanism to solve the problem of finger deformation during redirection, thereby improving the accuracy of hand redirection. In addition, compared with traditional methods, the present invention does not require the hierarchical structure of the source skeleton and the target skeleton to be consistent. As long as the skeleton is a human-shaped structure with a head, hands, and feet, the scope of application of motion transfer is wider. At the same time, it has significant advantages when the body proportions of the source skeleton and the target skeleton are significantly different. Not only are the gestures after transfer natural and coherent, but the position accuracy is high, and the sign language expression is reliable and accurate. In addition, in the reasoning stage, the present invention only requires a motion encoder, a static encoder, and a latent encoder of the source skeleton, and a static encoder and decoder of the target skeleton. There is no need for loops and adversarial processes, which reduces the complexity of the reasoning process, that is, reduces the execution time of sign language action transfer.
[0125] Furthermore, the above-described embodiments of the present invention can be applied to terminal devices that utilize the motion redirection-based sign language action transfer method. These terminal devices can include personal terminals and host terminals, and the present invention is not limited thereto. The terminal can support operating systems such as Windows, Android, iOS, and Windows Phone.
[0126] A sign language action transfer device based on motion redirection, a sign language action transfer method based on motion redirection can be applied to personal terminals and host terminal devices, which can be realized by Figure 2 The sign language motion transfer method based on motion redirection shown, the sign language motion transfer device based on motion redirection provided in the embodiment of the present application can realize each process realized by the above-mentioned sign language motion transfer method based on motion redirection.
[0127] A sign language action transfer device based on motion redirection, comprising at least:
[0128] An encoder module, configured to construct an encoder model, wherein the encoder model is configured as a motion encoder, a static encoder, and a latent encoder;
[0129] The motion encoder is configured to: input the original skeleton sign language animation data and output the encoded skeleton motion information;
[0130] The static encoder is configured as follows: input is skeleton space static data, and output is encoded skeleton structure information;
[0131] The latent encoder is coupled to the motion encoder and the static encoder, and is configured to: decouple the skeleton action information from the skeleton structure information to extract the abstract sign language action;
[0132] A decoder module is used to construct a decoder model, the decoder model is coupled to the encoder model, and is configured to: redirect the sign language abstract action and skeleton structure information to generate skeleton-reconstructed sign language animation data;
[0133] The discriminator module is used to construct a discriminator model, which is coupled to the encoder model and the decoder model and configured to: input the skeleton original sign language animation data, the skeleton reconstructed sign language animation data, and the skeleton space static data, and output identification results including a first identification result regarding the skeleton original sign language animation data and the skeleton space static data, and a second identification result regarding the skeleton reconstructed sign language animation data and the skeleton space static data.
[0134] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be applied, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0135] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A sign language action transfer method based on motion redirection, characterized in that: include: Constructing an encoder model, wherein the encoder model is configured as a motion encoder, a static encoder, and a latent encoder; The motion encoder is configured to: input the original skeleton sign language animation data and output the encoded skeleton motion information; The static encoder is configured as follows: input is skeleton space static data, and output is encoded skeleton structure information; The latent encoder is coupled to the motion encoder and the static encoder, and is configured to: decouple the skeleton action information from the skeleton structure information to extract the abstract sign language action; Constructing a decoder model, the decoder model being coupled to the encoder model and configured to: redirect the sign language abstract action and skeleton structure information to generate skeleton-reconstructed sign language animation data; A discriminator model is constructed, and the discriminator model is coupled to the encoder model and the decoder model, and is configured to: input the skeleton original sign language animation data, the skeleton reconstructed sign language animation data, and the skeleton space static data, and output identification results including a first identification result regarding the skeleton original sign language animation data and the skeleton space static data, and a second identification result regarding the skeleton reconstructed sign language animation data and the skeleton space static data.
2. The method according to claim 1, characterized in that Constructing a target loss function, wherein the target loss function includes a shallow loss function, and the shallow loss function is used to constrain skeleton motion information of both the source skeleton and the target skeleton; The shallow loss function is determined according to the skeleton action information generated by inputting the original sign language animation data of the source skeleton into the motion encoder and the skeleton action information generated by inputting the original sign language animation data of the redirected target skeleton into the motion encoder.
3. The method according to claim 2, characterized in that The shallow loss function is expressed as: in, The original sign language animation data Q representing the source skeleton A A Input the skeleton motion information generated by the motion encoder, The original sign language animation data Q representing the target skeleton B B Input the skeleton motion information generated by the motion encoder, l ltc Represents a shallow loss function.
4. The method according to claim 2 or 3, characterized in that The target loss function also includes a reconstruction loss function, which is used to constrain the reconstruction information of both the source skeleton and the target skeleton; The reconstruction loss function is determined according to the original sign language animation data and the reconstructed sign language animation data of the source skeleton, and the original sign language animation data and the reconstructed sign language animation data of the target skeleton.
5. The method according to claim 4, characterized in that The reconstruction loss function is expressed as: Among them, the Q A 、 They represent the original sign language animation data and the reconstructed sign language animation data of the source skeleton A, Q B 、 They represent the original sign language animation data and the reconstructed sign language animation data of the target skeleton B, respectively. rec represents the reconstruction loss function.
6. The method according to claim 5, characterized in that The target loss function also includes an adversarial loss function, which is used to constrain the adversarial information between the source skeleton and the target skeleton; The adversarial loss function is determined based on the first identification result and the second identification result of the source skeleton and the first identification result and the second identification result of the target skeleton.
7. The method according to claim 6, characterized in that The adversarial loss function is expressed as: in, Represents the skeleton reconstruction sign language animation data of the source skeleton A With the skeleton space static data S A The second identification result, C A (Q A ,S A ) represents the original sign language animation data Q of the source skeleton A A With the skeleton space static data S A The first identification result; Represents the target skeleton B's skeleton reconstruction sign language animation data With the skeleton space static data S B The second identification result, C B (Q B ,S B ) represents the original sign language animation data Q of the target skeleton B B With the skeleton space static data S B The first identification result, l adv represents the adversarial loss function.
8. The method according to claim 7, characterized in that The target loss function also includes an end loss function, which is used to constrain the movement speed of each bone joint at the end of the skeleton; The terminal loss function is determined according to the movement speed of each skeletal joint at the end of the source skeleton and the movement speed of each skeletal joint at the end of the target skeleton.
9. The method according to claim 8, characterized in that The terminal loss function is expressed as: in, Indicates the speed of each bone joint at the end of source skeleton A, Indicates the velocity of each bone joint at the end of the target skeleton B; h A Indicates the height of the source skeleton A, h B Indicates the skeleton height of the target skeleton B, l ee represents the terminal loss function.
10. A sign language action transfer device based on motion redirection, characterized in that: The method for sign language action transfer based on motion redirection according to any one of claims 1 to 9 comprises at least: An encoder module, configured to construct an encoder model, wherein the encoder model is configured as a motion encoder, a static encoder, and a latent encoder; The motion encoder is configured to: input the original skeleton sign language animation data and output the encoded skeleton motion information; The static encoder is configured as follows: input is skeleton space static data, and output is encoded skeleton structure information; The latent encoder is coupled to the motion encoder and the static encoder, and is configured to: decouple the skeleton action information from the skeleton structure information to extract the abstract sign language action; A decoder module is used to construct a decoder model, the decoder model is coupled to the encoder model, and is configured to: redirect the sign language abstract action and skeleton structure information to generate skeleton-reconstructed sign language animation data; The discriminator module is used to construct a discriminator model, which is coupled to the encoder model and the decoder model and configured to: input the skeleton original sign language animation data, the skeleton reconstructed sign language animation data, and the skeleton space static data, and output identification results including a first identification result regarding the skeleton original sign language animation data and the skeleton space static data, and a second identification result regarding the skeleton reconstructed sign language animation data and the skeleton space static data.