Construction method and device of sports style migration system
By converting a single label into a strong text description and extracting features using the CLIP model, semantic alignment in action style transfer is achieved, and the problems of unnatural action style transfer and style expression deviation in the existing technology are solved, which significantly improves the naturalness of the action and the accuracy of style transfer.
Patent Information
- Application Number
- CN202510144641.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-23
AI Technical Summary
In the transfer of movement style, the prior art has problems of unreasonable joint angle rotation and style expression deviation, making it difficult to effectively maintain the naturalness and consistency of the movement.
By converting a single label into a stronger text description, and using the CLIP model to extract high-dimensional feature vectors, calculating the similarity between semantic features and action features, achieving semantic alignment, thereby improving the effect of style transfer.
It significantly improves the naturalness of the generated action and the accuracy of style transfer, and is better than the performance of the prior art in content consistency and style consistency.
Smart Images

Figure CN120032425A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a method and device for constructing a motion style transfer system. Background Art
[0002] Human motion style plays a crucial role in conveying a character's emotions and personality through their movements. Humans are highly sensitive to subtle changes in motion style. For example, one can often infer a person's emotional state - such as happiness or sadness - simply from their walking state. In recent years, there has been a growing interest in generating stylized motion due to the importance of stylization in animated characters and avatars, especially in the fields of computer graphics and virtual reality. However, the process of manually creating diverse stylized motion is time-consuming and laborious, making it an impractical approach for many applications. A more efficient solution to this problem is to apply style transfer techniques between different motion sequences, leveraging existing motion capture databases. The task of motion style transfer is to apply the style features of one motion sequence (the style motion) to a target motion sequence (the content motion). Several successful approaches have been proposed to address this problem, each of which has contributed to improving the effectiveness of human motion style transfer.
[0003] However, ensuring that the generated actions are natural, consistent, and maintain the style efficiently in action style transfer remains a major challenge. To create realistic and expressive animations, the feature encoding of the transferred actions should fully preserve the authenticity and rich expressiveness of the input actions. Although the action content can be quantified and represented by pose data, style is a more abstract concept that is difficult to describe by direct quantification. Leveraging the success of the contrastive language-image pre-training (CLIP) model, it is found that introducing semantic information can significantly enhance the expressiveness of the feature generation process, making the generated actions more distinctive and expressive. Semantic information plays a key role in feature encoding, enabling the model to capture the emotions and intentions behind the input actions. Nevertheless, most current studies tend to underestimate the integration of semantic information. For example, MoST combines the Siamese action encoder and the attention mechanism, but does not fully utilize the semantic information behind the actions. Although MoST shows strong capabilities in feature extraction, the lack of semantic fusion may be one of the reasons for some unreasonable input results (such as joint angles). FineStyle achieves good results by embedding the content labels of content actions into the encoding process of style actions. However, it does not fully utilize the style labels of content actions or labels related to style actions. It mainly transfers the motion pattern from the source action to the target action without fully capturing the semantic style behind the action, resulting in deviation and inconsistency in style expression in the generated actions. Summary of the invention
[0004] The technical problem to be solved by the present invention is to improve the problems that still exist in the prior art, such as the unreasonable rotation of the joint angle of MoST and the deviation problem of the expression of the action style generated by FineStyle. The present invention converts a single label into a more expressive text description, and then feeds these descriptions into the text encoder of CLIP to obtain a high-dimensional feature vector. Next, the similarity between the semantic feature vector and the action feature vector is calculated, thereby successfully utilizing the powerful feature extraction capability of CLIP. The advantage of this method is that, with the help of the powerful semantic extraction function of CLIP, the feature vector generated by the model also has extremely strong semantic information. This not only improves the naturalness of the generated action, but also makes the effect of style transfer more accurate. In the two evaluation indicators of content consistency and style consistency, the present invention is superior to the best existing model in the field (the content consistency and style consistency indicators of the present invention are [3.0, 13.3], better than [8.5, 63.0] of MoST and [4.8, 14.0] of FineStyle), and has a significant improvement.
[0005] The solution of the present invention:
[0006] A method for constructing a motion style transfer system, which transfers the style of a target motion sequence by using a CLIP semantic alignment-assisted method, includes the following steps:
[0007] Step 1: Capture accurate 3D human body postures through multiple cameras and high-precision motion capture systems, annotate the action content and style of each action sequence, and splice the action sequence annotations into text descriptions ;
[0008] Step 2: Represent the action style input of the action sequence in step 1 with joint positions and use it as the Transformer encoder The input is encoded and further activated by a standard three-layer MLP to obtain the action style feature ;
[0009] Step 3: Represent the action content of the action sequence in step 1 with joint rotations and use it as a one-dimensional temporal convolutional encoder Input, generate action content features ;
[0010] Step 4: Describe the text in step 1 As a pre-trained CLIP text encoder The input is encoded and then goes through a linear mapping layer to obtain the content semantic features and style semantic features ;
[0011] Step 5: The feature alignment module further calculates the content semantic features in step 4 and the action content features in step 3 The cosine similarity between them, and the style semantic features in step 4 and the action style features in step 2 The cosine similarity between them is used to achieve semantic alignment and output the semantic alignment loss. Used to optimize the training of motion style transfer systems;
[0012] Step 6: The action generator G integrated with the AdaIN layer obtains output features with style transfer by adjusting the mean and variance of the action content features and action style features, thereby outputting an action sequence that meets the target style and content features;
[0013] Step 7: The discriminator distinguishes the generated action sequence from the real sequence by analyzing the spatiotemporal consistency, motion characteristics and style consistency.
[0014] A device for constructing a motion style transfer system, used for performing style transfer on a target motion sequence, comprising:
[0015] The action sequence dataset processing module is used to capture accurate 3D human body postures through multiple cameras and high-precision motion capture systems, and annotate the action content and action style of each action sequence, and splice the action sequence annotations into a text description that can more accurately reflect the action sequence. ;
[0016] The Transformer encoder module is used to encode the input action style features. After encoding, it is further activated by a standard three-layer MLP to obtain the action style features. ;
[0017] One-dimensional temporal convolutional encoder module, used to extract and encode the input action content to generate action content features;
[0018] The CLIP text encoder module is used to receive text descriptions, encode them, and then pass them through a linear mapping layer to obtain content semantic features. and style semantic features ;
[0019] Feature alignment module, which is used to calculate the cosine similarity between content semantic features and action content features, and the cosine similarity between style semantic features and action style features to achieve semantic alignment, and output the semantic alignment loss Used to optimize the training of motion style transfer systems;
[0020] The action generator module is used to adjust the mean and variance of the action content features and action style features through the AdaIN layer to obtain output features with style transfer, thereby outputting the generated action sequence;
[0021] The discriminator optimization training module is used to distinguish generated action sequences from real action sequences to improve the authenticity and style transfer effect of the generated action sequences.
[0022] The beneficial effects of the present invention are as follows:
[0023] 1. The present invention shows a significantly better effect than the prior art in the quantitative evaluation of model performance. A comprehensive and scientific quantitative index system is adopted, focusing mainly on the three core dimensions of generation quality, content retention and style performance, and a detailed quantitative evaluation of the model is carried out. From the experimental results, the method shows extremely low values in the two key indicators of content consistency (CC) and style consistency (SC++). Specifically, the significant reduction in content consistency (CC) shows that the present invention can effectively maintain the integrity of the input action content during the style transfer process and avoid the loss of content information; while the significant reduction in the value of style consistency (SC++) shows that the generation result accurately reflects the target style characteristics and demonstrates the model's powerful style transfer ability. The present invention can stably complete the style transfer task whether the input action content is the same or when there are significant differences in the input action content. The experimental results show that the model can not only achieve a smooth and natural transition between different styles, but also maintain a high degree of consistency in the action content information during the transfer process, which effectively avoids the common content distortion problem.
[0024] The following table shows the key indicator results of various methods when evaluated on the Xia dataset. These indicators are used to comprehensively measure the comprehensive performance of the model in terms of generation quality, content retention, and style expression. The following are the detailed data:
[0025]
[0026] 2. The present invention has made important contributions to the field of action style transfer in terms of the construction and expansion of datasets, broadened the selection range of available datasets in this field, and provided richer resource support for model training and research. The present invention has deeply optimized the existing data, especially reconstructed the BFA dataset. By dividing long sequence actions into smaller subsequences and accurately labeling action labels and style labels for each subsequence, a new dataset that is highly adapted to the training needs of dual-label models is formed. This improvement not only improves the structured level of the data, but also enhances the flexibility of the data, enabling it to more efficiently support model learning and training. The expanded BFA dataset is closer to actual application needs in terms of the comprehensiveness and complexity of the data, and provides a reliable data foundation for in-depth research in the field of style transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A flow chart of a method for constructing a motion style transfer system of the present invention;
[0028] Figure 2 The structure framework diagram of the construction device of the motion style transfer system of the present invention.
[0029] Figure numerals: action sequence dataset processing module 71, Transformer encoder module 72, one-dimensional temporal convolution encoder module 73, CLIP text encoder module 74, feature alignment module 75, action generator module 76, discriminator optimization training module 77. DETAILED DESCRIPTION
[0030] The embodiment of the present invention provides a method for constructing a motion style transfer system. In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] The present invention relates to a method for constructing a motion style transfer system, comprising: capturing accurate 3D human body postures through multiple cameras and a high-precision motion capture system, and annotating the action content and action style of each action sequence. Before training, the annotations of the action sequence will be spliced into a text description that can more accurately reflect the action sequence. ; The style action input of the action sequence is represented by joint positions and used as the Transformer encoder The input is encoded and further activated by a standard three-layer MLP to obtain the style action feature ; Represent the content action input as joint rotation and use it as a one-dimensional temporal convolutional encoder Input, generate content action features ; Text description As a pre-trained CLIP text encoder The input is encoded and passed through a linear mapping layer to obtain semantic features and ; The feature alignment module further calculates the semantic features and action characteristics Cosine similarity between , to achieve semantic alignment and output semantic alignment loss Used to optimize network training; AdaIN decoding achieves style transfer by adjusting the mean and variance of the feature vector so that the generated motion sequence can conform to the specified style characteristics.
[0032] The present invention can be applied in various fields such as animation and games, and accurate style migration tasks can be achieved by inputting the content action input and style action input to be migrated. Specifically: given a content input such as "running happily" and a style input such as "walking sadly", the action output by the network will be "running sadly", that is, it retains the original content expression (running) while combining the new style expression (sadness). In the experiment, it is assumed that each motion clip in the training set is assigned clear content and style labels, but there is no requirement that these clips appear in pairs. This means that there is no need to rely on motion pairs with the same content but different styles for training, which greatly simplifies the data requirements.
[0033] The execution environment of the invention uses a Core 16-core computer with a 5.2GHZ central processing unit and 32G bytes of memory. It is trained on the RTX 3090GPU using PyTorch using the Adam optimizer with a batch size of 16, and the number of training iterations of the model is 300,000 times. When the Xia dataset is used for training, the training time is about 5.4 hours, the inference time is 316.5ms, and the variance is 5075.7. At the same time, Python, C++ and other languages are used to construct the method program. Under the premise that the computer memory and video memory allow, the present invention can also be based on other execution environments, such as GeForce GTX 1080TI and other environments.
[0034] Embodiment 1
[0035] like Figure 1 As shown, the present invention provides a method for constructing a motion style transfer system, and the steps are as follows:
[0036] Step 1: Use multiple cameras and a high-precision motion capture system to capture accurate 3D human body postures, annotate the action content and style of each action sequence, and splice the action sequence annotations into a text description that more accurately reflects the action sequence. .
[0037] The above specifically includes: After capturing the 3D human posture, the action sequence is represented by the SMPL model, which is the most widely used model in the current field. The SMPL model (Skinned Multi-Person Linear Model) is a 3D human model based on linear blended shapes, which can generate three-dimensional human bodies of different shapes and postures using shape parameters (β) and posture parameters (θ). The model controls the body shape differences of the human body through 20 shape parameters, and 72 posture parameters represent joint rotations. SMPL is widely used in human modeling, posture estimation and animation production, and is efficient and flexible, adapting to a variety of application scenarios.
[0038] The action sequence is represented by the SMPL model as: , where a sequence of length T is used to represent motion data, etc. all represent motion data; j represents the number of joints in the action; the value of d depends on the type of input. is a textual description of the action sequence:
[0039] When representing an action, this parameter is defined using joint positions, which refer to the coordinates of each joint in 3D space, usually represented by a 3-dimensional vector, i.e. d=3;
[0040] When representing the action content, joint rotation is used and defined by unit quaternion. Quaternion is a commonly used method to represent 3D rotation. Compared with Euler angle, quaternion can avoid the universal joint lock problem. Therefore, in this case, the dimension d=4.
[0041] The action content label s and action style label c of each input action are processed and spliced into a natural language text description that can fully describe the action and posture , this process can be expressed by the following formula:
[0042] ,
[0043] in, Used to guide video description, indicating "a video about...", where person refers to the subject in the video. The text description shown above is only an example, and the present invention uses the above method to generate multiple different text descriptions.
[0044] The present invention experimentally tests the motion style transfer system on the two most commonly used datasets in the field of style transfer: the Xia dataset and the BFA dataset. The BFA dataset contains multiple action clips with the same style but different action content. To evaluate the performance of the model, each action clip is cropped and labeled according to the action content for use in the test. The action representation of both datasets contains 24 joints. The Xia dataset covers 8 styles and 6 actions, while the BFA dataset covers 16 styles and 10 actions. To ensure consistency, the frame rate of the action data is downsampled from 120fps to 30fps.
[0045] In order to explore the effectiveness of the CLIP-guided method in the model, two datasets containing actions and text descriptions, Xia-T and BFA-T, were constructed based on the Xia dataset and the BFA dataset.
[0046] The BFA dataset contains various action sequences with consistent action style features but diverse action contents, which are cropped and annotated according to the specific action content of each action clip.
[0047] The text description is generated by combining the action content label c and the action style label s. Specifically, the content and style labels are converted into natural language descriptions in the form of “a video of (s) person (c)”.
[0048] The text descriptions and action sequences together form the extended Xia-T and BFA-T datasets.
[0049] The present invention fully considers the specific requirements of the CLIP text encoder for input, and focuses on the accuracy and adaptability of the text content during design. The text description used can specifically and clearly express the action features and style features of the action sequence, so that the CLIP text encoder can efficiently parse and generate semantic features that match the input action and style. This text description not only improves the representativeness of the semantic features, but also effectively enhances the migration performance of the model under different action and style combinations, providing a reliable foundation for subsequent feature alignment and style migration tasks.
[0050] Step 2: Represent the action style sequence of the action sequence in step 1 with joint positions and use it as the Transformer encoder The input is encoded and further activated by a standard three-layer MLP (multi-layer perceptron) to obtain the action style feature. .
[0051] Combined with existing research results, the roles of joint angles and joint positions are deeply analyzed and reasonably distinguished to optimize the choice of content representation and style representation. Specifically, joint angles are proven to be more suitable for capturing the content information of motion because they can clearly describe the geometric relationship and overall structural characteristics of skeleton motion, ensuring that the model accurately models the core structure of the action. Joint positions, on the other hand, can capture subtle changes in motion style more meticulously due to their sensitivity to motion trajectory and detail changes, thus playing an important role in style representation.
[0052] In the above step 2: The Transformer encoder is also called a style encoder, which is a simplified Transformer designed by the present invention. The encoder is designed to better capture the style information in motion, thereby generating richer and more attractive style features. Compared with the commonly used Transformer model, it is optimized and the position encoding part is removed. Through this modification, unnecessary complexity is avoided and the model's ability to characterize motion style is improved. Experimental results also show that adding position encoding reduces the generalization ability of the model and affects the quality of the final generated results. Therefore, this part is omitted in the design to ensure the stability and efficiency of the model.
[0053] The Transformer encoder consists of multiple Transformer layers to capture long-range dependencies of style actions.
[0054] The Transformer encoder accepts the action sequence after the action style linear transformation. The encoder contains 4 layers and 8 attention heads without position encoding to extract action features. Then, the extracted initial features are further generated through a standard three-layer multi-layer perceptron (MLP) to generate the final action style features.
[0055] MLP is a standard three-layer multilayer perceptron. Each layer has a linear transformation. The first two layers use ReLU as the activation function, and the last layer uses log_softmax to calculate the logarithm of the classification.
[0056] Step 3: Represent the action content of the action sequence in step 1 with joint rotations and use it as a one-dimensional temporal convolutional encoder Input, generate action content features .
[0057] The 1D temporal convolution encoder consists of multiple layers of temporal 1D convolutions, which are followed by several residual blocks, which are used to map the input motion sequence to the representation code in the temporal content latent space, that is, project the motion content into a latent space to extract the temporal series features of the content action. During the encoding process, the input motion sequence first passes through the temporal 1D convolution layer to extract preliminary temporal features, and then further captures the complex temporal dependencies and local features through the residual block. Specifically, the 1D temporal convolution layer extracts the short-term dependencies of the motion content by capturing the local patterns of the time series data, while the residual block further enhances the feature expression ability, ensuring that the encoder can capture the global structural information of the motion. In order to ensure that the encoding process focuses on the content information of the motion itself, the intermediate temporal feature maps obtained by each layer of temporal 1D convolution will be subjected to instance normalization (IN). This normalization operation can effectively remove style-related deviations and strip the style characteristics of the input motion, so that the content encoding can focus on the core structural information describing the motion.
[0058] Step 4: Describe the text in step 1 As a pre-trained CLIP text encoder The input is encoded and then goes through a linear mapping layer to obtain the content semantic features and style semantic features .
[0059] Specifically, we preprocess the semantic labels of action content and action style and convert them into natural language text descriptions that can fully describe the action and style characteristics. The generated text description is further processed by the CLIP text encoder and a fully connected (FC) layer to obtain semantic content features. and semantic style features Specifically, given the text description of the action content, the content semantic features Learn as follows:
[0060] ,
[0061] in, Refers to the pre-trained CLIP text encoder.
[0062] Given a text description of an action style, obtain the style semantic features as follows:
[0063] .
[0064] The CLIP text encoder is a pre-trained text encoder used to map text representations to a latent space. The CLIP text encoder uses the ViT-B / 32 version, which is pre-trained on a large-scale dataset of more than 400 million image-text pairs and demonstrates excellent feature extraction capabilities. Based on the Transformer architecture, the CLIP text encoder can capture deep semantic information from a large amount of text and image data, and accurately extract the motion and style features contained in the input text. Although the model is not specifically designed for motion style generation tasks, its strong generalization ability provides a solid technical foundation for the present invention, enabling it to efficiently adapt to the style transfer requirements of motion sequences. This wide applicability and deep semantic understanding capabilities make the CLIP text encoder indispensable in the motion style transfer system.
[0065] Step 5: The feature alignment module further calculates the content semantic features in step 4 and the action content features in step 3 The cosine similarity between them and the style semantic features in step 4 and the action style features in step 2 The cosine similarity between them is used to achieve semantic alignment and output the semantic alignment loss. Used to optimize motion style transfer system training.
[0066] Calculate the semantic features in step 4 And the action features in steps 2 and 3 The cosine similarity between them includes: using cosine distance to calculate the feature similarity between semantic features and action features generated from action sequences. As described in step 1, a text description is generated by annotating the action sequence, and then input into the CLIP text encoder to obtain semantic features. Depending on the action sequence, a one-dimensional temporal convolutional encoder or a Transformer encoder is used to extract action features. Semantic alignment loss The calculation formula is as follows:
[0067] , ,
[0068] ,
[0069] C represents the action content, and S represents the action style. Represents the loss value obtained by calculating feature similarity; represents semantic features. When i=c, it represents content semantic features. When i=s, it represents style semantic features. Indicates action characteristics; and is the corresponding weight parameter used to balance the similarity constraints between content and style.
[0070] The design of the generated action loss refers to existing work and adopts a variety of effective loss calculation methods to ensure the quality and consistency of the generated results. Specifically, it includes the following three core losses: content consistency loss , used to measure whether the generated action retains the original content characteristics; adversarial loss , improving the realism of generated actions through adversarial training; and feature matching loss , which is used to enhance the consistency between the generated action features and the target features.
[0071] The final generated action loss is composed of the above loss terms:
[0072] ,
[0073] in, , , This weight design ensures that the various loss functions complement each other while balancing the effects of content preservation and style transfer, thereby generating high-quality action sequences that meet the target requirements.
[0074] Step 6: The action generator G obtains the output features with style transfer and generates a new action sequence by adjusting the mean and variance of the action content features and action style features.
[0075] The action generator integrates the AdaIN layer, which is responsible for transferring style information to the content vector in the latent space and achieving style transfer by adjusting the mean and variance of the content vector so that the generated action sequence can meet the specified style characteristics.
[0076] AdaIN refers to Adaptive Instance Normalization (AdaIN), which is a normalization method used to dynamically map content features to the target style feature space and is widely used in style migration and generation tasks. The process includes the following steps:
[0077] Input: The action generator receives two sets of input features: action content features and action style features , representing the structural information of the action and the target style attributes respectively.
[0078] Style adjustment: characteristics of action content Normalize the instance, remove its original mean and standard deviation, and use the action style features Providing target mean and standard deviation, adjust the instance-normalized action content features: the final output features combine the structure of action content features and the statistical properties of action style features.
[0079] After the above adjustments, the generated features not only retain the structural information of the action content features, but also integrate the statistical properties of the action style features. By dynamically adjusting the feature distribution, the AdaIN decoder achieves a deep fusion of content and style, providing an effective mechanism for generating action sequences that meet the target style.
[0080] The style transfer process is as follows:
[0081] ,
[0082] in, is a one-dimensional temporal convolutional encoder, is a sequence of action contents, is the Transformer encoder, is an action style sequence, MLP is a multi-layer perceptron, and G is an action generator.
[0083] Step 7: The discriminator distinguishes the generated action sequence from the real sequence by analyzing the spatiotemporal consistency, motion characteristics and style consistency. The specific methods include:
[0084] Spatiotemporal consistency: Real action sequences usually have smooth temporal transitions and spatial consistency (for example, the angle changes between human joints are continuous without abnormal jumps or posture distortions). The discriminator can determine whether the action sequence is real data by analyzing these details.
[0085] Consistency of motion features and style: In addition to being consistent with the content of the real action (i.e., the type of motion, rhythm, etc.), the generated action sequence must also be able to maintain the target style (e.g., the motion style of a specific person). The discriminator determines the effect of style transfer by comparing the similarity between the generated action and the target style. For 3D action style transfer, style transfer involves the posture of the joints, the rhythm of the gait, the range of motion of the limbs, etc. The discriminator can use these features to evaluate the effect of style transfer.
[0086] By distinguishing the generated motion sequences from the real motion sequences, the feedback of the discriminator will prompt the motion style transfer system to optimize in terms of enhancing spatiotemporal consistency, maintaining content consistency and strengthening the style transfer effect to improve the authenticity and style transfer effect of the new motion sequences.
[0087] In the motion style transfer system, all model parameters that need to be optimized will be trained using the Adam optimizer. Specifically, the Adam optimizer automatically adjusts the learning step size of each parameter based on the gradient information of each parameter, thereby accelerating the optimization process, reducing training time, and improving performance. The learning rate is set to 0.0001. This learning rate has been debugged and verified experimentally to balance the convergence speed and stability, avoiding instability in model training due to excessively large learning rates, or slow training progress due to excessively small learning rates.
[0088] Embodiment 2
[0089] An embodiment of the present invention provides a device for constructing a motion style transfer system, which is used to transfer the style of a target action sequence. The target action sequence is defined using a bvh file that is most commonly used in the animation field. Figure 2 The construction framework diagram of the device is shown, and the device includes an action sequence data set processing module 71, a Transformer encoder module 72, a one-dimensional time convolution encoder module 73, a CLIP text encoder module 74, a feature alignment module 75, an action generator module 76, and a discriminator optimization training module 77.
[0090] The action sequence data set processing module 71 is used to capture accurate 3D human body postures through multiple cameras and a high-precision motion capture system, and to annotate the action content and action style of each action sequence, and to splice the action sequence annotations into a text description that can more accurately reflect the action sequence. ;
[0091] Transformer encoder module 72 is used to encode the input action style features, and after encoding, it is further activated by a standard three-layer MLP to obtain the action style features. ;
[0092] A one-dimensional temporal convolutional encoder module 73, used to extract and encode the input action content to generate action content features;
[0093] The CLIP text encoder module 74 is used to receive the text description, encode it, and obtain the content semantic features through a linear mapping layer after encoding. and style semantic features ;
[0094] The feature alignment module 75 is used to calculate the cosine similarity between the content semantic features and the action content features, and the cosine similarity between the style semantic features and the action style features to achieve semantic alignment, and output the semantic alignment loss Used to optimize the training of motion style transfer systems;
[0095] An action generator module 76, configured to obtain an output feature with style transfer and output a generated action sequence by adjusting the mean and variance of the action content feature and the action style feature;
[0096] The discriminator optimization training module 77 is used to distinguish new motion sequences from real motion sequences to improve the authenticity and style transfer effect of the generated motion sequences.
[0097] Action sequence dataset processing module 71: The main function is to pre-process the input raw action sequence data to meet the needs of subsequent model training. Specifically, it includes: standardizing the captured action sequence into a unified representation, such as a time series of joint positions or joint angles. Adding clear action content labels and action style labels to each action sequence to provide dual-label support for the training process. Expanding the dataset through operations such as segmentation, splicing, rotation or interpolation to improve the diversity and robustness of samples. Generate text descriptions based on the annotation information for CLIP text encoder input. These descriptions can clearly express action and style features.
[0098] Transformer encoder module 72: It is used to encode the input action style and obtain the action style feature. It captures the global dependency between actions through a multi-head attention mechanism and generates an action style feature representation with rich semantic information, thereby accurately describing the intrinsic characteristics of the target style action.
[0099] One-dimensional temporal convolutional encoder module 73: used to extract and encode the input action content. This module models the local and global dependencies in the time series through multiple layers of one-dimensional temporal convolution and residual blocks, extracts the latent feature representation of the content action, and provides an accurate description of the core structure of the motion.
[0100] CLIP text encoder module 74: used to extract semantic features of text descriptions. The text descriptions are generated based on the style of the target action sequence and the annotations of the actions, ensuring a comprehensive and accurate representation of the style and action information of the target motion. By encoding the natural language text, the CLIP text encoder generates a high-dimensional semantic feature representation, which provides a basis for subsequent feature alignment.
[0101] Feature alignment module 75: aligns the action style feature or action content feature with the corresponding action feature by using the semantic features generated by the CLIP text encoder. Specifically, this module ensures the effective combination of style and content features by calculating the similarity between the semantic features and the action style features generated by the Transformer encoder or the action content features generated by the one-dimensional temporal convolutional encoder.
[0102] Action Generator Module 76: The mean and variance of the action content features and the action style features are adjusted by integrating the adaptive instance normalization technique to dynamically combine the action style features with the action content features. This module uses the normalization operation to strip the style information from the action content features and reassigns the feature distribution through the target style features to generate the final migration result that conforms to the target style.
[0103] Discriminator optimization training module 77: used to distinguish the generated action sequence from the real action sequence to improve the authenticity and style transfer effect of the generated action sequence.
[0104] The above description is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the technical solution content of the present invention still belongs to the technology of the present invention.
Claims
1. A method for constructing a motion style transfer system, characterized in that: The style transfer of the target motion sequence is performed by using the CLIP semantic alignment assistance method, including the following steps: Step 1: Capture accurate 3D human body postures through multiple cameras and high-precision motion capture systems, annotate the action content and style of each action sequence, and splice the action sequence annotations into text descriptions ; Step 2: Represent the action style input of the action sequence in step 1 with joint positions and use it as the Transformer encoder The input is encoded and further activated by a standard three-layer MLP to obtain the action style feature ; Step 3: Represent the action content of the action sequence in step 1 with joint rotations and use it as a one-dimensional temporal convolutional encoder Input, generate action content features ; Step 4: Describe the text in step 1 As a pre-trained CLIP text encoder The input is encoded and then goes through a linear mapping layer to obtain the content semantic features and style semantic features ; Step 5: The feature alignment module further calculates the content semantic features in step 4 and the action content features in step 3 The cosine similarity between them, and the style semantic features in step 4 and the action style features in step 2 The cosine similarity between them is used to achieve semantic alignment and output the semantic alignment loss. Used to optimize the training of motion style transfer systems; Step 6: The action generator G integrated with the AdaIN layer obtains output features with style transfer by adjusting the mean and variance of the action content features and action style features, thereby outputting an action sequence that meets the target style and content features; Step 7: The discriminator distinguishes the generated action sequence from the real sequence by analyzing the spatiotemporal consistency, motion characteristics and style consistency.
2. The method for constructing a motion style transfer system according to claim 1, characterized in that: The CLIP text encoder is a pre-trained text encoder used to map text representations to a latent space. The CLIP text encoder uses the ViT-B / 32 version.
3. The method for constructing a motion style transfer system according to claim 1, characterized in that: The Transformer encoder includes multiple Transformer layers for capturing long-range dependencies of style actions; The Transformer encoder accepts the action sequence after the target style is linearly transformed. The encoder contains 4 Transformer layers and 8 attention heads without position encoding.
4. The method for constructing a motion style transfer system according to claim 1, characterized in that: The 1D temporal convolutional encoder consists of multiple 1D temporal convolutional layers followed by a series of residual blocks that project the action content into a latent space for extracting the time series features of the action content.
5. The method for constructing a motion style transfer system according to claim 1, characterized in that: Calculate the content semantic features in step 4 and the action content features in step 3 The cosine similarity between them, and the style semantic features in step 4 and the action style features in step 2 The cosine similarity between includes: using cosine distance to calculate the feature similarity between the semantic features and the action features generated from the action sequence; Semantic alignment loss The calculation formula is as follows: , , , Said C represents the action content, and S represents the action style; Represents the loss value obtained by calculating feature similarity; Represents semantic features, Indicates action characteristics; and is the corresponding weight parameter used to balance the similarity constraints between content and style.
6. The method for constructing a motion style transfer system according to claim 1, characterized in that: The action generator is a generator integrated with an AdaIN layer, which is responsible for transferring style information to the content vector in the latent space, and obtaining output features with style transfer by adjusting the mean and variance of the content vector, thereby outputting an action sequence that meets the target style and content features.
7. The method for constructing a motion style transfer system according to claim 6, characterized in that: The working process of the AdaIN layer includes: Input: The AdaIN decoder receives two sets of input features: action content features and action style features , representing the structural information of the action and the target style attributes respectively; Style adjustment: characteristics of action content Normalize the instance, remove its original mean and standard deviation, and use the action style features Providing target mean and standard deviation, adjust the instance-normalized action content features: the final output features combine the structure of action content features and the statistical properties of action style features.
8. The method for constructing a motion style transfer system according to claim 1, characterized in that: The discriminator is a multi-category discriminator method, which is used to determine whether the input action is a real sequence of a specific style, or a generated action sequence output by a motion style transfer system; when updating the discriminator, if the real action sequence is judged to be false, the discriminator will be punished; if the generated false action is judged to be real, it will also be punished; the discriminator will not be punished for misjudging actions of other styles.
9. A device for constructing a motion style transfer system, characterized in that: Used to transfer the style of target motion sequences, including: The action sequence dataset processing module is used to capture accurate 3D human body postures through multiple cameras and high-precision motion capture systems, and annotate the action content and action style of each action sequence, and splice the action sequence annotations into a text description that can more accurately reflect the action sequence. ; The Transformer encoder module is used to encode the input action style features. After encoding, it is further activated by a standard three-layer MLP to obtain the action style features. ; One-dimensional temporal convolutional encoder module, used to extract and encode the input action content to generate action content features; The CLIP text encoder module is used to receive text descriptions, encode them, and then pass them through a linear mapping layer to obtain content semantic features. and style semantic features ; Feature alignment module, which is used to calculate the cosine similarity between content semantic features and action content features, and the cosine similarity between style semantic features and action style features to achieve semantic alignment, and output the semantic alignment loss Used to optimize the training of motion style transfer systems; The action generator module is used to adjust the mean and variance of the action content features and action style features through the AdaIN layer to obtain output features with style transfer, thereby outputting the generated action sequence; The discriminator optimization training module is used to distinguish generated action sequences from real action sequences to improve the authenticity and style transfer effect of the generated action sequences.
10. The device for constructing a motion style transfer system according to claim 9, characterized in that: The discriminator is trained using the Adam optimizer with a learning rate set to 0.0001.
Citation Information
Cited By
Neural network-based motion capture actor fitness evaluation method and system
CN121170906A
Athlete action analysis and evaluation method and system based on computer vision
CN121354222A
Gesture generation model training method, gesture control method and related equipment
CN121438385A