Bone action style dynamic migration method and system based on generative adversarial network
Through the dynamic migration method of skeletal action style based on the adversarial generation network and combined with multi-constraint optimization technology, multiple shortcomings in skeletal action generation and migration in the existing technology are solved, and high-quality, diverse action sequence generation and accurate style transfer are achieved.
Patent Information
- Application Number
- CN202510076564.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing skeletal action generation and migration methods have the problems of strong data dependence, difficulty in generating diverse samples, difficulty in decoupling action content and style, insufficient spatial and temporal continuity of generated samples, and the difficulty in achieving high-quality action style transfer under limited data conditions.
The dynamic migration method of skeletal action style based on an adversarial generation network is adopted. By extracting skeletal action features, an action style migration network is built, and multi-constraint optimization is used, including adversarial loss, cyclic consistency loss, content maintenance loss, style maintenance loss, spatial consistency loss, time consistency loss and action flow smoothness loss to optimize the generator and discriminator.
It realizes the generation of high-quality and diverse skeletal action sequences under limited data conditions, which can accurately decouple the action content and style characteristics. The generated action sequences are continuous and natural in space and time, significantly improving the effect of action style transfer.
Smart Images

Figure CN119992024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of skeletal motion sequence generation and migration, and in particular to a skeletal motion style dynamic migration method and system based on a generative adversarial network. Background Art
[0002] In recent years, skeletal motion sequence analysis technology has made rapid progress in the fields of computer vision, artificial intelligence, and human motion research. The motion sequence analysis method based on skeletal data extracts the key joint features of the human skeleton and is widely used in animation production, virtual reality, behavior recognition, rehabilitation training and other scenarios. Traditional skeletal motion generation technology mainly relies on rule definition or physical simulation. Although it can generate reasonable motion sequences under specific conditions, the diversity and naturalness of its generated results are often limited. In recent years, the introduction of deep learning, especially generative adversarial networks (GANs), has provided a new idea for motion generation and transfer. Generative adversarial networks can achieve high-quality and diverse skeletal motion generation through adversarial learning between generators and discriminators. However, how to achieve accurate decoupled transfer between motion content and style remains an important research direction.
[0003] Existing skeletal motion generation and transfer technologies have the following shortcomings: First, traditional methods mostly rely on specific data distributions and cannot generate diverse motion sequences under limited sample conditions. This limits its application in scenarios with few samples. For example, when rare or special style motion samples need to be generated, the effect is often not ideal. Secondly, in deep learning-based motion transfer methods, the content and style of the action are often difficult to effectively decouple. The generated action sequence may retain the target style characteristics, but often loses the essential content characteristics of the source action, resulting in unstable transfer effects. In addition, in the process of generating action sequences, existing methods are prone to spatial structural inconsistencies (such as abnormal joint position) and temporal discontinuities (such as inter-frame jitter or action breaks), thereby reducing the naturalness and authenticity of the generated samples. Summary of the invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problems solved by the present invention are: the existing skeletal motion generation and migration methods have the problems of strong data dependence and difficulty in generating diversified samples, the problem of difficulty in decoupling motion content and style, the problem of insufficient continuity of generated samples in space and time, and the problem of how to achieve high-quality motion style migration under limited data conditions.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: a method for dynamic transfer of skeletal action style based on a generative adversarial network, comprising extracting skeletal action features; constructing an action style transfer network; and performing multi-constraint optimization based on a loss function.
[0007] As a preferred solution of the method for dynamic transfer of skeletal action style based on a generative adversarial network described in the present invention, wherein: the extraction of skeletal action features includes using a human action data set, the data set includes a three-dimensional coordinate sequence of skeletal joint points of an action category;
[0008] Actions consist of frames, each of which records the three-dimensional coordinates (x, y, z) of the human skeleton joints;
[0009] Classifying data includes source action sequence and target action sequence;
[0010] The source action sequence represents the content information of the action and is used to preserve the essential characteristics of the action;
[0011] The target action sequence represents the style information of the action and is used to transfer style features.
[0012] As a preferred solution of the method for dynamic transfer of skeletal action style based on a generative adversarial network described in the present invention, wherein: the extracted skeletal action features include a unified joint point coordinate range;
[0013] Divide the three-dimensional coordinate values of all bone joint points by the maximum value of all related node coordinates in the sample;
[0014] Map the coordinate value of each joint point to the interval [0, 1];
[0015] Convert the relative coordinates of the bones into absolute coordinates;
[0016] Select the center point of the pelvis as the coordinate origin and translate the coordinates of the joint points, which is expressed as:
[0017] ″′
[0018] X=XX center ,Y=YY center ,Z=ZZ center
[0019] Among them, (X, Y, Z) are the original coordinates, (X center ,Y center ,Z center ) are the coordinates of the bone center point.
[0020] As a preferred solution of the method for dynamic transfer of skeletal action style based on a generative adversarial network described in the present invention, wherein: the construction of the action style transfer network includes designing a generator, inputting a source action sequence {X1, X2, ..., XK}, use the convolutional layer to extract local features, extract global temporal dependency features, and output the content feature vector C of the source action;
[0021] Input target action sequence {Y1,Y2,…,Y K}, combine the convolutional layer and Transformer to extract style features and output the style feature vector S of the target action;
[0022] Input C and S, decode and generate the action sequence after migration;
[0023] Distinguish between generated samples and real samples.
[0024] As a preferred solution of the method for dynamic transfer of skeletal action style based on adversarial generative network described in the present invention, wherein: the multi-constraint optimization based on loss function includes that the generator makes the generated samples close to the real data distribution, the discriminator aims to distinguish the generated samples from the real samples, and in the skeletal action transfer, the adversarial loss ensures that the generated action sequence distribution is consistent with the real action sequence, and the generator adversarial loss and the discriminator adversarial loss are constructed, which are expressed as:
[0025]
[0026] in, represents the adversarial loss of the generator, represents the adversarial loss of the discriminator, represents the generated sample, output by the generator, X represents the real sample, from the data distribution, P dota represents the real data distribution, P G Represents the generated data distribution, D represents the discriminator, and the output sample is the probability of being true;
[0027] In the process of action style transfer, it is necessary to ensure that the generated action sequence conforms to the target style, restore it to the source action sequence, build a cycle consistency loss model, ensure that the generated action samples do not deviate from the source action content, and further optimize the style transfer effect, which is expressed as:
[0028]
[0029] in, represents the cycle consistency loss, C represents the content feature vector of the source action sequence, S represents the style feature vector of the target action sequence, G(C,S) represents the action sequence with the target style output by the generator, Represents the norm, which is used to calculate the Euclidean distance between vectors.
[0030] As a preferred solution of the method for dynamic migration of skeletal action style based on a generative adversarial network described in the present invention, wherein: the multi-constraint optimization based on the loss function includes that the content information in the action sequence needs to remain unchanged during the migration process, and the content features of the generated samples are consistent with the source samples through constraints, and the content preservation loss model is constructed as follows:
[0031]
[0032] in, represents the content preservation loss, C represents the content feature vector of the source action sequence, G(C,S) represents the action sequence with the target style reported by the generator, and F content (·) represents the feature extraction function, which is used to extract the content features of the sample. represents the norm;
[0033] Through the style preservation loss, the style features of the generated samples are constrained to be consistent with the target action sequence. The style features of the samples are extracted using the feature extractor, and the style preservation loss model is constructed as follows:
[0034]
[0035] in, represents the style preservation loss, S represents the style feature vector of the target action sequence, G(C,S) represents the action sequence with the target style output by the generator, and F style (·) represents the feature extraction function, which is used to extract the style features of the sample. Represents the norm, which is used to calculate the distance between the style features of the generated sample and the target sample.
[0036] As a preferred solution of the method for dynamic transfer of skeletal action style based on a generative adversarial network described in the present invention, wherein: the multi-constraint optimization based on the loss function includes the need to maintain spatial consistency and temporal consistency;
[0037] For spatial consistency, the relative positions of the skeletal joints constitute the spatial structure of the action, and the structure needs to remain consistent during the migration process;
[0038] For temporal consistency, the temporal sequence characteristics of the skeletal motion need to be smooth and natural, avoiding jitter or breakage of the generated motion between time frames. A spatial and temporal consistency loss model is constructed, which is expressed as:
[0039]
[0040] in, represents the spatial consistency loss, represents the temporal consistency loss, represents the generated sample, which is output by the generator G(C,S), and X represents the real sample, which comes from the real data distribution P data , represents the Euclidean distance between joint points i and j in the generated sample, d(X i ,X j ) represents the Euclidean distance between joint points i and j in the real sample, t represents the index of the time frame, represents the position change of the generated sample between time frames t and t+1, and represents the amount of action change between frames, X t+1 -X t Represents the position change of the real sample in time frames t and t+1;
[0041] By constraining the smoothness of velocity and acceleration, we can avoid the generation of samples with broken or incoherent motions. The motion flow smoothness loss model is constructed as follows:
[0042]
[0043] in, represents the motion flow smoothness loss, represents the joint point position of the generated sample in time frame t, X t represents the joint point position of the real sample in time frame t, Indicates the speed of generating samples at time t, calculated as the position difference between frames, represents the velocity of the real sample at time t, Represents the acceleration of the generated sample at time t, calculated as the speed difference between frames, represents the acceleration of the real sample at time t;
[0044] Taking into account the constraints of confrontation, content, style, space, time and smoothness, the weight parameters are adjusted according to the task requirements to construct the total loss, which is expressed as:
[0045]
[0046] in, represents the total loss of the generator, λ1,λ2,…,λ7 represent the weight parameters of each part of the loss, which are used to adjust the contribution ratio of different losses to the overall optimization target.
[0047] Another object of the present invention is to provide a skeletal motion style dynamic transfer system based on a generative adversarial network, which can perform multi-constraint optimization through a loss function, thereby solving the problem that current skeletal motion generation and transfer methods cannot achieve high-quality motion style transfer under limited data conditions.
[0048] As a preferred solution of the skeletal motion style dynamic migration system based on the adversarial generative network described in the present invention, it includes: an initialization module, a motion style migration network construction module, and a multi-constraint optimization module; the initialization module is used to collect skeletal motion features; the motion style migration network construction module is used to design a generator and a discriminator; the multi-constraint optimization module is used to output the optimization targets of the generator and the discriminator.
[0049] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a method for dynamic transfer of skeletal action styles based on a generative adversarial network.
[0050] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for dynamic migration of skeletal action styles based on a generative adversarial network.
[0051] Beneficial effects of the present invention: The method for dynamic transfer of skeletal action style based on adversarial generative network provided by the present invention provides high-quality input for the action transfer network by extracting skeletal action features and data structuring, thereby solving the inconsistency problem caused by different data distributions and scale differences. By decoupling action content and style characteristics, the possibility of independent optimization is provided for subsequent network modeling. The action style transfer network is constructed to improve the effect of action style transfer, and the generated action sequence can not only retain the content characteristics of the source action, but also accurately reflect the target style characteristics. At the same time, the generator is continuously optimized by using discriminator feedback to make the generated samples more realistic and natural. Multi-constraint optimization is performed based on the loss function. Through the multiple loss functions of comprehensive constraints, this step comprehensively optimizes the quality of the generated samples from multiple dimensions of content, style, space and time. The generated results are highly authentic and smooth, overcoming the shortcomings of traditional action generation methods in sample authenticity and consistency. The present invention achieves better results in terms of the accuracy of decoupling action content and style, the naturalness and authenticity of generated samples, and the generalization ability under few sample conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0053] Figure 1 An overall flow chart of a method for dynamic transfer of skeletal action style based on a generative adversarial network provided for the first embodiment of the present invention.
[0054] Figure 2 An overall flow chart of a skeletal action style dynamic transfer system based on a generative adversarial network provided for the third embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0056] Example 1, reference Figure 1 , is an embodiment of the present invention, and provides a method for dynamic migration of skeletal action styles based on a generative adversarial network, comprising:
[0057] S1: Extract skeleton motion features.
[0058] Furthermore, data acquisition uses public action datasets (such as the NTU RGB+D dataset) to extract the three-dimensional coordinate information of the skeletal joints in each frame.
[0059] It is divided into source action sequence and target action sequence.
[0060] Source action sequence: contains action content information (such as running, jumping).
[0061] Target action sequence: contains action style information (such as speed and softness).
[0062] Normalize all joint coordinates to the interval [0, 1] to ensure consistent coordinate scales. Convert the relative coordinates of the skeleton to absolute coordinates, using the center point of the skeleton as the origin. Use 70% of the data as the training set and 30% of the data as the test set.
[0063] It should be noted that the center point of the pelvis is selected as the coordinate origin, and the coordinates of the joint points are translated, which is expressed as:
[0064] ″′
[0065] X=XX center ,Y=YY center ,Z=ZZ center
[0066] Among them, (X, Y, Z) are the original coordinates, (X center ,Y center ,Z center ) are the coordinates of the bone center point.
[0067] S2: Build an action style transfer network.
[0068] Furthermore, the generator design includes:
[0069] The generator adopts an autoencoder structure, combined with Transformer and adversarial learning for feature extraction and migration.
[0070] Content encoder input, source action sequence {X1,X2,…,X K}.
[0071] Processing flow:
[0072] Use convolutional layers to extract local features.
[0073] Transformer extracts global temporal dependency features.
[0074] Output the content feature vector C of the source action.
[0075] The style encoder inputs the target action sequence {Y1,Y2,…,Y K}.
[0076] Processing flow:
[0077] Similarly, the convolutional layer and Transformer are combined to extract style features.
[0078] Output the style feature vector S of the target action.
[0079] The decoder inputs are C and S.
[0080] Processing flow:
[0081] The Adaptive Instance Normalization (AdaIN) module combines content features and style features.
[0082] Decoding generates the transferred action sequence
[0083] S3: Multi-constraint optimization based on loss function.
[0084] Furthermore, adversarial loss is the core of generative adversarial networks (GANs). The generator and the discriminator are optimized through a game relationship. The goal of the generator is to deceive the discriminator so that the generated samples are as close to the real data distribution as possible, while the goal of the discriminator is to distinguish the generated samples from the real samples as much as possible.
[0085] In skeletal action transfer, the adversarial loss ensures that the distribution of the generated action sequence is consistent with the real action sequence, improving the authenticity of the generated samples.
[0086] To ensure that the generated action sequences are realistic and difficult for the discriminator to identify as forged samples, and to drive the generator to learn more realistic action features.
[0087] Construct the generator adversarial loss and the discriminator adversarial loss, expressed as:
[0088]
[0089] in, represents the adversarial loss of the generator, represents the adversarial loss of the discriminator, represents the generated sample, output by the generator, X represents the real sample, from the data distribution, P dota represents the real data distribution, P G Represents the generated data distribution, D represents the discriminator, and the output sample is the probability of being true.
[0090] It should be noted that in the process of action style transfer, it is necessary to ensure that the generated action sequence not only conforms to the target style, but also can be restored to the original action sequence. Cycle consistency constraints can prevent the generator from losing key content information of the action. Drawing on the concept of cycle consistency (CycleGAN), the migrated data maintains content consistency. Ensure that the generated action samples do not deviate from the original action content, and at the same time can further optimize the effect of style transfer, expressed as:
[0091]
[0092] in, represents the cycle consistency loss, C represents the content feature vector of the source action sequence, S represents the style feature vector of the target action sequence, G(C,S) represents the action sequence with the target style output by the generator, Represents the norm, which is used to calculate the Euclidean distance between vectors.
[0093] Furthermore, the content information in the action sequence (such as jumping, running, etc.) needs to remain unchanged during the migration process. By constraining the content features of the generated samples to be consistent with the source samples, the destruction of the content features during the generation process can be avoided. The extractor ensures that the content features remain stable. In order to improve the accuracy of the generated samples in terms of content and ensure that the core features of the source action sequence will not be modified incorrectly, a content preservation loss model is constructed as follows:
[0094]
[0095] in, represents the content preservation loss, C represents the content feature vector of the source action sequence, G(C,S) represents the action sequence with the target style reported by the generator, and F content(·) represents the feature extraction function, which is used to extract the content features of the sample. Represents the norm.
[0096] It should be noted that the key to action style transfer is that the generated samples have the target style. Through the style preservation loss, the style features of the generated samples are constrained to be consistent with the target action sequence. The style features of the samples are extracted using a feature extractor to ensure that the style of the generated results is accurate. To ensure that the generated samples can accurately reflect the style characteristics of the target action (such as speed, softness, etc.), the style preservation loss model is constructed as:
[0097]
[0098] in, represents the style preservation loss, S represents the style feature vector of the target action sequence, G(C,S) represents the action sequence with the target style output by the generator, and F style (·) represents the feature extraction function, which is used to extract the style features of the sample. Represents the norm, which is used to calculate the distance between the style features of the generated sample and the target sample.
[0099] Furthermore, spatial consistency and temporal consistency are maintained.
[0100] Spatial consistency is the relative positions of skeletal joints that constitute the spatial structure of the action, which needs to remain consistent during the migration process.
[0101] Temporal consistency refers to the time series characteristics of skeletal motion, which needs to be smooth and natural, avoiding jitter or breakage of the generated motion between time frames.
[0102] Ensure that the generated samples conform to the human body structure in space and remain smooth in time, and build a spatial and temporal consistency loss model, which is expressed as:
[0103]
[0104] in, represents the spatial consistency loss, represents the temporal consistency loss, represents the generated sample, which is output by the generator G(C,S), and X represents the real sample, which comes from the real data distribution P data , represents the Euclidean distance between joint points i and j in the generated sample, d(X i ,X j ) represents the Euclidean distance between joint points i and j in the real sample, t represents the index of the time frame, represents the position change of the generated sample between time frames t and t+1, and represents the amount of action change between frames, X t+1-X t Represents the position change of the real sample in time frames t and t+1.
[0105] It should be noted that the smoothness of the action is an important indicator to measure the quality of the generated samples. By constraining the smoothness of velocity and acceleration, the generated samples can avoid the occurrence of action breaks or incoherence. To make the generated actions smooth and natural, in line with the dynamic characteristics of the real action, by constraining the smoothness of velocity and acceleration, the generated samples can avoid the occurrence of action breaks or incoherence. The action flow smoothness loss model is constructed as:
[0106]
[0107] in, represents the motion flow smoothness loss, represents the joint point position of the generated sample in time frame t, X t represents the joint point position of the real sample in time frame t, Indicates the speed of generating samples at time t, calculated as the position difference between frames, represents the velocity of the real sample at time t, Represents the acceleration of the generated sample at time t, calculated as the speed difference between frames, Represents the acceleration of the real sample at time t.
[0108] Furthermore, by comprehensively considering the constraints of confrontation, content, style, space, time and smoothness, the weight parameters are adjusted according to the task requirements to ensure the comprehensive optimization of the generated samples in terms of authenticity, accuracy and smoothness. The generator is optimized in all aspects, and the generated samples have the target style, maintain content consistency, and conform to the characteristics of human motion in space and time. The total loss is constructed and expressed as:
[0109]
[0110] in, represents the total loss of the generator, λ1,λ2,…,λ7 represent the weight parameters of each part of the loss, which are used to adjust the contribution ratio of different losses to the overall optimization target.
[0111] Example 2, an embodiment of the present invention, provides a method for dynamic migration of skeletal action styles based on a generative adversarial network. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0112] First, we set up the experimental hardware environment. The computing device is equipped with an NVIDIA RTX 3090 GPU accelerator, an Intel i9-11900K CPU, and 64GB of memory. Based on the Python programming language, we use the TensorFlow and PyTorch deep learning frameworks. The data comes from the publicly available NTU RGB+D dataset, which contains rich skeletal motion data (such as running, jumping, waving, etc.), and each frame records the three-dimensional coordinates of the joint points.
[0113] Import the NTU RGB+D dataset, normalize and transform the joint coordinates. Classify the data into source action sequences and target action sequences, and divide them into training sets and test sets. Input the source action sequences and target action sequences of the training set into the generator. The generator outputs the migrated action sequences, which are input into the discriminator together with the real samples. The discriminator calculates the authenticity of the generated samples and feeds them back to the generator for parameter optimization. During the training process, the generator and the discriminator are continuously optimized through adversarial learning. Combined with multiple loss functions, the authenticity, style consistency and dynamic fluency of the generated samples are further improved. After each round of training, the performance of the generated samples on the validation set is recorded, including spatial consistency error, temporal consistency error, etc. Input the test set into the trained generator to generate an action sequence with the target style. Calculate the spatial consistency error, temporal consistency error, style preservation accuracy, content preservation accuracy and other indicators of the generated samples. Score the fluency of the dynamic performance of the generated actions.
[0114] Table 1 Experimental data table
[0115]
[0116]
[0117] The following trends can be observed from the experimental data:
[0118] Average spatial consistency error: The spatial consistency errors of all test subjects were within a reasonable range (<0.04 meters), with the minimum value being squat-stable (0.015 meters) and the maximum value being turn-soft (0.035 meters). This shows that by constraining the spatial consistency loss, the generated skeletal motions are basically consistent with the human body structure in terms of the relative positions of the joints.
[0119] Temporal consistency error: The temporal consistency error is also low, with the minimum value of squatting-stable (0.010 seconds) and the maximum value of turning-soft (0.025 seconds). This shows that the inter-frame motion of the generated samples has good continuity and avoids obvious motion breaks.
[0120] Style-preserving accuracy: The accuracy of different action styles ranges from 94.0% to 97.5%, especially squat-steady (97.5%) performs well, indicating that the style-preserving loss effectively transfers the style characteristics of the target action.
[0121] Content Preservation Accuracy: The content preservation accuracy was stable between 95.5% and 97.8%, proving that the source action content information was fully preserved, with squat-steady (97.8%) performing best.
[0122] Generated sample fluency score: The scores are concentrated between 9.2 and 9.7, and the generated samples are highly natural and smooth in dynamic movement performance.
[0123] Adversarial discrimination authenticity: The authenticity of all test objects exceeded 93%, indicating that the optimization of the generative adversarial network (GAN) effectively improved the authenticity of the generated samples, with the squat-steady authenticity reaching the highest (97%).
[0124] Compared with traditional action generation and migration methods, the present invention has significant advantages in the following aspects:
[0125] Spatial consistency optimization: Traditional methods do not fully constrain the relative positions of joints, and the generated actions are prone to abnormal bone positions. This invention significantly reduces the spatial error (the minimum value is only 0.015 meters) through spatial consistency loss.
[0126] Temporal consistency optimization: Traditional methods often lead to discontinuity of motion between frames. This invention introduces temporal consistency and motion flow smoothness loss, which significantly reduces the temporal consistency error and generates more natural dynamic performance of motion (the minimum error is 0.010 seconds).
[0127] Decoupling of style and content: Traditional methods find it difficult to preserve action content features while transferring style. This paper successfully achieves precise decoupling of content and style through independently designed content preservation loss and style preservation loss, with accuracies of 97.8% and 97.5% respectively.
[0128] Fluency of generated samples: Traditional methods lack optimization of dynamic continuity, and the generated actions may appear jerky or uncoordinated. By integrating multiple loss constraints, the generated action sequences of this invention are significantly ahead in fluency scores, with scores concentrated between 9.2-9.7.
[0129] It can be clearly seen from the experimental data in this embodiment that the present invention has significantly improved the authenticity, style accuracy, content retention and dynamic fluency of action generation compared with the traditional method, verifying the feasibility and superiority of the present invention in practical applications. These advantages make it widely applicable to animation production, behavior analysis, virtual reality and medical rehabilitation.
[0130] Example 3, reference Figure 2 , as an embodiment of the present invention, provides a skeletal action style dynamic migration system based on a generative adversarial network, including an initialization module, an action style migration network construction module, and a multi-constraint optimization module.
[0131] The initialization module is used to collect skeletal motion features, the motion style transfer network construction module is used to design the generator and discriminator, and the multi-constraint optimization module is used to output the optimization objectives of the generator and discriminator.
[0132] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0134] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0135] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc. It should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and are not limited. Although the present invention is described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.
[0136] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for dynamic transfer of skeletal action style based on a generative adversarial network, characterized in that: include: Extract skeleton motion features; Build an action style transfer network; Multi-constraint optimization based on loss function; The generator makes the generated samples close to the real data distribution. The goal of the discriminator is to distinguish between generated samples and real samples. In the skeletal action transfer, the adversarial loss ensures that the generated action sequence distribution is consistent with the real action sequence. The generator adversarial loss and the discriminator adversarial loss are constructed and expressed as: in, represents the adversarial loss of the generator, represents the adversarial loss of the discriminator, represents the generated sample, output by the generator, X represents the real sample, from the data distribution, P dota represents the real data distribution, P G Represents the generated data distribution, D represents the discriminator, and the output sample is the probability of being true; In the process of action style transfer, it is necessary to ensure that the generated action sequence conforms to the target style, restore it to the source action sequence, build a cycle consistency loss model, ensure that the generated action samples do not deviate from the source action content, and further optimize the style transfer effect, which is expressed as: in, represents the cycle consistency loss, C represents the content feature vector of the source action sequence, S represents the style feature vector of the target action sequence, G(C,S) represents the action sequence with the target style output by the generator, :L2 represents the norm, which is used to calculate the Euclidean distance between vectors; The content information in the action sequence needs to remain unchanged during the migration process. By constraining the content features of the generated samples to be consistent with the source samples, the content preservation loss model is constructed as follows: in, represents the content preservation loss, C represents the content feature vector of the source action sequence, G(C,S) represents the action sequence with the target style reported by the generator, and F content (·) represents the feature extraction function, which is used to extract the content features of the sample. :L2 represents the norm; Through the style preservation loss, the style features of the generated samples are constrained to be consistent with the target action sequence. The style features of the samples are extracted using the feature extractor, and the style preservation loss model is constructed as follows: in, represents the style preservation loss, S represents the style feature vector of the target action sequence, G(C,S) represents the action sequence with the target style output by the generator, and F style (·) represents the feature extraction function, which is used to extract the style features of the sample. : L2 represents the norm, which is used to calculate the distance between the style features of the generated sample and the target sample; The need to maintain spatial consistency and temporal consistency; For spatial consistency, the relative positions of the skeletal joints constitute the spatial structure of the action, and the structure needs to remain consistent during the migration process; For temporal consistency, the temporal sequence characteristics of the skeletal motion need to be smooth and natural, avoiding jitter or breakage of the generated motion between time frames. A spatial and temporal consistency loss model is constructed, which is expressed as: in, represents the spatial consistency loss, represents the temporal consistency loss, represents the generated sample, which is output by the generator G(C,S), and X represents the real sample, which comes from the real data distribution P data , represents the Euclidean distance between joint points i and j in the generated sample, d(X i ,X j ) represents the Euclidean distance between joint points i and j in the real sample, t represents the index of the time frame, represents the position change of the generated sample between time frames t and t+1, and represents the amount of action change between frames, X t+1 -X t Represents the position change of the real sample in time frames t and t+1; By constraining the smoothness of velocity and acceleration, we can avoid the generation of samples with broken or incoherent motions. The motion flow smoothness loss model is constructed as follows: in, represents the motion flow smoothness loss, represents the joint point position of the generated sample in time frame t, X t represents the joint point position of the real sample in time frame t, Indicates the speed of generating samples at time t, calculated as the position difference between frames, represents the velocity of the real sample at time t, Represents the acceleration of the generated sample at time t, calculated as the speed difference between frames, represents the acceleration of the real sample at time t; Taking into account the constraints of confrontation, content, style, space, time and smoothness, the weight parameters are adjusted according to the task requirements to construct the total loss, which is expressed as: in, represents the total loss of the generator, λ1,λ2,…,λ7 represent the weight parameters of each part of the loss, which are used to adjust the contribution ratio of different losses to the overall optimization target.
2. The method for dynamic transfer of skeletal action style based on a generative adversarial network as claimed in claim 1, characterized in that: Extracting skeletal motion features includes using a human motion data set, the data set comprising a three-dimensional coordinate sequence of skeletal joint points of the motion category; Actions consist of frames, each of which records the three-dimensional coordinates (x, y, z) of the human skeleton joints; Classifying data includes source action sequence and target action sequence; The source action sequence represents the content information of the action and is used to preserve the essential characteristics of the action; The target action sequence represents the style information of the action and is used to transfer style features.
3. The method for dynamic transfer of skeletal action style based on a generative adversarial network as claimed in claim 2, characterized in that: The extracting of skeletal motion features includes unifying the range of joint point coordinates; Divide the three-dimensional coordinate values of all bone joint points by the maximum value of all related node coordinates in the sample; Map the coordinate value of each joint point to the interval [0, 1]; Convert the relative coordinates of the bones into absolute coordinates; Select the center point of the pelvis as the coordinate origin and translate the coordinates of the joint points, which is expressed as: ″′ X=X-X center ,Y=Y-Y center ,Z=Z-Z center Among them, (X, Y, Z) are the original coordinates, (X center ,Y center ,Z center ) are the coordinates of the bone center point.
4. The method for dynamic transfer of skeletal action style based on a generative adversarial network as claimed in claim 3, characterized in that: The construction of the action style transfer network includes designing a generator, inputting a source action sequence X1, X2, ..., X K }, use the convolutional layer to extract local features, extract global temporal dependency features, and output the content feature vector C of the source action; Input target action sequence {Y1,Y2,…,Y K }, combine the convolutional layer and Transformer to extract style features and output the style feature vector S of the target action; Input C and S, decode and generate the action sequence after migration; Distinguish between generated samples and real samples.
5. A system using the method for dynamic transfer of skeletal action style based on a generative adversarial network as claimed in any one of claims 1 to 4, characterized in that: Including initialization module, action style transfer network construction module, multi-constraint optimization module; The initialization module is used to collect skeletal motion features; The action style transfer network building module is used to design the generator and the discriminator; The multi-constraint optimization module is used to output optimization objectives of the generator and the discriminator.
6. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method for dynamic transfer of skeletal action style based on a generative adversarial network according to any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for dynamic transfer of skeletal action style based on a generative adversarial network according to any one of claims 1 to 4 are implemented.
Citation Information
Cited By
Athlete action analysis and evaluation method and system based on computer vision
CN121354222A
Intelligent interactive system for generating and interacting with dynamic content of intangible cultural heritage based on generative adversarial network
CN122530396A