Action generation method and device, equipment and storage medium
Through the combination of codec and diffusion models, efficient migration of action style is achieved, and the problem of equipment and personnel dependence in the process of action stylization in the prior art is solved, which improves generation efficiency and reduces costs.
Patent Information
- Application Number
- CN202510134120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
AI Technical Summary
When the prior art realizes movement stylization, professional equipment, venues and personnel are required, resulting in low generation efficiency and high migration costs.
By inputting the original action data into the pre-trained codec model, the action coded features are extracted, and the action style description information is combined into the diffusion model, and the action style description is used to perform denoising processing to achieve action style migration. Finally, the target action coded features are input to the decoder to generate the target action data.
Without changing the action content, efficient migration of action style is achieved, reducing dependence on professional equipment, venues and personnel, improving the generation efficiency of stylized action data and reducing migration costs.
Smart Images

Figure CN120070684A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology. Specifically, it relates to a method, device, equipment and storage medium for action generation. Background Art
[0002] In fields such as games, animations, and films, it is often necessary to convert an animation into an animation of other styles, that is, action stylization. Among them, action stylization requires that the action after stylization only changes the action style information (such as feminization, masculinization, old age, infancy, etc.), while keeping the action content information (such as walking, running, jumping, etc.) unchanged. That is, the goal of the action stylization task is to keep the action content of the specified reference action unchanged and change the action style of the reference action to the specified target action style.
[0003] Currently, the commonly used action stylization methods include motion capture and art production. Among them, the motion capture method requires a motion capture actor to perform the original reference action according to the specified target action style, and use professional motion capture equipment to collect the motion data of the motion capture actor in a professional motion capture scene to obtain the reference action data of the target action style. The art production method is to let professional art personnel manually modify the action style of the reference action in the animation production software to obtain the reference action data of the target action style. However, the above two solutions involve investments in professional equipment, venues, and personnel, and the process is cumbersome. Therefore, the generation efficiency of the reference action data of the target action style is low and the action style migration cost of the reference action data is high. Summary of the Invention
[0004] In view of this, the present application provides a method, device, equipment and storage medium for action generation, which can support the migration of action styles without changing the action content of the original action data, and does not require the assistance of professional venues, equipment, and personnel, which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data.
[0005] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and are described in detail as follows.
[0006] In a first aspect, an embodiment of the present application provides a method for action generation, and the method for action generation includes:
[0007] Input the original action data into the encoder of the pre-trained encoding and decoding model to obtain the action encoding features output by the encoder for the original action data;
[0008] Input the action style description information into the pre-trained target encoder to obtain the style description features output by the target encoder for the action style description information;
[0009] Input the action encoding feature and the style description feature into a pre-trained diffusion model. Use the style description feature as the guiding control condition for the diffusion model during denoising processing, and output a target action encoding feature that matches the action encoding feature and the style description feature through the diffusion model.
[0010] Input the target action encoding feature into the decoder of the codec model to obtain target action data output by the decoder for the target action encoding feature; wherein, the action content of the target action data matches the action content of the original action data, and the action style of the target action data matches the action style description information.
[0011] In a second aspect, an embodiment of the present application provides an action generation device, which includes:
[0012] A first encoding module for inputting original action data into the encoder of a pre-trained codec model to obtain an action encoding feature output by the encoder for the original action data;
[0013] A second encoding module for inputting action style description information into a pre-trained target encoder to obtain a style description feature output by the target encoder for the action style description information;
[0014] A prediction module for inputting the action encoding feature and the style description feature into a pre-trained diffusion model. Use the style description feature as the guiding control condition for the diffusion model during denoising processing, and output a target action encoding feature that matches the action encoding feature and the style description feature through the diffusion model.
[0015] A decoding module for inputting the target action encoding feature into the decoder of the codec model to obtain target action data output by the decoder for the target action encoding feature; wherein, the action content of the target action data matches the action content of the original action data, and the action style of the target action data matches the action style description information.
[0016] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above action generation method are implemented.
[0017] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the above-mentioned action generation method.
[0018] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0019] An action generation method, device, equipment and storage medium provided by an embodiment of the present application input original action data into an encoder of a pre-trained encoding and decoding model to obtain action encoding features output by the encoder for the original action data; input action style description information into a pre-trained target encoder to obtain style description features output by the target encoder for the action style description information; input the action encoding features and the style description features into a pre-trained diffusion model, use the style description features as the guiding control condition during denoising processing of the diffusion model, and output target action encoding features that match the action encoding features and the style description features through the diffusion model; input the target action encoding features into a decoder of the encoding and decoding model to obtain target action data output by the decoder for the target action encoding features. In this way, the present application can support the migration of action styles without changing the action content of the original action data, and does not require professional venues, equipment, personnel, etc., which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 It shows a schematic flowchart of an action generation method provided by an embodiment of the present application;
[0022] Figure 2 It shows a schematic flowchart of a process for generating target action data provided by an embodiment of the present application;
[0023] Figure 3 It shows a schematic flowchart of a model training method for an encoding and decoding model provided by an embodiment of the present application;
[0024] Figure 4 It shows a schematic flowchart of a model training method for a diffusion model provided by an embodiment of the present application;
[0025] Figure 5The figure shows a schematic flowchart of a training method for an action style extraction network provided by an embodiment of the present application;
[0026] Figure 6 The figure shows a schematic structural diagram of an action generation device provided by an embodiment of the present application;
[0027] Figure 7 It is a schematic structural diagram of an electronic device 700 provided by an embodiment of the present application. Detailed implementation manners
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application illustrate operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0029] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the protection scope of the present application.
[0030] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.
[0031] Currently, the commonly used action stylization methods include motion capture and art production. Among them, the motion capture method requires a motion capture actor to perform the original reference action according to the specified target action style, and use professional motion capture equipment to collect the actions of the motion capture actor in a professional motion capture scene, obtaining reference action data of the target action style. The art production method is to let professional artists manually modify the action style of the reference action in animation production software to obtain reference action data of the target action style. However, the above two solutions involve inputs such as professional equipment, venues, and personnel, and the process is cumbersome. Therefore, the generation efficiency of the reference action data of the target action style is low, and the action style migration cost of the reference action data is high.
[0032] Based on this, the embodiments of the present application provide an action generation method, device, equipment, and storage medium, which can support the migration of the action style without changing the action content of the original action data, and do not require the assistance of professional venues, equipment, and personnel, which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data.
[0033] In one embodiment of the present application, an action generation method can run on a terminal device or a server. Among them, the terminal device can be a local terminal device. When the action generation method runs on the server, the action generation method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device (i.e., the terminal device).
[0034] For the convenience of understanding the embodiments of the present application, a detailed introduction will be given below to an action generation method, device, equipment, and storage medium provided by the embodiments of the present application.
[0035] Refer to Figure 1 as shown Figure 1 The flowchart of an action generation method provided by the embodiments of the present application is shown, where the action generation method includes steps S101 - S104; specifically:
[0036] S101, input the original action data into the encoder of the pre-trained codec model, and obtain the action encoding features output by the encoder for the original action data.
[0037] S102, input the action style description information into the pre-trained target encoder, and obtain the style description features output by the target encoder for the action style description information.
[0038] S103. Input the action encoding feature and the style description feature into a pre-trained diffusion model. Use the style description feature as the guiding control condition for the diffusion model during denoising processing, and output, through the diffusion model, a target action encoding feature that matches the action encoding feature and the style description feature.
[0039] S104. Input the target action encoding feature into the decoder of the encoder-decoder model to obtain target action data output by the decoder for the target action encoding feature.
[0040] For the above action generation method provided by the embodiments of the present application, the original action data is input into the encoder of a pre-trained encoder-decoder model to obtain an action encoding feature output by the encoder for the original action data; the action style description information is input into a pre-trained target encoder to obtain a style description feature output by the target encoder for the action style description information; the action encoding feature and the style description feature are input into a pre-trained diffusion model. Using the style description feature as the guiding control condition for the diffusion model during denoising processing, a target action encoding feature that matches the action encoding feature and the style description feature is output through the diffusion model; the target action encoding feature is input into the decoder of the encoder-decoder model to obtain target action data output by the decoder for the target action encoding feature. In this way, the present application can support the migration of the action style of the original action data without changing the action content, and does not require professional venues, equipment, personnel, etc., which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data.
[0041] The following is an exemplary description of each step in the above action generation method provided by the embodiments of the present application:
[0042] S101. Input the original action data into the encoder of a pre-trained encoder-decoder model to obtain an action encoding feature output by the encoder for the original action data.
[0043] Here, the original action data can be an action image sequence composed of multiple frames of action images of an action object. Among them, the action object can be a real person, or a virtual character in a game or a film and television work, etc. The action of the action object in each frame of the action image can be actions such as "walking", "running", "jumping", etc. The present application embodiments do not make any limitations on the specific object type to which the above action object belongs and the specific action content of the action object in each frame of the action image (that is, the action content of the above original action data).
[0044] Here, the encoding and decoding model consists of an encoder and a decoder; among them, in the encoding and decoding model, the encoder is used to encode the original action data input into the encoder, map the original action data into the latent space of the encoding and decoding model, and obtain the feature representation of the original action data in the latent space as the action encoding feature of the original action data.
[0045] Specifically, through the above-mentioned pre-trained encoding and decoding model, the above-mentioned original action data with a higher feature dimension can be mapped into a more compact feature space (i.e., the above-mentioned latent space), so that the encoder in the encoding and decoding model can output an action encoding feature with a lower feature dimension (i.e., lower than the feature dimension of the original feature data), which is beneficial to reducing the feature dimension and saving computing resources.
[0046] It should be noted that the above-mentioned encoding and decoding model can be a MAE (Masked Autoencoders, a computer vision model based on self-supervised learning) model, or other models with an encoding and decoding structure. That is, in the embodiments of the present application, it is only necessary to ensure that the above-mentioned encoding and decoding model includes at least one set of encoder and decoder. For the specific model structure of the above-mentioned encoding and decoding model, the embodiments of the present application do not make any limitations.
[0047] S102, input the action style description information into the pre-trained target encoder to obtain the style description feature output by the target encoder for the action style description information.
[0048] Here, the action style description information is used to represent the relevant description information that can reflect the target action style; among them, the target action style is the action style to which the above-mentioned original action data specified by the user needs to be migrated. The user can choose to use text-type action style description information to reflect the above-mentioned target action style, or choose to use image-type action style description information to reflect the above-mentioned target action style, or choose to use voice-type action style description information to reflect the above-mentioned target action style. For the specific information type of the above-mentioned action style description information, the embodiments of the present application do not make any limitations.
[0049] Specifically, when choosing to use text-type action style description information to reflect the above-mentioned target action style, one or more keywords matching the above-mentioned target action style can be determined from various different types of style description keywords to form the above-mentioned text-type action style description information.
[0050] Exemplary illustration. Taking the style description keywords included in the gender style as "male" and "female", and the style description keywords included in the age style as "juvenile", "adolescent", "adult", and "elderly" as examples, if the target action style is a character action style applicable to an adult male virtual character, then "male gender and adult age" can be used as the action style description information of the text type.
[0051] Specifically, when choosing to use the action style description information of the image type to reflect the above target action style, an action image sequence composed of multiple frames of action images whose action styles match the above target action style can be obtained as the action style description information of the above image type; among them, since the action style description information is only used to reflect the action style that the above original action data needs to migrate, and does not involve changing the specific action content in the above original action data, therefore, in the action style description information of the image type, the action content can be the same as the above original action data (for example, the original action data is an action image sequence of a virtual character of a young girl walking, and the action style description information can be an action image sequence of a virtual character of an adult male walking), or different from the above original action data (for example, the original action data is an action image sequence of a virtual character of a young girl walking, and the action style description information can be an action image sequence of a virtual character of an adult male jumping).
[0052] It should be noted that since the information type of the above action style description information is not unique, therefore, in the embodiments of the present application, the specific encoder type of the above target encoder can be determined according to the information type of the above action style description information selected in actual applications. For the specific encoder type of the above target encoder, the embodiments of the present application also do not make any limitations.
[0053] Exemplary illustration. In actual applications, if the action style description information of the text type is selected, a pre-trained text encoder can be selected as the above target encoder (for example, a pre-trained CLIP text encoder can be selected), and the action style description information of the text type is input into the text encoder to obtain the text features output by the text encoder as the above style description features.
[0054] Exemplary illustration. In actual applications, if the action style description information of the image type is selected, a pre-trained image encoder can be selected as the above target encoder, and the action style description information of the image type is input into the image encoder to obtain the image features output by the image encoder as the above style description features.
[0055] Exemplary illustration, in practical applications, if it is selected to use the action style description information of the voice type, the action style description information of the voice type can be first converted into the action style description information of the text type, and then the converted action style description information of the text type is input into the pre-trained text encoder (i.e., the target encoder), and the text features output by the text encoder are obtained as the above-mentioned style description features.
[0056] S103, input the action encoding feature and the style description feature into the pre-trained diffusion model, use the style description feature as the guiding control condition for the diffusion model during denoising processing, and output the target action encoding feature that matches the action encoding feature and the style description feature through the diffusion model.
[0057] Here, there are two processes, a forward process and a reverse process, inside the diffusion model. Among them, in the forward process, the diffusion model performs noise addition processing on the input action encoding feature (for example, adding T steps of noise to the input action encoding feature, where the specific value of T can be set and adjusted according to actual needs), and obtains the noise-added action encoding feature.
[0058] In the reverse process, using the above-mentioned style description feature as the guiding control condition, the diffusion model (mainly the denoising module in the diffusion model) can predict the denoising result of the above-mentioned noise-added action encoding feature (that is, predict the action encoding feature without noise under the condition that the above-mentioned style description feature is used as the guiding control condition), and output the predicted denoising result as the above-mentioned target action encoding feature.
[0059] S104, input the target action encoding feature into the decoder of the encoder-decoder model, and obtain the target action data output by the decoder for the target action encoding feature.
[0060] Here, referring to the relevant description content in the above-mentioned step S101, it can be known that the encoder-decoder model is composed of an encoder and a decoder; among them, in the encoder-decoder model, corresponding to the encoder in step S101, in step S104, the decoder is used to perform decoding processing on the above-mentioned target action encoding feature input into the decoder, map the above-mentioned target action encoding feature from the latent space back to the data space corresponding to the above-mentioned original action data, and obtain the action data corresponding to the above-mentioned target action encoding feature in this data space, which is the above-mentioned target action data.
[0061] It should be noted that since the above-mentioned target action encoding features are obtained by the diffusion model denoising the above-mentioned noise-added action encoding features on the basis of the above-mentioned action encoding features and relying on the above-mentioned style description features as the guiding control conditions, therefore, in the target action data obtained by restoring the above-mentioned target action encoding features through the decoder, the action content of the target action data matches the action content of the original action data (equivalent to the action content being determined by the above-mentioned action encoding features that are the data basis for the diffusion model processing), and the action style of the target action data matches the action style description information (equivalent to the action style being affected by the above-mentioned style description features relied on by the diffusion model during the denoising process).
[0062] Regarding the above steps S101 - S104, it should be noted that in the embodiments of the present application, the overall model architecture composed of the above-mentioned encoding and decoding model, the above-mentioned target encoder, and the above-mentioned diffusion model can be used as a pre-trained action style conversion model; wherein, the above-mentioned original action data and the above-mentioned action style description information are the source data input by the user to the above-mentioned action style conversion model (which is also equivalent to the model input data of the above-mentioned action style conversion model). Inside the model of the above-mentioned action style conversion model, the action style conversion model is used to convert the action style of the above-mentioned original action data from the original action style (i.e., the original action style of the original action data) to the target action style matching the above-mentioned action style description information under the condition of keeping the action content of the above-mentioned original action data unchanged. Furthermore, through the above-mentioned action style conversion model, the above-mentioned target action data whose action content matches the original action data and whose action style matches the action style description information can be output (i.e., the above-mentioned target action data is the model output result of the above-mentioned action style conversion model).
[0063] Specifically, for the specific implementation process of the above steps S101 - S104, taking the above-mentioned action style description information as the action style description information of the image type as an example, Figure 2 shows a schematic flow diagram of generating target action data provided by the embodiments of the present application, as Figure 2 shown, the encoding and decoding model consists of an encoder and a decoder; wherein, from the source data input by the user, the original action data is obtained, and the obtained original action data is input into the encoder. The encoder can output the action encoding features of the original action data and input the output action encoding features into the diffusion model.
[0064] As Figure 2 shown, from the source data input by the user, the action style description information is obtained, and the obtained action style description information is input into the target encoder (which can be an image encoder at this time). The target encoder can output the style description features of the action style description information and input the output style description features into the diffusion model.
[0065] As Figure 2 shown, according to the above-mentioned action encoding features and the above-mentioned style description features input, the diffusion model can output target action encoding features that match the above-mentioned action encoding features and the above-mentioned style description features, and input the output above-mentioned target action encoding features into the decoder of the encoder-decoder model. By decoding the input above-mentioned target action encoding features, the decoder can output target action data whose action content matches the original action data (refer to Figure 2 shown, the movement amplitude of the legs of the action object in the target action data and the distance between the two feet are basically consistent with the original action data), and the action style matches the action style description information; thus enabling the present application to support the migration of the action style without changing the action content of the original action data (refer to Figure 2 shown, the action style of the action object in the target action data is migrated from the action style in the previous original action data to the action style of the action style description information), and it does not require the aid of professional venues, equipment, personnel, etc., which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data.
[0066] Next, the specific implementation processes of the above steps in the embodiments of the present application will be described in detail respectively:
[0067] Regarding the encoder-decoder model in the above steps S101-S104, in an alternative embodiment, Figure 3 shows a schematic flowchart of a model training method for an encoder-decoder model provided by an embodiment of the present application. As Figure 3 shown, the model training method includes steps S301-S303; specifically:
[0068] S301, input a plurality of first sample action data into the encoder, and output the first action encoding features of each first sample action data through the encoder.
[0069] Here, as an alternative embodiment, similar to the original action data in step S101, an action image sequence of a plurality of different action contents (such as different action contents such as running, jumping, walking, etc.) of one or more action objects can be collected as the above-mentioned plurality of first sample action data.
[0070] Here, as another alternative embodiment, for the convenience of model training of the codec model and to make the codec model converge more easily, after collecting the action image sequences of the above-mentioned multiple different action contents, the action image sequences of the above-mentioned multiple different action contents collected can be preprocessed according to the method shown in the following steps a1-a2, and the preprocessed action image sequences are used as the above-mentioned first sample action data in the input encoder. Specifically:
[0071] Step a1: For each collected action image sequence, extract action features in a specific data feature format suitable for model training of the codec model from the action image sequence.
[0072] Exemplarily, action features in the commonly used invariant features format can be adopted. This feature has a total of 626 dimensions. Among them, the following features can be mainly extracted to represent the action features of the action object in the action image sequence:
[0073] Root node features: the vertical height of the root node of the skeleton of the action object, the linear velocity in the x-z plane, the angular velocity of rotation around the y-axis, and 6D rotation features, etc.;
[0074] Body joint features: 6D rotation of the body joints of the action object, local position coordinates (i.e., the position coordinates of each joint in the coordinate system of the joint itself), local linear velocity, etc.;
[0075] Double-foot contact label: a label indicating whether the two feet of the action object are in contact with the ground. Among them, if the two feet of the action object are in contact with the ground, the double-foot contact label can be set to 1, otherwise it can be set to 0.
[0076] Step a2: To enrich the number of training data, data augmentation processing (such as mirror processing of the action image sequence, etc.) can also be performed on each collected action image sequence. Extract action features in a specific data feature format suitable for model training of the codec model from the action image sequence after data augmentation processing, and use all the action features extracted in steps a1 and a2 as the above-mentioned multiple first sample action data to be input into the encoder.
[0077] Here, the specific encoding method of the encoder for the input first sample action data in step S301 can refer to the encoding process of the encoder for the original action data in the foregoing step S101, and the repeated parts will not be elaborated here.
[0078] S302. For each first sample action data, input the first action encoding feature of the first sample action data into the decoder, and output the restored action data of the first action encoding feature through the decoder.
[0079] Here, the specific decoding method performed by the decoder on the input first action encoding feature in step S302 can refer to the decoding process of the decoder on the target action encoding feature in the foregoing step S104. Repetitive parts will not be elaborated here.
[0080] S303. Adjust the model parameters of the encoding and decoding model according to the loss between the first sample action data and the restored action data, and obtain the encoding and decoding model including the adjusted model parameters.
[0081] Here, since the model training objective of the encoding and decoding model is mainly to train the decoder to output the restored action data that can restore the first sample action data before encoding as much as possible according to the first action encoding feature output by the encoder, during the model training process of the encoding and decoding model, the training loss of the model belongs to the reconstruction loss. The L2 loss function can be preferentially used to calculate the loss between the first sample action data and the restored action data. That is, the mean square error between the first sample action data and the restored action data can be calculated as the loss between the first sample action data and the restored action data.
[0082] It should be noted that in addition to the L2 loss function, other types of loss functions can also be used to calculate the loss between the first sample action data and the restored action data. The specific type of loss function actually used is not limited in the embodiments of the present application.
[0083] Specifically, after calculating the loss between each group of first sample action data and the restored action data, the sum of the losses between each group of first sample action data and the restored action data can be calculated as the overall model loss of the encoding and decoding model, and the encoding and decoding model can be trained until the encoding and decoding model converges (equivalent to continuously adjusting the model parameters of the encoding and decoding model until the above overall model loss reaches the minimum), and the converged encoding and decoding model is obtained as the trained encoding and decoding model (that is, the encoding and decoding model including the adjusted model parameters is obtained).
[0084] For the target encoder in the above step S102, similar to the model training method of the above encoding and decoding model, multiple initial action style description information can also be collected as the training data set. Each initial action style description information is input into the target encoder, and the initial style description features of each initial action style description information are output. Then, the initial style description features of each initial action style description information are input into the target decoder, and the restored action style description information of each initial style description feature is output. Then, according to the loss between the restored action style description information of each initial style description feature and the initial action style description information, the target encoder and the target decoder are trained until the target encoder and the target decoder converge. The repetitive parts are not elaborated here.
[0085] For the diffusion model in the above steps S103 - S104, on the basis of the above trained encoding and decoding model, in an optional implementation manner, Figure 4 shows a schematic flow chart of a model training method for a diffusion model provided by an embodiment of the present application, as Figure 4 shown, the model training method includes steps S401 - S405; specifically:
[0086] S401, input multiple second sample action data into the encoder, and output the second action encoding features of each second sample action data through the encoder.
[0087] Here, the data acquisition method of the second sample action data in step S401 can refer to the data acquisition method of the first sample action data in the foregoing step S301, and the specific execution manner of step S401 can refer to the specific execution manner of the foregoing step S301, except that the encoder in step S401 is the already trained encoder. The repetitive parts are not elaborated here.
[0088] S402, input the sample action style description information into the target encoder, and obtain the sample style description features output by the target encoder for the sample action style description information.
[0089] Here, the specific information format of the sample action style description information can refer to the relevant description content of the action style description information in the foregoing step S102, and the specific execution manner of step S402 can refer to the specific execution manner of the foregoing step S102. The repetitive parts are not elaborated here.
[0090] S403. Input the second action coding feature and the sample style description feature into the diffusion model. Use the sample style description feature as the guiding control condition for the diffusion model during denoising processing, and output a sample action coding feature that matches the second action coding feature and the sample style description feature through the diffusion model.
[0091] Here, the specific execution method of step S403 can refer to the specific execution method of the foregoing step S103, and the repeated parts will not be elaborated here.
[0092] Specifically, in the above steps S401 - S403, the above sample action coding feature can be output according to the method shown in the following steps b1 - b4:
[0093] Step b1. For each second sample action data m 0 , input the second sample action data m 0 into the trained above-mentioned encoder. After being encoded by the above-mentioned encoder, a second action coding feature f 0 is obtained. Input the second action coding feature f 0 into the diffusion model. Inside the diffusion model, after T steps of adding noise in a forward process, an action sequence feature f T with added noise is obtained.
[0094] Step b2. For the above sample action style description information, input the sample action style description information into the trained target encoder. After being encoded by the above-mentioned target encoder, a sample style description feature f text is obtained.
[0095] Step b3. Inside the diffusion model, different noise levels can be set according to different time steps during the model training process. Therefore, the time step is also an important conditional input. In an optional embodiment of the present application, the time step t can pass through an MLP layer to obtain a time step feature vector f mlp(t) .
[0096] Step b4. Inside the diffusion model, the denoising module θ is composed of a Transformer Encoder (i.e., Transformer encoder) (including 12 Transformer layers). Among them, the input of the denoising module θ is the concatenation of the above second action coding feature f 0 , the above sample style description feature f text and the above time step feature vector f mlp(t) . Therefore, the sample action coding feature f' 0 output by the denoising module θ (i.e., with the above sample style description feature f textAs a guiding control condition, the denoising result of the second action encoded feature after adding noise is predicted through the denoising module θ, and the predicted denoising result) output can be expressed as: f′ 0 = θ(f T , f mlp(t) , f text ).
[0097] S404. Input the sample action encoded feature into the decoder to obtain the target sample action data output by the decoder for the sample action encoded feature.
[0098] Here, the specific execution manner of step S404 can refer to the specific execution manner of the foregoing step S104, and the repeated parts will not be elaborated here.
[0099] S405. Adjust the model parameters of the diffusion model according to the loss between the second action encoded feature and the sample action encoded feature to obtain the diffusion model including the adjusted model parameters.
[0100] Here, the L2 loss function can also be preferentially used to calculate the loss between the second action encoded feature and the sample action encoded feature, that is, the mean square error between the second action encoded feature and the sample action encoded feature can be calculated as the loss between the second action encoded feature and the sample action encoded feature.
[0101] It should be noted that in addition to the L2 loss function, other types of loss functions can also be used to calculate the loss between the second action encoded feature and the sample action encoded feature. The specific type of loss function actually used is not limited in the embodiments of the present application.
[0102] Specifically, after calculating the loss between each group of second action encoded features and sample action encoded features, the sum of the losses between each group of second action encoded features and sample action encoded features can be calculated as the overall model loss of the diffusion model, and the diffusion model is trained until the diffusion model converges (equivalent to continuously adjusting the model parameters of the diffusion model until the above overall model loss reaches the minimum), and the converged diffusion model is obtained as the trained diffusion model (that is, the diffusion model including the adjusted model parameters) is obtained.
[0103] Here, referring to the model training method described in the above steps S401 - S405, it can be known that when training the diffusion model according to the model training method described in the above steps S401 - S405, based on the loss between the second action encoding feature and the sample action encoding feature, the diffusion model can be trained to learn how to improve the matching degree of the sample action encoding feature output by the model and the second action encoding feature output by the encoder on the action content side, which is beneficial to improving the consistency of the finally generated target action data and the original action data before action style transfer in terms of action content.
[0104] On this basis, also for the purpose of improving the consistency of the finally generated target action data and the original action data before action style transfer in terms of action content, the forward kinematics process can be applied to the input action (i.e., the second sample action data) and the generated action (i.e., the target sample action data) respectively according to the method shown in the following steps c1 - c2. According to the difference between the input action and the generated action after applying the forward kinematics process, the model parameters of the diffusion model can be further adjusted to further improve the consistency of the finally generated target action data and the original action data before action style transfer in terms of action content. Specifically:
[0105] Step c1: Apply the forward kinematics process to the second sample action data and the target sample action data respectively to obtain the first global position information of each joint of the action object in the second sample action data and the second global position information of each joint of the action object in the target sample action data.
[0106] Here, the forward kinematics process involves mapping the joint parameters (such as angles) of the action object to the three - dimensional pose information and position coordinates of the end - effector (i.e., in the three - dimensional space scene). By applying the forward kinematics process, the local position coordinates of each joint of the action object (i.e., the position coordinates of each joint in its own coordinate system) can be converted to the unified world coordinate system (i.e., the world coordinate system corresponding to the above three - dimensional space scene), and the coordinates of each joint in the above world coordinate system are used as the global position coordinates of each joint.
[0107] Specifically, after applying the forward kinematics process to the second sample action data, the above - mentioned global position coordinates of each joint of the action object in the second sample action data can be obtained as the above - mentioned first global position information.
[0108] Specifically, after applying the forward kinematics process to the target sample action data, the above - mentioned global position coordinates of each joint of the action object in the target sample action data can be obtained as the above - mentioned second global position information.
[0109] Step c2: Adjust the model parameters of the diffusion model according to the loss between the first global position information and the second global position information, to obtain the diffusion model including the adjusted model parameters.
[0110] Here, the L2 loss function can also be preferentially used to calculate the loss between the first global position information and the second global position information, that is, the mean square error between the first global position information and the second global position information can be calculated as the loss between the first global position information and the second global position information.
[0111] It should be noted that in addition to the L2 loss function, other types of loss functions can also be used to calculate the loss between the first global position information and the second global position information. The specific type of loss function actually used is not limited in the embodiments of the present application.
[0112] Specifically, after calculating the loss between each group of first global position information and second global position information, the sum of the losses between each group of first global position information and second global position information can be calculated as the overall model loss of the diffusion model. On the basis of the above steps S401 - S405, the diffusion model trained in the above steps S401 - S405 is continuously trained until the diffusion model converges again (equivalent to continuously adjusting the model parameters of the diffusion model until the above overall model loss reaches the minimum), and the diffusion model after re - convergence is obtained as the diffusion model trained after the above steps c1 - c2 (that is, the diffusion model including the adjusted model parameters).
[0113] Here, in addition to improving the consistency of the finally generated target action data and the original action data before action style transfer in terms of action content, on the action style side, it is also necessary to improve the matching degree between the finally generated target action data and the action style description information in terms of action style. At this time, as an optional embodiment, an additional pre - trained action style extraction network can be added on the basis of the training method described in the above steps S401 - S405, and the model parameters of the diffusion model can be further adjusted according to the method shown in the following steps d1 - d3, so as to improve the matching degree between the finally generated target action data and the action style description information in terms of action style. Specifically:
[0114] Step d1: From multiple third - sample action data, obtain target third - sample action data whose action style label matches the sample action style description information.
[0115] Here, the data acquisition method of the third - sample action data in step d1 can refer to the data acquisition method of the first - sample action data in the foregoing step S301, and the repeated parts will not be elaborated here.
[0116] Specifically, multiple action style categories can be preset. For each third sample action data, an action style category that matches the action style of the third sample action data is determined from the above-mentioned multiple preset action style categories, and the action style label corresponding to the action style category is determined as the action style label of the third sample action data.
[0117] Exemplarily, taking the above-mentioned multiple preset action style categories including "juvenile male", "juvenile female", "adolescent male", "adolescent female", "adult male", "adult female", "elderly male", and "elderly female" as an example, if the third sample action data is an action image sequence of a virtual character of a young girl walking, the action style label of the third sample action data can be determined as "adolescent female".
[0118] Specifically, when performing step d1, according to the above-mentioned sample action style description information determined by the action style, a third sample action data whose action style label matches the above-mentioned sample action style description information can be randomly sampled from multiple third sample action data as the above-mentioned target third sample action data.
[0119] Step d2: Respectively extract the action style features of the target sample action data and the target third sample action data through a pre-trained action style extraction network.
[0120] Here, the above-mentioned action style extraction network can be a pre-trained encoding network. For example, it can be a transformer Encoder network containing multiple Transformer encoding layers.
[0121] Specifically, after inputting the above-mentioned target sample action data into the action style extraction network, the feature vector output by the last Transformer encoding layer in the action style extraction network can be obtained as the action style feature of the target sample action data.
[0122] Specifically, after inputting the above-mentioned target third sample action data into the action style extraction network, the feature vector output by the last Transformer encoding layer in the action style extraction network can be obtained as the action style feature of the target third sample action data.
[0123] Step d3: Adjust the model parameters of the diffusion model according to the loss between the action style features of the target sample action data and the action style features of the target third sample action data, and obtain the diffusion model including the adjusted model parameters.
[0124] Here, the L2 loss function can also be preferentially used to calculate the loss between the action style feature of the target sample action data and the action style feature of the target third sample action data. That is, the mean square error between the above two action style features can be calculated as the loss between the above two action style features.
[0125] It should be noted that in addition to the L2 loss function, other types of loss functions can also be used to calculate the loss between the above two action style features. The specific type of loss function actually used is not limited in any way in the embodiments of the present application.
[0126] Specifically, based on the above steps S401 - S405 or the above steps c1 - c2, the diffusion model trained in the above steps S401 - S405 or the diffusion model trained in the above steps c1 - c2 can be continuously trained according to the loss calculated in the above step d3 until the diffusion model converges again, and the diffusion model after re - convergence is obtained as the diffusion model trained after the above steps d1 - d3 (that is, a diffusion model including adjusted model parameters).
[0127] Here, in an alternative embodiment, Figure 5 shows a schematic flowchart of a training method for an action style extraction network provided by an embodiment of the present application, as Figure 5 shown, the training method includes steps S501 - S503; specifically:
[0128] S501, input a plurality of third sample action data into the action style extraction network, and output the action style feature of each third sample action data through the action style extraction network.
[0129] Here, the specific execution manner of step S501 can refer to the specific execution manner of the foregoing step d2, and the repeated parts will not be elaborated here.
[0130] S502, for each third sample action data, input the action style feature of the third sample action data into a classifier, and predict the action style label to which the action style feature of the third sample action data belongs through the classifier, and output the predicted action style label of the third sample action data.
[0131] Here, after inputting the action style feature of the third sample action data into the classifier, the classifier can perform multi - classification prediction on the action style label to which the action style feature of the third sample action data belongs, predict the probability that the action style feature of the third sample action data belongs to each preset action style, and determine the preset action style with the highest probability as the predicted action style label of the third sample action data.
[0132] S503. Adjust the model parameters of the action style extraction network and the classifier according to the classification prediction loss between the predicted action style label of the third sample action data and the action style label to which the third sample action data belongs, so as to obtain the action style extraction network and the classifier including the adjusted model parameters.
[0133] Here, the cross-entropy loss function can be used to calculate the classification prediction loss between the predicted action style label of the third sample action data and the action style label to which the third sample action data belongs. Among them, in addition to the cross-entropy loss function, other types of classification loss functions can also be used to calculate the classification prediction loss between the predicted action style label of the third sample action data and the action style label to which the third sample action data belongs.
[0134] Specifically, after calculating the above classification prediction loss, the model parameters of the action style extraction network and the classifier can be adjusted according to the calculated classification prediction loss until the classification network composed of the action style extraction network and the classifier converges, so as to obtain the action style extraction network and the classifier including the adjusted model parameters.
[0135] Based on the above action generation method provided by the embodiments of the present application, input the original action data into the encoder of the pre-trained codec model to obtain the action encoding features output by the encoder for the original action data; input the action style description information into the pre-trained target encoder to obtain the style description features output by the target encoder for the action style description information; input the action encoding features and the style description features into the pre-trained diffusion model, use the style description features as the guiding control condition during the denoising process of the diffusion model, and output the target action encoding features that match the action encoding features and the style description features through the diffusion model; input the target action encoding features into the decoder of the codec model to obtain the target action data output by the decoder for the target action encoding features. In this way, the present application can support the migration of the action style of the original action data without changing the action content, and does not require professional venues, equipment, personnel, etc., which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data.
[0136] Based on the same inventive concept, the present application also provides an action generation device corresponding to the above action generation method. Since the principle of solving problems by the action generation device in the embodiments of the present application is similar to that of the above action generation method in the embodiments of the present application, the implementation of the action generation device can refer to the implementation of the above action generation method, and the repeated parts will not be described again.
[0137] Refer to Figure 6 as shown Figure 6The figure shows a schematic structural diagram of an action generation device provided by an embodiment of the present application. Among them, the action generation device includes:
[0138] A first encoding module 601, configured to input the original action data into an encoder of a pre-trained codec model, and obtain action encoding features output by the encoder for the original action data;
[0139] A second encoding module 602, configured to input action style description information into a target encoder, and obtain style description features output by the target encoder for the action style description information;
[0140] A prediction module 603, configured to input the action encoding features and the style description features into a pre-trained diffusion model, use the style description features as a guiding control condition during denoising processing of the diffusion model, and output target action encoding features that match the action encoding features and the style description features through the diffusion model;
[0141] A decoding module 604, configured to input the target action encoding features into a decoder of the codec model, and obtain target action data output by the decoder for the target action encoding features; wherein, the action content of the target action data matches the action content of the original action data, and the action style of the target action data matches the action style description information.
[0142] In an optional implementation manner, when using the style description features as a guiding control condition during denoising processing of the diffusion model and outputting target action encoding features that match the action encoding features and the style description features through the diffusion model, the prediction module 603 is configured to:
[0143] Perform noise addition processing on the action encoding features through the diffusion model to obtain the noise-added action encoding features;
[0144] Use the style description features as a guiding control condition, predict the denoising result of the noise-added action encoding features through the diffusion model, and output the predicted denoising result as the target action encoding features.
[0145] In an optional implementation manner, the action generation device further includes a first training module, where the first training module is used to train the codec model through the following method:
[0146] Input a plurality of first sample action data into the encoder, and output first action encoding features of each first sample action data through the encoder;
[0147] For each first sample action data, input the first action encoding feature of the first sample action data into the decoder, and output the restored action data of the first action encoding feature through the decoder;
[0148] Adjust the model parameters of the codec model according to the loss between the first sample action data and the restored action data, and obtain the codec model including the adjusted model parameters.
[0149] In an alternative embodiment, the action generation device further includes a second training module, wherein the second training module is used to train the diffusion model by the following method:
[0150] Input multiple second sample action data into the encoder, and output the second action encoding features of each second sample action data through the encoder;
[0151] Input the sample action style description information into the target encoder, and obtain the sample style description features output by the target encoder for the sample action style description information;
[0152] Input the second action encoding features and the sample style description features into the diffusion model, use the sample style description features as the guiding control conditions during the denoising process of the diffusion model, and output sample action encoding features that match the second action encoding features and the sample style description features through the diffusion model;
[0153] Input the sample action encoding features into the decoder, and obtain the target sample action data output by the decoder for the sample action encoding features;
[0154] Adjust the model parameters of the diffusion model according to the loss between the second action encoding features and the sample action encoding features, and obtain the diffusion model including the adjusted model parameters.
[0155] In an alternative embodiment, the second training module is further used for:
[0156] Apply a forward kinematics process to the second sample action data and the target sample action data respectively, and obtain the first global position information of each joint of the action object in the second sample action data and the second global position information in the target sample action data;
[0157] Adjust the model parameters of the diffusion model according to the loss between the first global position information and the second global position information, and obtain the diffusion model including the adjusted model parameters.
[0158] In an alternative implementation, the second training module is further configured to:
[0159] Obtain target third sample action data whose action style label matches the sample action style description information from multiple third sample action data;
[0160] Extract the action style features of the target sample action data and the target third sample action data respectively through a pre-trained action style extraction network;
[0161] Adjust the model parameters of the diffusion model according to the loss between the action style features of the target sample action data and the action style features of the target third sample action data, and obtain the diffusion model including the adjusted model parameters.
[0162] In an alternative implementation, the second training module is further configured to train the action style extraction network through the following method:
[0163] Input multiple third sample action data into the action style extraction network, and output the action style features of each third sample action data through the action style extraction network;
[0164] For each third sample action data, input the action style features of the third sample action data into a classifier, and predict the action style label to which the action style features of the third sample action data belong through the classifier, and output the predicted action style label of the third sample action data;
[0165] Adjust the model parameters of the action style extraction network and the classifier according to the classification prediction loss between the predicted action style label of the third sample action data and the action style label to which the third sample action data belongs, and obtain the action style extraction network and the classifier including the adjusted model parameters.
[0166] Based on the above-mentioned action generation device provided by the embodiments of the present application, the original action data is input into the encoder of the pre-trained codec model to obtain the action encoding features output by the encoder for the original action data; the action style description information is input into the pre-trained target encoder to obtain the style description features output by the target encoder for the action style description information; the action encoding features and the style description features are input into the pre-trained diffusion model, and the style description features are used as the guiding control conditions during the denoising process of the diffusion model, and the target action encoding features matching the action encoding features and the style description features are output through the diffusion model; the target action encoding features are input into the decoder of the codec model to obtain the target action data output by the decoder for the target action encoding features. In this way, the present application can support the transfer of the action style without changing the action content of the original action data, and does not require professional venues, equipment, personnel, etc., which is beneficial to improving the generation efficiency of stylized action data and reducing the action style transfer cost of the original action data.
[0167] Based on the same inventive concept, the present application also provides an electronic device corresponding to the above-mentioned action generation method. Since the principle of solving problems by the electronic device in the embodiments of the present application is similar to that of the above-mentioned action generation method in the embodiments of the present application, the implementation of the electronic device can refer to the implementation of the above-mentioned action generation method, and the repeated parts will not be described again.
[0168] Figure 7 The following is a schematic structural diagram of an electronic device 700 provided by an embodiment of the present application, including: a processor 701, a memory 702, and a bus 703. The memory 702 stores machine-readable instructions executable by the processor 701. When the electronic device runs an action generation method as in the embodiment, the processor 701 communicates with the memory 702 through the bus 703, and the processor 701 executes the machine-readable instructions. Among them, when the processor 701 executes the machine-readable instructions, the following steps are implemented, specifically:
[0169] Input the original action data into the encoder of the pre-trained codec model to obtain the action encoding features output by the encoder for the original action data;
[0170] Input the action style description information into the pre-trained target encoder to obtain the style description features output by the target encoder for the action style description information;
[0171] Input the action encoding features and the style description features into the pre-trained diffusion model, and use the style description features as the guiding control conditions during the denoising process of the diffusion model, and output the target action encoding features matching the action encoding features and the style description features through the diffusion model;
[0172] Input the target action encoding feature into the decoder of the codec model to obtain the target action data output by the decoder for the target action encoding feature; wherein, the action content of the target action data matches the action content of the original action data, and the action style of the target action data matches the action style description information.
[0173] In an alternative embodiment, when using the style description feature as the guiding control condition for the diffusion model during denoising processing, and outputting the target action encoding feature that matches the action encoding feature and the style description feature through the diffusion model, the processor 701 is configured to:
[0174] Perform noise addition processing on the action encoding feature through the diffusion model to obtain the noise-added action encoding feature;
[0175] Using the style description feature as the guiding control condition, predict the denoising result of the noise-added action encoding feature through the diffusion model, and output the predicted denoising result as the target action encoding feature.
[0176] In an alternative embodiment, the processor 701 is configured to train the codec model through the following method:
[0177] Input multiple first sample action data into the encoder, and output the first action encoding feature of each first sample action data through the encoder;
[0178] For each first sample action data, input the first action encoding feature of the first sample action data into the decoder, and output the restored action data of the first action encoding feature through the decoder;
[0179] Adjust the model parameters of the codec model according to the loss between the first sample action data and the restored action data to obtain the codec model including the adjusted model parameters.
[0180] In an alternative embodiment, the processor 701 is configured to train the diffusion model through the following method:
[0181] Input multiple second sample action data into the encoder, and output the second action encoding feature of each second sample action data through the encoder;
[0182] Input the sample action style description information into the target encoder to obtain the sample style description feature output by the target encoder for the sample action style description information;
[0183] Input the second action encoding feature and the sample style description feature into the diffusion model. Use the sample style description feature as the guiding control condition for the diffusion model during denoising processing, and output, through the diffusion model, a sample action encoding feature that matches the second action encoding feature and the sample style description feature;
[0184] Input the sample action encoding feature into the decoder to obtain target sample action data output by the decoder for the sample action encoding feature;
[0185] Adjust the model parameters of the diffusion model according to the loss between the second action encoding feature and the sample action encoding feature to obtain the diffusion model including the adjusted model parameters.
[0186] In an alternative embodiment, the processor 701 is further configured to:
[0187] Apply a forward kinematics process to the second sample action data and the target sample action data respectively to obtain the first global position information of each joint of the action object in the second sample action data and the second global position information in the target sample action data;
[0188] Adjust the model parameters of the diffusion model according to the loss between the first global position information and the second global position information to obtain the diffusion model including the adjusted model parameters.
[0189] In an alternative embodiment, the processor 701 is further configured to:
[0190] Obtain target third sample action data whose action style label matches the sample action style description information from multiple third sample action data;
[0191] Extract the action style features of the target sample action data and the target third sample action data respectively through a pre-trained action style extraction network;
[0192] Adjust the model parameters of the diffusion model according to the loss between the action style features of the target sample action data and the action style features of the target third sample action data to obtain the diffusion model including the adjusted model parameters.
[0193] In an alternative embodiment, the processor 701 is further configured to train the action style extraction network through the following method:
[0194] Input multiple third sample action data into the action style extraction network, and output, through the action style extraction network, the action style features of each third sample action data;
[0195] For each third sample action data, input the action style feature of the third sample action data into a classifier, and use the classifier to predict the action style label to which the action style feature of the third sample action data belongs, and output the predicted action style label of the third sample action data;
[0196] According to the classification prediction loss between the predicted action style label of the third sample action data and the action style label to which the third sample action data belongs, adjust the model parameters of the action style extraction network and the classifier, and obtain the action style extraction network and the classifier including the adjusted model parameters.
[0197] By using the above electronic device provided in the embodiments of the present application, input the original action data into the encoder of the pre-trained codec model to obtain the action coding feature output by the encoder for the original action data; input the action style description information into the pre-trained target encoder to obtain the style description feature output by the target encoder for the action style description information; input the action coding feature and the style description feature into the pre-trained diffusion model, use the style description feature as the guiding control condition during the denoising process of the diffusion model, and output the target action coding feature that matches the action coding feature and the style description feature through the diffusion model; input the target action coding feature into the decoder of the codec model to obtain the target action data output by the decoder for the target action coding feature. In this way, the present application can support the migration of the action style of the original action data without changing the action content, and does not require the assistance of professional venues, equipment, and personnel, etc., which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data.
[0198] Based on the same inventive concept, the embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the processor performs the following steps:
[0199] Input the original action data into the encoder of the pre-trained codec model to obtain the action coding feature output by the encoder for the original action data;
[0200] Input the action style description information into the pre-trained target encoder to obtain the style description feature output by the target encoder for the action style description information;
[0201] Input the action encoding feature and the style description feature into a pre-trained diffusion model, use the style description feature as the guiding control condition for the diffusion model during denoising processing, and output a target action encoding feature that matches the action encoding feature and the style description feature through the diffusion model;
[0202] Input the target action encoding feature into the decoder of the encoding-decoding model to obtain target action data output by the decoder for the target action encoding feature; wherein, the action content of the target action data matches the action content of the original action data, and the action style of the target action data matches the action style description information.
[0203] In an optional implementation manner, when using the style description feature as the guiding control condition for the diffusion model during denoising processing and outputting a target action encoding feature that matches the action encoding feature and the style description feature through the diffusion model, the processor is configured to:
[0204] Perform noise addition processing on the action encoding feature through the diffusion model to obtain a noise-added action encoding feature;
[0205] Use the style description feature as the guiding control condition, and predict the denoising result of the noise-added action encoding feature through the diffusion model, and output the predicted denoising result as the target action encoding feature.
[0206] In an optional implementation manner, the processor is configured to train the encoding-decoding model through the following method:
[0207] Input multiple first sample action data into the encoder, and output a first action encoding feature of each first sample action data through the encoder;
[0208] For each first sample action data, input the first action encoding feature of the first sample action data into the decoder, and output the restored action data of the first action encoding feature through the decoder;
[0209] Adjust the model parameters of the encoding-decoding model according to the loss between the first sample action data and the restored action data to obtain the encoding-decoding model including the adjusted model parameters.
[0210] In an optional implementation manner, the processor is configured to train the diffusion model through the following method:
[0211] Input multiple second sample action data into the encoder, and output a second action encoding feature of each second sample action data through the encoder;
[0212] Input the sample action style description information into the target encoder to obtain the sample style description features output by the target encoder for the sample action style description information;
[0213] Input the second action encoding feature and the sample style description feature into the diffusion model, use the sample style description feature as the guiding control condition for the diffusion model during denoising processing, and output the sample action encoding feature that matches the second action encoding feature and the sample style description feature through the diffusion model;
[0214] Input the sample action encoding feature into the decoder to obtain the target sample action data output by the decoder for the sample action encoding feature;
[0215] Adjust the model parameters of the diffusion model according to the loss between the second action encoding feature and the sample action encoding feature to obtain the diffusion model including the adjusted model parameters.
[0216] In an optional implementation manner, the processor is further configured to:
[0217] Apply a forward kinematics process to the second sample action data and the target sample action data respectively to obtain the first global position information of each joint of the action object in the second sample action data and the second global position information in the target sample action data;
[0218] Adjust the model parameters of the diffusion model according to the loss between the first global position information and the second global position information to obtain the diffusion model including the adjusted model parameters.
[0219] In an optional implementation manner, the processor is further configured to:
[0220] Obtain the target third sample action data whose action style label matches the sample action style description information from multiple third sample action data;
[0221] Extract the action style features of the target sample action data and the target third sample action data respectively through a pre-trained action style extraction network;
[0222] Adjust the model parameters of the diffusion model according to the loss between the action style features of the target sample action data and the action style features of the target third sample action data to obtain the diffusion model including the adjusted model parameters.
[0223] In an alternative embodiment, the processor is further configured to train the action style extraction network by the following method:
[0224] Input a plurality of third sample action data into the action style extraction network, and output the action style features of each third sample action data through the action style extraction network;
[0225] For each third sample action data, input the action style features of the third sample action data into a classifier, and predict the action style label to which the action style features of the third sample action data belong through the classifier, and output the predicted action style label of the third sample action data;
[0226] According to the classification prediction loss between the predicted action style label of the third sample action data and the action style label to which the third sample action data belongs, adjust the model parameters of the action style extraction network and the classifier to obtain the action style extraction network and the classifier including the adjusted model parameters.
[0227] Through the above computer-readable storage medium provided by the embodiments of the present application, input the original action data into the encoder of the pre-trained codec model to obtain the action encoding features output by the encoder for the original action data; input the action style description information into the pre-trained target encoder to obtain the style description features output by the target encoder for the action style description information; input the action encoding features and the style description features into the pre-trained diffusion model, and use the style description features as the guiding control condition during the denoising process of the diffusion model, and output the target action encoding features that match the action encoding features and the style description features through the diffusion model; input the target action encoding features into the decoder of the codec model to obtain the target action data output by the decoder for the target action encoding features. In this way, the present application can support the migration of the action style of the original action data without changing the action content, and does not require professional venues, equipment, personnel, etc., which is beneficial to improving the generation efficiency of stylized action data and reducing the action style migration cost of the original action data.
[0228] In the embodiments of the present application, when the computer-readable storage medium is run by the processor, other machine-readable instructions may also be executed to execute the interaction method in the game as described in other embodiments. For the specific steps and principles of the executed interaction method, refer to the description of the method-side embodiments, which will not be elaborated here.
[0229] In the embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections of systems or units can be in electrical, mechanical or other forms.
[0230] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0231] In addition, each functional unit in the embodiments provided in the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0232] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0233] It should be noted that similar reference numerals and letters denote similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0234] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions described in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims described.
Claims
1. An action generation method, characterized in that: The action generation method comprises: Inputting the original motion data into an encoder of a pre-trained encoding and decoding model to obtain motion coding features output by the encoder for the original motion data; Inputting the action style description information into a pre-trained target encoder to obtain a style description feature output by the target encoder for the action style description information; Inputting the action coding feature and the style description feature into a pre-trained diffusion model, using the style description feature as a guiding control condition of the diffusion model during denoising, and outputting a target action coding feature that matches the action coding feature and the style description feature through the diffusion model; The target action coding feature is input into the decoder of the codec model to obtain the target action data output by the decoder for the target action coding feature; wherein the action content of the target action data matches the action content of the original action data, and the action style of the target action data matches the action style description information.
2. The action generation method according to claim 1, characterized in that: The method of using the style description feature as a guiding control condition of the diffusion model during denoising processing and outputting a target action coding feature that matches the action coding feature and the style description feature through the diffusion model includes: Performing noise processing on the action coding feature through the diffusion model to obtain the noisy action coding feature; The style description feature is used as a guiding control condition, and the denoising result of the noisy action coding feature is predicted through the diffusion model, and the predicted denoising result is output as the target action coding feature.
3. The action generation method according to claim 1, characterized in that: The encoding and decoding model is obtained by training using the following method: Inputting a plurality of first sample action data into the encoder, and obtaining a first action encoding feature of each first sample action data through the output of the encoder; For each first sample action data, input the first action coding feature of the first sample action data into the decoder, and output the restored action data of the first action coding feature through the decoder; According to the loss between the first sample action data and the restored action data, the model parameters of the encoding and decoding model are adjusted to obtain the encoding and decoding model including the adjusted model parameters.
4. The action generation method according to claim 1, characterized in that: The diffusion model is obtained by training using the following method: Inputting a plurality of second sample action data into the encoder, and obtaining a second action encoding feature of each second sample action data through the output of the encoder; Inputting the sample action style description information into the target encoder to obtain the sample style description features output by the target encoder for the sample action style description information; Inputting the second action coding feature and the sample style description feature into a diffusion model, taking the sample style description feature as a guiding control condition of the diffusion model during denoising, and outputting a sample action coding feature that matches the second action coding feature and the sample style description feature through the diffusion model; Inputting the sample action coding feature into the decoder to obtain target sample action data output by the decoder for the sample action coding feature; According to the loss between the second action coding feature and the sample action coding feature, the model parameters of the diffusion model are adjusted to obtain the diffusion model including the adjusted model parameters.
5. The action generation method according to claim 4, characterized in that: The training method of the diffusion model also includes: Applying a forward kinematics process to the second sample motion data and the target sample motion data respectively, to obtain first global position information of each joint of the motion object in the second sample motion data and second global position information in the target sample motion data; According to the loss between the first global position information and the second global position information, the model parameters of the diffusion model are adjusted to obtain the diffusion model including the adjusted model parameters.
6. The action generation method according to claim 4, characterized in that: The training method of the diffusion model also includes: Acquire, from a plurality of third sample action data, target third sample action data whose action style label matches the sample action style description information; Extracting the action style features of the target sample action data and the target third sample action data respectively through a pre-trained action style extraction network; According to the loss between the motion style feature of the target sample motion data and the motion style feature of the target third sample motion data, the model parameters of the diffusion model are adjusted to obtain the diffusion model including the adjusted model parameters.
7. The action generation method according to claim 6, characterized in that: The action style extraction network is trained by the following method: Inputting a plurality of third sample action data into the action style extraction network, and obtaining the action style feature of each third sample action data through the output of the action style extraction network; For each third sample action data, inputting the action style feature of the third sample action data into a classifier, predicting the action style label to which the action style feature of the third sample action data belongs through the classifier, and outputting the predicted action style label of the third sample action data; According to the classification prediction loss between the predicted action style label of the third sample action data and the action style label to which the third sample action data belongs, the model parameters of the action style extraction network and the classifier are adjusted to obtain the action style extraction network and the classifier including the adjusted model parameters.
8. An action generating device, characterized in that: The action generating device comprises: A first encoding module, used for inputting the original motion data into an encoder of a pre-trained encoding and decoding model to obtain motion encoding features output by the encoder for the original motion data; A second encoding module is used to input the action style description information into a pre-trained target encoder to obtain a style description feature output by the target encoder for the action style description information; A prediction module, configured to input the action coding feature and the style description feature into a pre-trained diffusion model, use the style description feature as a guiding control condition of the diffusion model during denoising, and output a target action coding feature that matches the action coding feature and the style description feature through the diffusion model; A decoding module is used to input the target action coding feature into the decoder of the encoding and decoding model to obtain the target action data output by the decoder for the target action coding feature; wherein the action content of the target action data matches the action content of the original action data, and the action style of the target action data matches the action style description information.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the action generation method as described in any one of claims 1 to 7 are performed.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the action generation method as described in any one of claims 1 to 7 are executed.