Action generation method and device, electronic equipment and storage medium
By extracting and encoding features of the bone files and music files of the virtual object, the action sequence of the virtual object is generated, which solves the problem of low action generation efficiency in the prior art, and achieves more efficient and high-quality action generation.
Patent Information
- Application Number
- CN202510179039.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
AI Technical Summary
In the prior art, the action generation efficiency of virtual objects is low, and a specific posture action is required to manually set in keyframes, and the interpolation process is cumbersome.
By obtaining the bone files and music files of the virtual object, bone information processing and music feature extraction are performed, the fused vector is encoded and decoder to generate the action sequence of the virtual object.
It improves the action generation efficiency of virtual objects, reduces manual intervention, and the generated action sequence is more in line with the music rhythm and emotion, and improves the quality and fluency of the animation.
Smart Images

Figure CN120125720A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to a method, an apparatus, an electronic device, and a storage medium for generating actions. Background Art
[0002] In the wave of the Internet, entertainment projects have become increasingly important in people's lives. To meet the requirements of presenting scene effects in some entertainment projects (such as animated movies and games), it is necessary to endow virtual objects in the scene with diverse action performances, especially the setting of the actions of virtual objects.
[0003] Currently, the actions of virtual objects often require relevant staff to set specific pose actions at key frames to define the object actions at important time nodes, and then perform interpolation based on the time intervals and action changes between key frames to obtain the transitional actions between key frames, thereby forming continuous animations. However, since it takes a lot of time and effort for humans to design actions for virtual objects, the efficiency of generating actions of virtual objects is relatively low. Summary of the Invention
[0004] Embodiments of the present application provide a method, an apparatus, an electronic device, and a storage medium for generating actions, which can improve the efficiency of generating actions of virtual objects.
[0005] In a first aspect, embodiments of the present application provide a method for generating actions, the method including:
[0006] Obtain a skeleton file of a virtual object, perform skeleton information processing on the skeleton file to obtain a skeleton information vector of the virtual object;
[0007] Obtain a music file, perform feature extraction on the music file to obtain a music feature vector of the music file;
[0008] Perform encoding processing on the skeleton information vector and the music feature vector through at least one encoder to obtain a fused vector;
[0009] Obtain an action sequence of the virtual object through at least one decoder based on the fused vector.
[0010] In a second aspect, embodiments of the present application provide an apparatus for generating actions, the apparatus including:
[0011] An information processing module, configured to obtain a skeleton file of a virtual object, perform skeleton information processing on the skeleton file to obtain a skeleton information vector of the virtual object;
[0012] A feature extraction module, configured to obtain a music file, perform feature extraction on the music file to obtain a music feature vector of the music file;
[0013] An encoding module, configured to encode the above-mentioned skeletal information vector and the above-mentioned music feature vector through at least one encoder to obtain a fused vector;
[0014] An action generation module, configured to obtain an action sequence of the above-mentioned virtual object through at least one decoder based on the above-mentioned fused vector.
[0015] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory storing multiple instructions; a processor loads the instructions from the memory to execute the steps of any one of the action generation methods provided in the embodiments of the present application.
[0016] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps of any one of the action generation methods provided in the embodiments of the present application.
[0017] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program or instructions, and when the computer program or instructions are executed by a processor, the steps in any one of the action generation methods provided in the embodiments of the present application are implemented.
[0018] By adopting the solution of the embodiment of the present application, the skeletal information vector of the virtual object can be obtained by acquiring the skeletal file of the virtual object and performing skeletal information processing on the above-mentioned skeletal file; the music file is acquired, and feature extraction is performed on the above-mentioned music file to obtain the music feature vector of the above-mentioned music file; the above-mentioned skeletal information vector and the above-mentioned music feature vector are encoded through at least one encoder to obtain a fused vector; an action sequence of the above-mentioned virtual object is obtained through at least one decoder based on the above-mentioned fused vector, so as to automatically generate an action sequence of the virtual object through the skeletal information vector of the skeleton and the music feature vector of the music, thereby improving the action generation efficiency of the virtual object. Description of the Drawings
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 It is a schematic flowchart of an embodiment of the action generation method provided in the embodiment of the present application;
[0021] Figure 2 It is a schematic diagram of the generation process of the skeletal information vector provided in the embodiment of the present application;
[0022] Figure 3 It is a schematic diagram of the generation process of the music attribute feature vector provided in the embodiment of the present application;
[0023] Figure 4 It is a schematic diagram of the generation process of the music emotion feature vector provided in the embodiment of the present application;
[0024] Figure 5 It is a schematic diagram of the generation process of the fused vector provided in the embodiment of the present application;
[0025] Figure 6 It is a schematic diagram of the structure of the action generation device provided in the embodiment of the present application;
[0026] Figure 7 It is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application. At the same time, in the description of the embodiments of the present application, terms such as "first" and "second" are only used for distinguishing descriptions, and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, "a plurality" means two or more, unless otherwise specifically defined.
[0028] The embodiment of the present application provides an action generation method, device, electronic device and computer-readable storage medium.
[0029] Specifically, this embodiment will be described from the perspective of the action generation device. The action generation device can be specifically integrated in an electronic device, that is, the action generation method in the embodiment of the present application can be executed by the electronic device. Optionally, the electronic device may include: a terminal device. The terminal device can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a game console, or a personal computer (PC) and other devices.
[0030] The action generation method provided by the embodiment of the present application can be applied to an action generation system. Among them, the action generation system can include a player terminal device and a server. The terminal can be a device that includes both receiving and transmitting hardware, that is, a device with receiving and transmitting hardware capable of performing two-way communication on a two-way communication link. The player terminal device and the server can perform two-way communication through a network.
[0031] Optionally, the server can be an independent server, or a server network or server cluster composed of servers, including but not limited to computers, network hosts, single network servers, multiple sets of network servers, or cloud servers composed of multiple servers. Among them, the cloud server is composed of a large number of computers or network servers based on cloud computing.
[0032] The following will be described in detail with reference to the accompanying drawings. In this embodiment, the execution entity is taken as the terminal device. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown in the drawings.
[0033] The action generation method of this embodiment obtains the skeleton file of the virtual object, processes the skeleton information of the above skeleton file to obtain the skeleton information vector of the above virtual object; obtains the music file, extracts the features of the above music file to obtain the music feature vector of the above music file; encodes the above skeleton information vector and the above music feature vector through at least one encoder to obtain a fused vector; and obtains the action sequence of the above virtual object through at least one decoder based on the above fused vector, which can improve the action generation efficiency of the virtual object.
[0034] Please refer to Figure 1 , taking the terminal as an example for illustration. This embodiment provides an action generation method. The specific process of this action generation method can be as follows in steps 101 to 104, where:
[0035] Step 101: Obtain the skeleton file of the virtual object, and process the skeleton information of the above skeleton file to obtain the skeleton information vector of the above virtual object.
[0036] Among them, the above virtual object needs to generate corresponding object actions, which can be dance actions or corresponding work actions, etc., and can be specifically set according to requirements and are not limited here.
[0037] Among them, the skeleton of the virtual object includes but is not limited to the skeleton parts and skeleton joints that make up the skeleton structure, and can be specifically set according to requirements and are not limited here.
[0038] In this embodiment, the terminal processes the bone information of the virtual object's bone file to obtain the bone information vector of the virtual object, so as to indicate the bone structure relationship of the virtual object and / or the bone movement relationship of the bone at at least one moment. Among them, the above-mentioned bone information vector is used to describe the motion characteristics of the bone, that is, based on the information representation of the pose information to be predicted at each motion moment, the bone information vector of the bone is clarified.
[0039] In some embodiments, the above-mentioned bone file includes the initial bone pose information of the virtual object's bone at the initial moment, the above-mentioned action sequence includes the pose information of the virtual object's bone at multiple moments, and the above-mentioned obtaining the virtual object's bone file and processing the bone information of the above-mentioned bone file to obtain the above-mentioned virtual object's bone information vector includes: through the motion feature association of the pose information of the bone before and after the motion in the first kinematic equation, based on the above-mentioned initial bone pose information, calculate the information representation of the pose information to be predicted of the above-mentioned bone at multiple moments, where the information representation of the pose information to be predicted at each moment is obtained based on the information representation of the pose information to be predicted before the corresponding moment; based on the above-mentioned initial bone pose information and the information representation of the pose information to be predicted at the above-mentioned multiple moments, obtain the bone information vector of the above-mentioned virtual object.
[0040] Among them, the above-mentioned pose information includes but is not limited to position, joint angle, direction, etc., and can be specifically set according to requirements and will not be limited here.
[0041] In this embodiment, the terminal can obtain the initial bone pose information of the virtual object's bone at the initial moment, so as to calculate the pose information of the bone at different moments based on the initial bone pose information of the virtual object's bone at the initial moment, so as to form the actions of the bone at different moments, that is, represent the object actions of the virtual object at the corresponding moments through the pose information of the bone, such as forming the dance actions of the virtual object.
[0042] Among them, the above-mentioned first kinematic equation can be used to characterize the correlation of the skeleton before and after movement, that is, the first kinematic equation can be an equation corresponding to the forward kinematics (FK) algorithm. The equation corresponding to the FK algorithm can calculate the position and posture of the end of the skeleton under given motion parameters based on the initial skeleton posture information, for example, the structure of the skeleton and the initial position and angle of the joints. Since at least some parameters in the given motion parameters need to be clarified based on subsequent steps, it is necessary to calculate the information representation of the predicted posture information of the skeleton at each motion moment, so that after obtaining the motion parameters, the posture information of the skeleton at the corresponding moment, such as the position of each joint of the skeleton, can be obtained based on the information representation of the predicted posture information, thereby ensuring the rationality and coherence of the object movement of the virtual object at the skeleton level.
[0043] Among them, the information representation of the above-mentioned posture information to be predicted is used to indicate the expression method of the posture information to be predicted, such as the information representation contains motion parameters that need to be filled in, so that after the motion parameters are filled in the information representation, the posture information at the corresponding moment can be obtained based on the filled-in motion parameters.
[0044] It should be noted that the equations corresponding to the FK algorithm can forward infer the position of the bone ends based on the initial state of the joints and the motion parameters, and construct the basic framework of the action, which is suitable for generating a coherent and natural preliminary motion trajectory. In the initial construction stage of the object's action, the equations corresponding to the FK algorithm can generate or manually set basic motion parameters, such as the direction of movement, based on characteristic information such as the rhythm of the music, and then, based on the generated basic motion parameters, calculate the posture information such as the position of each joint in the skeleton for use in forming the action outline.
[0045] Specifically, the above-mentioned motion feature association is based on the motion parameter representation, and the above-mentioned motion feature association of the posture information of the bones before and after the movement in the first kinematic equation, on the basis of the above-mentioned initial bone posture information, calculates the information representation of the posture information to be predicted of the above-mentioned bones at multiple moments, which may include: the association relationship between the motion parameters provided by the first kinematic equation and the posture information of the bones before and after the movement, and on the basis of the above-mentioned initial bone posture information, calculates the information representation of the posture information to be predicted of the above-mentioned bones at multiple moments.
[0046] Correspondingly, the information representation of the posture information to be predicted at each of the above-mentioned moments is obtained based on the information representation of the posture information to be predicted before the corresponding moment, and may include: the information representation of the posture information to be predicted at each moment is obtained based on the motion parameters of the corresponding moment, and the information representation of the posture information to be predicted before the corresponding moment.
[0047] Among them, some of the above motion parameters can be set by the user, such as the direction of motion, the direction of rotation, etc., while the other part of the parameters need to be obtained by model processing and can be unknown parameters. In this example, the given motion parameters are not only for the motion parameters corresponding to the initial bone pose information, but can also include the motion parameters corresponding to the bone pose information at each motion moment. The initial bone pose information can be used as a basis to drive the change of the bone pose at the initial moment. During the entire motion process, the motion parameters at each motion moment can be continuously changed and updated to prompt the corresponding action changes of the bone.
[0048] For example, at the start of an object action sequence, the shoulder joint angle is 0 degrees, but as the action progresses, a series of new angles corresponding to different moments, as well as new motion parameters such as translation parameters and speed parameters, will be generated to determine the position and posture of the bone at different moments. Thus, the exact position and posture of the bone end during the entire motion process can be calculated through the equations corresponding to the FK algorithm to obtain the information representation of the to-be-predicted pose information of the above bone at multiple motion moments.
[0049] Among them, the above motion parameters include but are not limited to joint angle change parameters, translation parameters, speed parameters, etc., and can be specifically set according to requirements and are not limited here.
[0050] Among them, the above joint angle change parameters include but are not limited to rotation angles, flexion and extension angles, etc.
[0051] Among them, the above rotation angle can be indicated as the angle of rotation around a certain axis, that is, it is used to describe the angle of joint rotation around the coordinate axis. For the human bone model, each joint has its specific rotation freedom. Taking the shoulder joint in the human bone model as an example, it can rotate around three coordinate axes (x, y, z). When the arm needs to be waved forward with the shoulder joint as the center, the shoulder joint will rotate around a certain axis (such as the y-axis) by a certain angle, which determines the position change of the bone (arm part) in space.
[0052] Among them, the above flexion and extension angle can be indicated as the angle of flexion and extension in a certain direction. For example, for some joints, such as the knee joint and the elbow joint, the flexion and extension angle is a key motion parameter. Taking the knee joint as an example, during actions such as walking or jumping, the knee joint will flex and extend within a certain range, and the magnitude and change rate of this flexion and extension angle will affect the position and posture of the entire lower limb bone.
[0053] Exemplarily, the above rotation angles can be set as the abduction and adduction angles around the y-axis and the internal rotation and external rotation angles around the z-axis, etc. The above flexion and extension angles can be set as the flexion and extension angles around the x-axis, and the value range of the angles can be determined according to the physiological structure and movement range of the bone. For the elbow joint and knee joint, their main rotation angles are the flexion and extension angles, and their value range is generally the angle corresponding to the straight state to the angle corresponding to the maximum bending degree.
[0054] Specifically, the above translation parameters include but are not limited to linear displacement parameters, translation directions, translation distances, etc., to indicate the translation mode of the joints of the bone in three-dimensional space. Among them, since some parts of the bone may have linear displacements. For example, in some object actions corresponding to the human body bone, the pelvis may translate in the horizontal direction (x-axis) or in the vertical direction (y-axis); for another example, for the bones of a virtual object, if you want to simulate actions such as skating, the whole may produce continuous linear displacements along the skating direction (assumed to be the z-axis), so as to accurately calculate the position of the bone end in space through the translation parameters.
[0055] Exemplarily, taking the hip joint as an example, when generating the bone vector of the walking action, there will be parameters of the translation distance in the x-axis (front-back direction), the translation distance in the y-axis (left-right direction) and / or the translation distance in the z-axis (up-down direction). For example, in a simple walking action, the hip joint may have a parameter of how many meters it translates per step in the x-axis direction (the specific value depends on factors such as the scale of the bone model), the translation parameter in the y-axis direction may be close to 0 (assuming the walking route is a straight line), and there will be a parameter of how many meters it slightly floats up and down in the z-axis direction to reflect the body undulation during walking.
[0056] Specifically, the above speed parameters include but are not limited to joint angular velocity, joint linear velocity, etc. The speed parameter determines the speed of the bone movement. The above joint angular velocity can be the speed when describing the rotation movement of the bone (such as the joint rotating around a certain axis); the above joint linear velocity can be the speed when describing the translation movement of the bone. Among them, for the rotation or translation of the joints of the bone, there will be corresponding speed parameters. For example, for the rotation movement, the speed parameter can be set as how many degrees it rotates per second. For example, if the speed parameter of the wrist joint rotating around a certain coordinate axis is 15° / s, then after 1 second, the angle of the wrist joint rotating around a certain coordinate axis is 15°; for the translation movement, the speed parameter can be set as how many meters it translates per second. For example, if the translation speed of the knee joint in the x-axis direction is 1 m / s, then after 1 second, the displacement of the knee joint in the x-axis direction is 1 m.
[0057] Exemplarily, if it is set that the human lower limb bones need to be controlled to move currently, the human lower bones may include the hip joint, knee joint, ankle joint, etc. Then, the initial position coordinates, initial angles, and lengths of each part of the limb bones in space relative to a certain reference coordinate system can be obtained, such as the length from the hip joint to the knee joint, the length from the knee joint to the ankle joint, etc. Moreover, in the corresponding motion parameters further set, for example, corresponding joint angle change parameters (including flexion to extension, abduction to adduction, internal rotation to external rotation angles) are set for the hip joint, as well as corresponding translation parameters (for example, indicating a certain distance of translation in the x-axis direction).
[0058] Then, based on the set joint angle change parameters and translation parameters, matrix transformations of rotation and translation are performed to update the position and pose of the hip joint, that is, by determining the rotation matrix corresponding to the flexion-to-extension angle that needs to rotate around the x-axis, the updated position of the hip joint is obtained. Among them, if it is assumed that the rotation center is at the initial position of the hip joint, then since the initial position vector is a zero vector, the result after rotation based on the initial position vector is only a formal representation.
[0059] Then, for the position of the knee joint, by establishing a local coordinate system with the updated hip joint as the origin and based on the length from the hip joint to the knee joint, the position coordinates of the knee joint in this local coordinate system are obtained. Among them, the position coordinates of the knee joint in this local coordinate system can be obtained by assuming that the lower limb moves near the xy plane in the initial state and the displacement in the z direction is 0, and more complex motion parameters can be introduced to adjust the actual situation.
[0060] Finally, for the position of the ankle joint, by setting the flexion and extension angles and translation parameters corresponding to the knee joint (for example, indicating a certain distance of translation in the negative y-axis direction), and then based on the set parameters, the length from the knee joint to the ankle joint, and the position of the knee joint, the position of the ankle joint is obtained.
[0061] In some embodiments, it may further include: calculating the pose adjustment information of the to-be-predicted poses of the above bones at multiple moments through the second kinematic equation. Then, based on the above initial bone pose information and the information representation of the to-be-predicted poses at the above multiple moments, the bone information vector of the above virtual object is obtained, including: based on the above initial bone pose information, and the information representation and pose adjustment information corresponding to the to-be-predicted poses at the above multiple moments, the bone information vector of the above bones is obtained.
[0062] Among them, the above second kinematic equation can be the equation corresponding to the inverse kinematics (IK) algorithm. The equation corresponding to this IK algorithm can know the target position and attitude of the end of the skeleton, and inversely deduce the movements and angle adjustments that each joint in the skeleton needs to make. For example, during the generation of dance movements, when it is necessary to make a certain part of the skeleton (such as the hand) reach a specific target position (such as touching a virtual object), the equation corresponding to the IK algorithm can obtain the corresponding pose adjustment information to precisely adjust the movements of each joint to achieve this goal and ensure the accuracy and naturalness of the game dance movements.
[0063] It should be noted that the equation corresponding to the IK algorithm can inversely solve the pose adjustment information of each part or joint in the skeleton according to the target position of the end of the skeleton, and can play a key role in scenarios where precise control of the end position is required. For example, in scenarios where precise control of the end position is required for specific interaction actions, such specific interaction actions can be hand-holding actions, shoulder-touching actions, or specific pose actions in a duet dance, etc.
[0064] Among them, in the process of obtaining the skeleton information vector of the above skeleton based on the initial skeleton pose information at the above initial moment, the information representation and pose adjustment information of the to-be-predicted pose information at the above multiple movement moments, a weighted average strategy can be adopted, and weights can be set according to the emphasis on indicators such as fluency and accuracy of the action to fuse the results output by the equation corresponding to the FK algorithm and the results output by the equation corresponding to the IK algorithm; or, the results output by the FK algorithm and the results output by the IK algorithm can be fused by means of information superposition to achieve local precise adjustment of the generation action corresponding to the information representation generated by the FK algorithm.
[0065] It can be understood that as Figure 2 shown, the terminal can introduce a skeleton motion analysis module, which integrates the FK algorithm and the IK algorithm to parse and analyze the input skeleton file through the FK algorithm and the IK algorithm to obtain the to-be-predicted pose information representation and pose adjustment information, so as to output a skeleton information vector that can accurately understand the skeleton motion. Among them, the above skeleton file includes the initial skeleton pose information of the skeleton at the initial moment.
[0066] Step 102: Obtain a music file, extract features from the above music file to obtain the music feature vector of the above music file.
[0067] In this embodiment, in order to better generate the object actions of virtual objects that conform to the virtual scene, the music feature vector of the music file in the virtual scene can be used to participate in the generation of the object actions of the virtual objects.
[0068] In some embodiments, the above-mentioned music feature vector includes a music attribute feature vector and a music emotion feature vector. The above-mentioned feature extraction of the above-mentioned music file to obtain the music feature vector of the above-mentioned music file may include: first, extracting attribute features of the above-mentioned music file to obtain the music attribute feature vector of the above-mentioned music file, and then, based on the above-mentioned music file and / or the above-mentioned music attribute feature vector, extracting emotion features to obtain the music emotion feature vector of the above-mentioned music file.
[0069] Specifically, the above-mentioned extraction of attribute features of the above-mentioned music file to obtain the music attribute feature vector of the above-mentioned music file may include: first, extracting spectral features of the music signal in the above-mentioned music file to obtain the music spectral feature vector of the above-mentioned music file, then, performing frequency domain conversion on the above-mentioned music signal to obtain the converted frequency domain signal, and extracting frequency domain features of the above-mentioned frequency domain signal to obtain the music frequency domain feature vector of the above-mentioned music file, and finally, based on the above-mentioned music spectral feature vector and the above-mentioned music frequency domain feature vector, the extracted spectral features and frequency domain features are merged (for example, feature splicing) to obtain the music attribute feature vector of the above-mentioned music file.
[0070] Specifically, the step of extracting spectral features of the music signal in the above-mentioned music file may include: processing the music signal through a preset spectral feature extraction algorithm to obtain a music spectral feature vector of the music file, and the music spectral feature vector may include but is not limited to features corresponding to key elements such as pitch, timbre, rhythm, etc., thereby providing a quantifiable basis for the subsequent generation of game dance movements according to the music.
[0071] Among them, the above-mentioned preset spectral feature extraction algorithm can be a Mel Frequency Cepstral Coefficient (MFCC) algorithm, which can process the music signal in the music data, convert it into Mel Frequency Cepstral Coefficients, and then capture the spectral characteristics of the music from the converted Mel Frequency Cepstral Coefficients. By analyzing the feature vectors extracted by the MFCC algorithm, the speed of the music rhythm can be determined, thereby providing a reference for generating a dance movement rhythm that matches it.
[0072] Specifically, the step of extracting frequency domain features of the above-mentioned frequency domain signal may include: processing the music data through a preset frequency domain conversion algorithm to obtain a frequency domain signal of the music, and determining the frequency domain feature vector of the music based on the frequency domain signal, so as to accurately grasp the melody, harmony and other characteristics of the music by analyzing the energy distribution of the music in different frequency bands, so as to better associate the music elements with the dance movements, such as determining the liveliness or amplitude of the dance movements according to the energy strength of the high frequency band of the music.
[0073] Among them, the above-mentioned preset frequency-domain conversion algorithm can be the Short-Time Fourier Transform (STFT) algorithm. This STFT algorithm can perform a short-time Fourier transform on the music signal in the music data, converting it from a time-domain signal to a frequency-domain signal.
[0074] Exemplarily, as Figure 3 shown, the terminal can introduce a music feature extraction module. After obtaining a music file, which stores music data, the spectral features and frequency-domain features of the music file can be obtained respectively, and then the obtained audio features and frequency-domain features can be fused to obtain a music attribute feature vector.
[0075] Specifically, performing emotional feature extraction based on the above-mentioned music file and / or the above-mentioned music attribute feature vector to obtain the music emotional feature vector of the above-mentioned music file may include: First, performing emotional recognition based on the above-mentioned music attribute feature vector to obtain the initial emotional feature vector of the above-mentioned music file. Then, performing text extraction on the above-mentioned music file to obtain music text data, such as music lyrics, so as to perform semantic analysis on the above-mentioned music text data to obtain a semantic feature vector. Finally, based on the above-mentioned semantic feature vector and the above-mentioned initial emotional feature vector, the music emotional feature vector of the above-mentioned music file is obtained.
[0076] Among them, the step of performing semantic analysis on the above-mentioned music text data to obtain a semantic feature vector may include: inputting the text data into a semantic analysis model to perform semantic analysis on the text data, and extracting keywords, themes, and emotional tendencies (such as positive, negative, neutral), and generating an emotional feature representation related to music.
[0077] Among them, the above-mentioned semantic analysis model can be a pre-trained language (Bidirectional Encoder Representations from Transformers, BERT) model. This BERT model can learn the general feature representation of language from large-scale text data through unsupervised learning and can be used in specific scenarios, such as the emotional analysis scenario, etc.
[0078] Specifically, the step of obtaining a music emotional feature vector based on the above-mentioned music semantic features and the above-mentioned music attribute features may include: an emotional analysis model based on deep learning can be used to recognize emotional features (such as cheerful, sad, exciting, etc.) of the music attribute feature vector and output an initial emotional feature vector. Then, the initial emotional feature vector is updated through the emotional feature representation obtained from the music semantic features to obtain a music emotional feature vector.
[0079] Exemplarily, as Figure 4 shown, the music attribute feature vector can be input into a hybrid structure of a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM). The hybrid structure of the CNN and LSTM is the above-mentioned sentiment analysis model based on deep learning. Then, semantic extraction is performed on the music data to obtain a music sentiment feature vector based on the sentiment feature representation and the initial sentiment feature vector.
[0080] Among them, the above-mentioned CNN is a deep learning model that has been widely used in many fields such as image recognition, object detection, and speech recognition; the above-mentioned LSTM is a special Recurrent Neural Network (RNN) designed to solve the problem of gradient vanishing or gradient explosion that easily occurs when traditional RNNs process long sequence data, so as to better model long-distance dependency relationships.
[0081] Step 103: Encode the above-mentioned skeletal information vector and the above-mentioned music feature vector through at least one encoder to obtain a fused vector.
[0082] In this embodiment, the skeletal information vector and the feature vector are integrated into one vector, that is, the fused vector, through at least one encoder, so as to use the integrated vector as an input to predict the object action of the virtual object.
[0083] In some embodiments, the above-mentioned encoding the above-mentioned skeletal information vector and the above-mentioned music feature vector through at least one encoder to obtain a fused vector may include: First, splice the above-mentioned skeletal information vector and the above-mentioned music feature vector to obtain a spliced vector, and then encode the spliced vector based on at least one encoder to obtain an encoding result, and obtain the fused vector based on the above-mentioned encoding result.
[0084] Exemplarily, as Figure 5 shown, the terminal can splice the above-mentioned skeletal information vector, the above-mentioned music attribute feature, and the above-mentioned music sentiment feature to obtain a spliced vector.
[0085] And, in Figure 5 the above-mentioned at least one encoder includes a Transformer encoder and a Variational Autoencoder (VAE), etc. The spliced vector is encoded by the above-mentioned encoders respectively to obtain the encoding results of the corresponding encoders, so as to obtain the fused vector.
[0086] In some embodiments, the at least one encoder includes a first encoder. Encoding the splicing vector based on the at least one encoder to obtain an encoding result may include: through the first encoder, performing attention processing on the splicing vector based on an attention mechanism to obtain a first encoding result, where the attention processing encodes the splicing vector based on the attention feature information of the splicing vector, and the attention feature information is used to characterize the correlation degree between different information segments in the splicing vector.
[0087] Specifically, the first encoder may include a Transformer encoder. Inputting the splicing vector into the Transformer encoder, the Transformer encoder, by means of its internal self-attention mechanism, comprehensively and deeply analyzes each element in the splicing vector and their mutual relationships, so as to be able to capture the long-range dependencies and complex context information in the input data to obtain the attention feature information about the splicing vector. For example, it can accurately discover the internal relationship between the music rhythm and the movement rhythm, the corresponding connection between the music emotion and the movement style, and the correlation between the movement state reflected by the bone information and the movement smoothness, etc., and then form a profound understanding of aspects such as music, emotion, and bone movement. After encoding processing, it outputs a feature representation containing the internal correlation degree, and the Transformer encoding feature can be used as the first encoding result.
[0088] In some embodiments, the at least one encoder includes a second encoder. Encoding the splicing vector based on the at least one encoder to obtain an encoding result includes: through the second encoder, performing compression processing on the splicing vector to obtain a second encoding result, where the compression processing encodes the splicing vector based on the spatial distribution feature of the splicing vector.
[0089] Specifically, the second encoder may include a VAE encoder. Inputting the splicing vector into the VAE encoder, the VAE encoder performs a compression operation on the splicing vector. Through a complex neural network mapping and learning process, the splicing vector is transformed into a low-dimensional vector. During the compression operation, the VAE encoder can learn the internal structure and distribution law in the input information, so that the generated low-dimensional vector can be used to indicate the distribution feature of the splicing vector in the latent space, to describe the complex splicing vector in a concise and informative representation manner, to show the positioning and discreteness of the input data in the latent space, and at the same time, it can also combine the mapping law of the movement features contained in the bone information vector in the latent space to prepare for the subsequent generated object actions.
[0090] In some embodiments, the step of obtaining the fused vector based on the above encoding results includes: concatenating the above first encoding result and the second encoding result in vectors, so as to connect the vectors corresponding to the two encoding results in sequence to form a new fused vector. Thus, when generating an object action based on the fused vector, it is possible to utilize both the feature representation with intrinsic correlation in the Transformer encoding features and the latent space distribution information determined by the low-dimensional vector output by the VAE encoder. Moreover, it also takes into account the important influence of the bone vector information in the process of generating the object action, laying a solid foundation for generating high-quality, diverse object actions that conform to the movement laws and the characteristics of music emotion and rhythm.
[0091] Step 104: Obtain the action sequence of the virtual object based on the above fused vector through at least one decoder.
[0092] In this embodiment, the terminal can perform processing on the generation of the action of the virtual object based on the above bone information vector and music feature vector to obtain the action sequence of the virtual object. The action sequence contains the pose information of the bones of the virtual object at multiple moments. Based on the pose information of the bones, the virtual object is driven to perform corresponding actions, avoiding the time-consuming and laborious manual creation of the object actions of the virtual object, greatly improving the generation efficiency and quality of the object actions, reducing the manual operation time, being suitable for mass production, and ensuring that the dance actions highly match the in-game music. At the same time, it can also be naturally and smoothly presented under different character bone structures, enhancing the expressiveness and realism of the game character in the game scene. Through the high-quality action sequence, a more rich and personalized experience is provided for the game.
[0093] In some embodiments, the above at least one decoder includes a first decoder. The step of obtaining the action sequence of the virtual object based on the above fused vector through at least one decoder includes: determining the pose information to be generated at the current moment based on the above fused vector and the pose information generated before the current moment through the above first decoder, so as to obtain the action sequence of the virtual object.
[0094] Among them, the above first decoder can be structurally matched with the above first encoder.
[0095] In some embodiments, the method of determining the pose information to be generated at the current moment based on the fused vector and the pose information generated before the current moment by the first decoder to obtain the action sequence of the virtual object may include: determining, by the first decoder, based on the fused vector and the pose information generated before the current moment, the target motion parameters required from the pose information before the current moment to the pose information at the current moment, and then determining, based on the target motion parameters and the information representation of the pose information at the current moment, the pose information to be generated at the current moment to obtain the action sequence of the virtual object.
[0096] Specifically, the first decoder may be a Transformer decoder. By inputting the fused vector into the Transformer decoder, the Transformer decoder can gradually generate the pose information corresponding to each moment in the action sequence according to the input fused vector and the self-attention mechanism and / or cross-attention mechanism configured in the Transformer decoder in complex situations involving multi-modal information interaction. Each moment's pose information contains a plurality of poses for indicating different parts and joints of the skeleton.
[0097] Among them, when generating the object action at each moment, the first decoder can dynamically infer what kind of object action should be generated at the next moment based on the music attribute features, skeleton motion features, music emotion features corresponding to the current moment in the fused vector, and the pose information of the skeleton that has been generated previously. By continuously repeating this process, the action sequence can be generated.
[0098] Among them, the motion parameters required from the pose information before the moment to the pose information at the moment, as described in the above embodiments for the motion parameters. For example, the motion parameter is a joint angle change parameter. For example, when the action changes, the angle of the shoulder joint changes from 0° to 30° when the arm swings, and the flexion and extension angle changes of the knee joint when kicking, which are used to reflect the action amplitude and posture change; Another example is that the motion parameter is a translation parameter. For example, when sliding, the displacement of the hip joint in the horizontal direction, and the change of the position of the end of the hand bone during jumping and rotating, which depict the spatial trajectory of the object action; Another example is that the motion parameter is a speed parameter. For example, for the action speed and acceleration, the number of swings of the wrist per second when waving the hand quickly reflects the speed, and the angle change acceleration at the starting stage of the arm extension, which enriches the action expressiveness and rhythm change.
[0099] It should be noted that the self-attention mechanism of the Transformer decoder can focus on the importance degrees of various input features according to the actual situation. For example, at a certain moment, it may focus more on the rhythm features of music and needs to generate matching object actions in combination with the current movement trend of the skeleton. At another moment, it will adjust the style of the action and the movement amplitude of the bone joints according to the emotional features, so as to ensure that the generated object actions not only match the music in rhythm, but also perfectly match in emotional expression, and at the same time meet the requirements of the rationality and smoothness of the movement, resulting in the final action sequence being more in line with the expected requirements.
[0100] In some embodiments, determining the target motion parameters required between the pose information before the current moment and the pose information at the current moment by the first decoder based on the fused vector and the pose information generated before the current moment includes: determining the target motion parameters required between the pose information before the current moment and the pose information at the current moment by the first decoder based on the fused vector, the weight indication information corresponding to the fused vector, and the pose information generated before the current moment, where the weight indication information is the attention weights generated for each feature vector in the fused vector by the attention mechanism corresponding to the first decoder.
[0101] Among them, the attention weights corresponding to each feature vector (such as music attribute feature vector, music emotion feature vector) in the fused vector corresponding to each moment can obtain the corresponding attention scores by calculating the feature correlation, and this attention score is used as the attention weight. Then, actions are generated using the feature vectors with weights assigned.
[0102] Exemplarily, at each moment, the Transformer decoder calculates the correlation between each input feature and the object action part to be generated currently. This is achieved by calculating the similarity between the feature vector and the action generation target vector (a vector representing the object action features expected to be generated). For example, if a fast footstep action is to be generated currently, the correlation between the music rhythm feature and this action may be relatively high, and the excitement level feature of the emotion may also have a certain correlation. Then, the attention scores are calculated based on the correlation. Using the dot product attention mechanism, the dot product operation is performed between the feature vector and the action generation target vector, and then normalized by the softmax function to obtain the attention scores of each feature. For example, after calculation, the attention score of the music rhythm feature is 0.8, and the emotion feature is 0.2. Finally, according to the updated weights, the input features are weighted and summed to obtain a comprehensive feature representation. For example, the music rhythm feature vector is multiplied by its weight 0.8, the emotion feature vector is multiplied by 0.2, and then these results are added to obtain a comprehensive feature vector. The comprehensive feature vector is input into the generation layer of the Transformer decoder to generate the corresponding part of the object action. For example, this comprehensive feature vector can help the decoder determine parameters such as the speed and amplitude of the current footstep action, so as to generate an object action that conforms to the current emphasized feature (such as music rhythm).
[0103] Moreover, when generating the object action at the next moment, the above steps are repeated. According to the object action requirements and the changes of the input features at different time steps, the feature weights are dynamically adjusted. For example, if a slow arm action is to be generated at the next time step, the soothing level feature of the emotion may obtain a higher attention score, so as to get more emphasis after the weight update, making the generated arm action conform to the soothing emotion and other relevant input features.
[0104] Exemplarily, first, the self-attention mechanism is used to analyze the previous action sequence to determine the influence weights of each part of the action on the current step. For example, if the previous arm movement is closely related to the strong beat of the music and the current music is about to enter the strong beat, the weight related to the arm movement will be relatively high. At the same time, according to the generated probability distribution, consider the probabilities of different joint actions under the current music emotion and action coherence. For example, the probability of knee joint bending is high in lively music, and this probability is adjusted in combination with the previous action state. Then, combined with multi-modal input, fuse music features, emotions, and action styles, etc. Under strong rhythm electronic music, excited emotion, and hip-hop style, consider the drumbeat rhythm, excitement level, and hip-hop action elements. Next, weights (such as 0.4, 0.3, 0.3) are set for the fusion of the self-attention mechanism, the generated probability distribution, and multi-modal input respectively, and the weighted sum of each inferred action element such as joint angle change and position movement is calculated. For example, for the shoulder joint angle change, the self-attention mechanism infers 30°, the probability distribution infers 20°, and the multi-modal inference is 25°. The final angle change is 0.4×30° + 0.3×20° + 0.3×25° = 25.5°. After that, the weighted sum result is used as the preliminary action and substituted into the three inference methods for iterative optimization. The self-attention mechanism checks the coherence. If the preliminary action is abrupt, the weights are adjusted and re-inferred; the probability distribution inference re-evaluates the probability based on the preliminary action and corrects the unreasonable action probability; the multi-modal input fusion adjusts the action time, amplitude, etc. according to the matching degree of the preliminary action with the music and style. Through continuous iteration, until the object action at the next moment that meets the requirements is generated.
[0105] In some embodiments, at least one of the above decoders includes a second decoder. The obtaining of the action sequence of the virtual object by the at least one decoder based on the fused vector may include: adjusting the action sequence generated by the first decoder by the second decoder based on the spatial distribution features corresponding to the fused vector to obtain an adjusted action sequence.
[0106] Wherein, the structure of the second decoder matches that of the second encoder.
[0107] In this embodiment, the second decoder may be a VAE decoder. After obtaining the action sequence of the virtual object generated by the first decoder, the action sequence may be input into the VAE decoder, or the processed action sequence may be input into the VAE decoder by performing appropriate feature extraction or encoding operations on the action sequence, so that the VAE decoder can further refine, improve, and adjust the style and amplitude of the input action sequence by using the knowledge of the latent space distribution learned previously, and at the same time, fully consider the motion characteristics reflected by the skeletal information, so that the object motion is more natural and smooth and fits the actual motion situation. For example, it can make the object motion more in line with the skeletal kinematics principle according to the distribution law of the latent space, or change the style of the object motion by adjusting the relevant parameters in the encoding vector to make it more dynamic or more soothing, so as to meet different specific needs, and finally output a complete object motion sequence after optimization and supplementation.
[0108] The encoding vector is an abstract representation of object-oriented motion information. It contains the characteristics of the object motion, such as the amplitude and speed of the motion. This information is stored in the encoding vector in the form of numerical values, and different numerical combinations represent different motion styles. For example, a vigorous street dance style has higher motion speed and amplitude-related parameters in the encoding vector, while an elegant classical dance style has lower speed and softer motion amplitude-related parameters.
[0109] Exemplarily, in a scenario where adjustments are made based on amplitude-related parameters, when changing from a soothing classical dance style to an unrestrained street dance style, the parameters representing the amplitude of the movement in the encoding vector are adjusted. For example, the parameter corresponding to the arm swing amplitude in classical dance may be 0.3. When converted to street dance style, this parameter is increased to 0.7. Related parameters such as the body rotation amplitude are also increased and adjusted accordingly. Accordingly, the parameter adjustment strategy can be to use a linear function to adjust according to the degree of difference between the target style and the original style. If the target style movement amplitude is twice that of the original style, the amplitude parameter is multiplied by 2 using a linear function.
[0110] It is understandable that when the amplitude parameters are adjusted in the encoding vector, the object motion generation system will recalculate the limb motion range according to the new amplitude parameters. For example, after the arm motion amplitude parameters are increased, the maximum arm swing angle increases from the original 45° to 90° within the allowable range of human bones and muscles. At the same time, the associated shoulder shrug amplitude, body tilt amplitude, etc. will also be updated accordingly according to the coordination relationship to ensure natural and integrated movements.
[0111] Exemplarily, in a scenario where adjustments are made based on speed-related parameter adjustments, for the transition from a slow ethnic dance to a fast-paced jazz dance, the action speed parameter is found in the encoding vector. Suppose the action speed parameter of the ethnic dance is 0.2, and when transitioning to the jazz dance style, it is adjusted to 0.6. Correspondingly, the parameter adjustment strategy can be to use a scaling method and adjust according to the speed ratio relationship between the target style and the original style. For example, if the average action speed of the target style is three times that of the original style, the speed-related parameter is multiplied by 3.
[0112] It can be understood that after the speed-related parameter is adjusted, the object action generation system will re-plan the displacement of the limb actions within each time step. For example, for the foot actions, after the speed parameter increases, the original movement of 5 centimeters per time step may now increase to 10 centimeters. At the same time, the system will also consider the changes in acceleration and deceleration to simulate the gradual change process of speed in real object actions. For example, when transitioning from a slow to a fast style, the action acceleration is gradually increased to make the speed change more natural.
[0113] Specifically, the distribution law of the above-mentioned latent space includes the Gaussian distribution law, that is, in the latent space, the joint angles approximately follow the Gaussian distribution law. For example, for the elbow joint angle, assume that its mean value is a reasonable intermediate value within the normal action range, such as 90°, and the standard deviation is 15°. This indicates that in reasonable object actions, the elbow joint angle will mostly be concentrated within the range of the mean value plus or minus several standard deviations, that is, the probability within the range of 60° to 120° is relatively high. Then, when the VAE decoder finds that the elbow joint angle in the generated object action does not conform to this Gaussian distribution law, for example, the angle is 150°, it can be adjusted according to the distribution law, that is, by calculating the deviation degree of the current angle from the reasonable distribution range and linearly mapping the angle back to the reasonable range. For example, according to the distance between the current angle and the mean value and the standard deviation, a adjustment coefficient is calculated to adjust the 150° angle to a value close to the reasonable range, such as 105°. At the same time, during the adjustment process, the coordination of adjacent joint angles will also be considered to ensure the coherence of the entire limb action.
[0114] Specifically, the distribution law of the above-mentioned latent space includes the distribution law of bone position movement, that is, in the latent space, there is a certain law for bone position movement. For example, the vector length of bone position movement within a unit time approximately conforms to the Gaussian distribution law. Taking the foot position movement as an example, assuming that under the normal movement rhythm, the mean value of the position movement vector length of the foot within one time step is 10 cm and the standard deviation is 3 cm. This means that under normal circumstances, the distance of foot position movement will mostly be in the range of 7 - 13 cm. Then, when the VAE decoder detects that the generated foot position movement does not conform to this distribution law, for example, the movement distance reaches 30 cm, a reasonable adjustment target value can be determined according to the properties of the Gaussian distribution law to approach the mean value, such as adjusting to 12 cm. Then, for the direction of position movement, it will also be corrected according to the overall style and coherence of the object's movement. If the movement direction deviates too much from the normal movement direction, it will be adjusted to a more reasonable direction, and through a smooth transition method, it is ensured that the trajectory of bone position movement is continuous and conforms to the principles of bone movement mechanics.
[0115] In some embodiments, after obtaining the above-mentioned action sequence of the virtual object, it may further include: sending the action sequence of the virtual object to the target user to display the action sequence to the target user. Then, receiving the action feedback information of the target user on the action sequence of the virtual object, where the action feedback information includes but is not limited to the user's preference degree for the object's action style, the feeling of the match between the action and the music, the overall visual effect evaluation, and whether it conforms to the natural movement state, etc. Finally, the terminal can convert the above-mentioned action feedback information into a corresponding reward mechanism based on the reinforcement learning model, and update the action sequence of the virtual object based on the above-mentioned reward mechanism.
[0116] Among them, the above-mentioned reinforcement learning model can be constructed based on a reinforcement learning algorithm (for example, the Proximal Policy Optimization (PPO) algorithm).
[0117] Exemplarily, converting the above-mentioned action feedback information into a corresponding reward mechanism can be: if the user is satisfied with the object's action, a positive reward is given; if not satisfied, a negative reward is given. The PPO algorithm explores and learns through the feedback information, adjusts and optimizes the policy for generating the object's action sequence according to the reward mechanism, so as to better adapt to the diverse needs of different users, continuously improve the quality of the generated object's action and user satisfaction, and at the same time ensure that the generated object's action always conforms to the scientific laws of movement.
[0118] Exemplarily, since the action feedback information can be that the user expresses their preference for the action style of the object, such as liking a more lively, more elegant, more powerful or more ethnic style, etc. It may also point out elements in the current object action that do not meet the expected style, such as the action being too complex or simple, or not conforming to the typical action postures of a specific object style. Therefore, optimization can be carried out based on the action feedback information. That is, if the user hopes that the object action style is more lively, the PPO algorithm will adjust the object action generation strategy, so that in the initial stage of action generation by the Transformer decoder, more attention is paid to the fast-paced parts of the music, and the generation of action elements related to the lively style is enhanced. For example, actions such as jumping and rapid rotation are added. In the optimization and supplementation stage of the VAE decoder, the style-related parameters in the latent space will be adjusted to make the generated object actions more in line with the characteristics of the lively style in terms of action amplitude, speed and posture combination.
[0119] Exemplarily, since the action feedback information can be that the user will feedback whether the rhythm of the object action matches the music, and whether the emotional expression of the action matches the emotional atmosphere of the music, etc. For example, the user may feel that the object action is not exciting enough in the climax part of the music, or the action is too intense in the slow music passage. Therefore, optimization can be carried out based on the action feedback information. That is, when the user feedbacks that the rhythm of the object action does not match the music, the PPO algorithm will adjust the strategy. In the action generation stage of the Transformer decoder, the connection between the music rhythm feature vector and the object action rhythm elements will be strengthened. For example, by adjusting the weights in the self-attention mechanism, the decoder will refer more to the beat features of the music when generating actions. In the VAE decoder stage, according to the speed of the music rhythm, the speed and pauses of the action are optimized, so that the object action can accurately follow the rhythm changes of the music, generating more compact and fast actions in the fast-paced music part and more soothing and smooth actions in the slow-paced part.
[0120] In some embodiments, after obtaining the action sequence of the above virtual object, the terminal can use the action file output module to encode the action sequence of the above virtual object in a preset format through a preset software toolkit in the action file output module, and obtain an action file that conforms to the above preset format.
[0121] Specifically, the action sequence can be encoded and output according to the specifications of the FBX format, ensuring that the generated action sequence can be saved in the FBX file format and is convenient for subsequent use in the game development environment. Among them, relevant functions provided in the FBX software development kit (SDK) can be used to complete this task.
[0122] Among them, the above FBX is a common 3D model file format used to save game object actions for convenient use in game development.
[0123] It can be understood that the above action generation method enables users to produce an action file in a specific format for the object actions of game characters that meet the user's needs through a simple operation process, thereby reducing the usage threshold, having good adaptability and scalability, and being able to meet the diverse needs of different users and different application scenarios.
[0124] In some embodiments, in order to implement the action generation method, the encoder and decoder need to be trained in advance.
[0125] First, in the data preparation stage, a large number of game music files, skeletal model files of game characters corresponding to the game music files, and corresponding matching game action sequences can be collected as training data. The training data needs to include various different game music styles, different character skeletal types, and different game object action styles. Then, in order to improve the accuracy of the training data, the training data can be preprocessed. For example, feature extraction can be performed on the collected game music files. For example, music attribute feature vectors can be extracted through the MFCC algorithm and the STFT algorithm. Then, based on the music attribute feature vectors and music data, music emotion feature vectors can be obtained. And the skeletal files can be processed to parse out information such as the structure of the skeleton, joint positions, and angles, and thus be converted into skeletal information vectors. Finally, for the game action sequence, it can be decomposed into pose information such as the positions and motion states of each joint at different time points.
[0126] Secondly, in the system training stage, the two core components, Transformer and VAE, can be trained separately.
[0127] In some embodiments, in the Transformer training, the data generated in the above data preparation stage can be first labeled as a training set. These labeled data clarify the matching relationship between the music and the corresponding object actions in terms of rhythm, emotional expression, style, etc. Then, it is processed into the form of the above-mentioned music attribute feature vectors, skeletal information vectors, and music emotion feature vectors, and combined into corresponding training sample pairs (the sample is the fused vector of the music attribute feature vector, skeletal information vector, and music emotion feature vector, and the label is the corresponding standard game action sequence).
[0128] Specifically, the above-mentioned annotation method for the rhythm can be: dividing the music data by beats, annotating the start time and intensity (for example, the strong beat and weak beat time points in each measure in 4 / 4 time), and decomposing the object action into units and annotating the start time and duration. The corresponding matching relationship is manifested as: the strong beat of the music is synchronized with the heavy beat of the action, and the fast rhythm of the music corresponds to the object action sequence to achieve the matching in terms of rhythm.
[0129] Specifically, the annotation method corresponding to the above emotional expression can be: determining the emotional category of music data according to melody, harmony, rhythm, etc. using an emotion analysis model (for example, lively, sad, etc.), and annotating the corresponding emotional category for the object actions in terms of posture, amplitude, speed, etc. (for example, a large jumping action corresponds to an exciting emotional category, and a gentle and slow action corresponds to a soothing emotional category). The corresponding matching relationship is manifested as: lively music matches the jumping and lively object actions, and sad music matches the slow and introverted object actions to reflect the matching in terms of emotional expression.
[0130] Specifically, the annotation method corresponding to the above style can be: dividing the music data according to the type (for example, classical, pop, etc.) and elements (for example, musical instruments, melody characteristics), and dividing the object actions according to the type (for example, ballet, hip-hop, etc.) and action characteristics (for example, standing on tiptoe, rhythm, etc.). The corresponding matching relationship is manifested as: classical music matches ballet actions, and hip-hop music matches typical hip-hop actions to present the matching in terms of style.
[0131] Then, the definition of the loss function is carried out, that is, the cross-entropy loss function can be adopted to measure the degree of difference between the object action sequence generated by the Transformer decoder and the standard game action sequence. For example, the cross-entropy loss will calculate the difference between the probability distribution of each generated action element and the real action element, and then accumulate to obtain the loss value of the entire sequence.
[0132] Specifically, in the scenario where Transformer is trained for object action generation, for each action element (for example, the specific action angle, action direction, etc. of a certain limb joint), there will be two probability distributions, one is the prediction distribution and the other is the target distribution.
[0133] The generated probability distribution (prediction distribution): Output by the Transformer model, which gives the probability situation corresponding to the value of each possible action element based on the model's processing of the input information (such as music features, skeletal information, etc.). For example, for an action element of a certain knee bending angle, the probabilities corresponding to the possible value angles of 30°, 45°, and 60° output by the model are [0.2, 0.5, 0.3], that is, the model believes that the probability of the knee bending angle being 45° is 0.5.
[0134] The probability distribution of the real action element (target distribution): Derived from the real action situation pre-annotated in the training dataset, which is a deterministic distribution situation. For example, for an action element of a certain knee bending angle, if the real knee bending angle is 45°, then its probability distribution is [0, 1, 0], that is, the probability at the position corresponding to the real value is 1, and the probabilities at other value positions are 0.
[0135] Furthermore, the cross-entropy loss function is used to measure the difference between these two probability distributions, and its calculation formula is:
[0136]
[0137] Where P represents the true probability distribution (target distribution), P = {p1, p2, ···, pn}; Q represents the generated probability distribution (predicted distribution), Q = {q1, q2, ···, qn}; n represents the number of all possible values of the action element, and i represents different value indices; log is the logarithmic function with the natural constant e as the base.
[0138] Exemplarily, assume that there are 4 possible values for a certain object action element (such as the wrist rotation angle), which are angles A1, A2, A3, and A4 respectively.
[0139] The generated probability distribution Q: The probability distribution output by the model after calculation is Q = {0.1, 0.3, 0.4, 0.2}, which means that the model believes that the probability of the wrist rotation angle being A3 is the highest, with a probability of 0.4.
[0140] The true probability distribution P: If it is known from the labeled data that the true wrist rotation angle is A2, then its true probability distribution is P = {0, 1, 0, 0}, and only the probability at the position corresponding to the true angle A2 is 1, and the others are 0.
[0141] Calculate the difference between the two according to the cross-entropy calculation formula:
[0142] H(P, Q) = -(0log(0.1) + 1log(0.3) + 0log(0.4) + 0log(0.2)) = -log(0.3)
[0143] This calculation result is the cross-entropy loss value, which quantifies the degree of difference between the generated probability distribution and the true probability distribution. When training the Transformer, the above cross-entropy loss calculation will be performed for each action element, and then the loss situations of all action elements will be integrated (for example, by averaging, etc.) to obtain the total loss of the generation of the entire game action sequence. The model will adjust its own parameters according to this total loss through the backpropagation algorithm, so that the generated probability distribution can continuously approach the true probability distribution, thereby improving the accuracy and quality of the generated object actions, and making the generated object actions more in line with the true expected situation.
[0144] Finally, model training is carried out. In each training iteration, the input vectors of the training sample pairs are fed into the encoder of the Transformer. After encoding, they are passed to the decoder to generate the object action sequence. Then, the loss value is calculated according to the defined loss function, and the loss value is backpropagated to the parameters of each layer of the Transformer through the backpropagation algorithm. The parameters are updated according to a certain optimization algorithm (such as Adam, SGD, etc.), and this process is continuously repeated until the loss value of the Transformer on the training set converges or reaches the preset number of training epochs, enabling the Transformer to learn the mapping relationship between music, skeletal information, and emotions and object actions, and having the ability to generate a reasonable action sequence according to the input features.
[0145] In some embodiments, in VAE training, the loss function can be defined first: The loss function of VAE consists of two parts. One part is the reconstruction loss, which measures the difference between the object action sequence generated by the VAE decoder and the real object action sequence. The mean squared error (MSE) is used to calculate the sum of the squares of the differences between each element of the generated action sequence and the corresponding element of the real sequence. The other part is the KL divergence (Kullback-Leibler Divergence), which measures the difference between the posterior distribution (determined by the mean and variance) of the latent variables output by the VAE encoder and the preset prior distribution, so as to constrain the VAE to learn that the latent space distribution conforms to the expectation. The total loss function is the weighted sum of these two parts of losses, and the relationship between the reconstruction accuracy and the rationality of the latent space distribution can be balanced by adjusting the weights.
[0146] Among them, the presentation of the above elements can be represented by the angles and positions of joints, etc. For example, for an action sequence, at a specific moment, for the upper limb part of the human body, the action elements can be represented by vectors, such as [shoulder joint angle, elbow joint angle, wrist joint angle, finger joint angle, shoulder joint x coordinate, shoulder joint y coordinate, shoulder joint z coordinate, elbow joint x coordinate,...], where each value in the vector is an action element, which represents the posture and position of the limb in three-dimensional space. For the generated action sequence and the real action sequence, the action elements can be represented by the above vectors at each moment.
[0147] Exemplarily, assume there is a simple action sequence, only considering two moments (t1 and t2), and only considering two action elements, namely the shoulder joint angle and the elbow joint angle, at each moment. Then, the real object action sequence is represented as:
[0148] At t1 moment: A(t1,real) = [30°, 60°] (indicating that the shoulder joint angle is 30° and the elbow joint angle is 60°);
[0149] At time t2: A(t2,real) = [40°, 70°];
[0150] The generated object action sequence is represented as:
[0151] At time t1: A(t1,gen) = [25°, 55°];
[0152] At time t2: A(t2,gen) = [38°, 68°].
[0153] Then, calculate the mean squared error (MSE): First, calculate the square of the difference for each element at each time step, as follows:
[0154] At time t1:
[0155] For the square of the difference in shoulder joint angle: (30 - 25)^2 = 25;
[0156] For the square of the difference in elbow joint angle: (60 - 55)^2 = 25;
[0157] At time t2:
[0158] For the square of the difference in shoulder joint angle: (40 - 38)^2 = 4;
[0159] For the square of the difference in elbow joint angle: (70 - 68)^2 = 4.
[0160] Then, sum up the squares of all these differences: MSE = 1 / (2×2)×(25 + 25 + 4 + 4), where, since there are two time steps and two action elements in each time step, we need to divide by 2×2 to find the average), that is, the calculation result is MSE = 58 / 4 = 14.5.
[0161] It should be noted that by calculating the sum of squares, the overall difference between the generated object action sequence and the real object action sequence is measured. And by squaring the difference of each element, it can ensure that the difference is positive, and, the contribution of larger differences to the loss function is amplified, so that larger differences can be emphasized. For example, if we simply calculate the absolute value of the difference, then larger differences and smaller differences can cancel each other out when summing, while squaring can more prominently reflect larger differences, making the model pay more attention to the inaccurate action elements during the training process. And by squaring the difference of each element, it can ensure good mathematical properties. Since the square operation has good mathematical properties and is convenient for operations such as taking derivatives, when training the model based on optimization algorithms such as gradient descent, the gradient of the loss function with respect to the model parameters can be calculated conveniently, so as to effectively adjust the parameters and make the model optimize in the direction of reducing the loss, that is, the generated object action sequence is closer to the real sequence.
[0162] It is understandable that in VAE training, it is necessary to clarify how different the distribution of the latent variables output by the VAE encoder is from the preset distribution. The distribution of this latent variable is determined by its mean and variance, and the range can be used to indicate the interval where the latent variable is likely to appear.
[0163] Regarding the assumption of the distribution, usually the preset distribution is set to a relatively standard normal distribution. For example, the mean of the normal distribution is 0 and the variance is 1. The distribution of the latent variables output by the VAE encoder is also a normal distribution, and its mean and variance are output by the encoder itself and will change with different input data.
[0164] Among them, the method of measuring the difference between the posterior distribution (determined by the mean and variance) of the latent variables output by the VAE encoder and the preset prior distribution: the KL divergence can be used to measure the difference between these two distributions. For example, if the mean of the distribution of the latent variables output by the VAE encoder is 0.5 and the variance is 0.8, since it is different from the preset normal distribution with a mean of 0 and a variance of 1, in order to calculate the difference between them, a series of comparison steps can be used to clarify the difference in shape (i.e., variance) and position (i.e., mean) between this output distribution and the preset distribution. That is, finally, the corresponding difference result can be obtained, and this difference result can represent the degree of difference between the two distributions.
[0165] Then, the difference value (KL divergence) is regarded as a signal to be backpropagated to the VAE encoder to prompt the VAE encoder to clarify the difference value between the output latent variable distribution and the expected distribution. When this difference value is relatively large, the encoder will adjust its internal parameters to change the distribution of the output latent variables. Thus, through continuous training, the distribution of the latent variables output by the VAE encoder approaches the preset distribution, which promotes better constraint of the distribution of the latent variables in space, and the model will also be more reasonable when generating action sequences.
[0166] It should be noted that in the training iteration, the input vector is fed into the VAE encoder to obtain the mean and variance. After generating the latent variable through the reparameterization trick, the decoder generates the object action sequence. Then, the reconstruction loss and KL divergence are calculated respectively, and the loss value is calculated according to the total loss function. The loss is backpropagated to the parameters of each layer of the VAE using the backpropagation algorithm, and the parameters are updated according to the optimization algorithm. The training continues until the loss converges or reaches the preset number of training epochs, so that the VAE can learn a reasonable latent space representation of the input data, accurately reconstruct the object action sequence according to the latent variable, and ensure the good properties of the latent space distribution. During the training process, by adjusting the parameters of the encoder and decoder, the VAE can accurately learn the latent distribution law of the input information, thereby generating more diverse and natural game object actions.
[0167] It can be understood that since the encoder of the VAE is a neural network. When the input vector (from a sample in the training set, such as the music feature vector corresponding to the object action sequence) enters the encoder, this neural network will perform a series of linear and non-linear transformations on the input vector. At the last layer of the encoder network, there are two output branches. One branch outputs the mean vector, and the other branch outputs the variance vector. These outputs are obtained by performing the forward propagation calculation of the neural network on the input vector. For example, for a simple fully connected neural network encoder, after the input vector is processed by the weighted sum activation function of multiple hidden layers, the mean and variance are calculated through a specific weight matrix and bias term at the last layer.
[0168] Moreover, after obtaining the mean and variance, the latent variable can be generated by the reparameterization method. Because directly sampling from the mean and variance will cause the gradient to not be backpropagated (because the sampling operation is non-differentiable). The reparameterization trick is to sample a noise vector from a standard normal distribution, and then generate the latent variable through the formula (latent variable = mean + square root of variance x noise vector), so that during the training process, the gradient can be backpropagated from the mean and variance to the encoder.
[0169] Among them, after generating the latent variable, it can be input into the decoder of the VAE, and the decoder will generate the reconstructed output (such as the reconstructed object action sequence) according to the latent variable. Then, the loss function is calculated. The loss function usually includes two parts: the reconstruction loss (such as using the mean square error to measure the difference between the reconstructed output and the original input) and the KL divergence loss (used to measure the difference between the posterior distribution of the latent variable determined by the mean and variance and the preset prior distribution). The VAE is trained by minimizing this loss function, so that the encoder can better extract the distribution characteristics of the latent variable, and the decoder can better reconstruct the input data according to the latent variable.
[0170] As can be seen from the above, by obtaining the skeleton file of the virtual object, processing the skeleton information of the above skeleton file, the skeleton information vector of the above virtual object is obtained; obtaining the music file, extracting the features of the above music file, the music feature vector of the above music file is obtained; encoding the above skeleton information vector and the above music feature vector through at least one encoder to obtain a fused vector; obtaining the action sequence of the above virtual object through at least one decoder based on the above fused vector, so as to automatically generate the action sequence of the virtual object through the skeleton information vector of the skeleton and the music feature vector of the music, thereby improving the action generation efficiency of the virtual object.
[0171] This embodiment also provides an action generation device, which can be specifically integrated in a terminal device. For example, as Figure 6 shown, the action generation device may include:
[0172] An information processing module 601, configured to obtain the skeleton file of the virtual object, process the skeleton information of the above skeleton file, and obtain the skeleton information vector of the above virtual object;
[0173] A feature extraction module 602, configured to obtain the music file, extract the features of the above music file, and obtain the music feature vector of the above music file;
[0174] An encoding module 603, configured to encode the above skeleton information vector and the above music feature vector through at least one encoder to obtain a fused vector;
[0175] An action generation module 604, configured to obtain the action sequence of the above virtual object through at least one decoder based on the above fused vector.
[0176] In some embodiments, the above skeleton file includes the initial skeleton pose information of the skeleton of the above virtual object at the initial moment, the above action sequence includes the pose information of the skeleton of the above virtual object at multiple moments, and the information processing module 601 is specifically configured to:
[0177] Through the motion feature association of the pose information of the skeleton before and after motion in the first kinematic equation, based on the above initial skeleton pose information, calculate the information representation of the predicted pose information of the above skeleton at multiple moments, where the information representation of the predicted pose information at each moment is obtained based on the information representation of the predicted pose information before the corresponding moment;
[0178] Based on the above initial skeleton pose information and the information representation of the predicted pose information at the above multiple moments, obtain the skeleton information vector of the above virtual object.
[0179] In some embodiments, the above-mentioned motion feature association is based on the representation of motion parameters. Specifically, the information processing module 601 is configured to:
[0180] Based on the association relationship between the motion parameters provided by the first kinematic equation and the pose information of the bone before and after the motion, calculate the information representation of the predicted pose information of the bone at multiple moments on the basis of the above-mentioned initial bone pose information;
[0181] The information representation of the predicted pose information at each moment is obtained based on the information representation of the predicted pose information before the corresponding moment, including:
[0182] The information representation of the predicted pose information at each moment is obtained based on the motion parameters at the corresponding moment and the information representation of the predicted pose information before the corresponding moment.
[0183] In some embodiments, the above-mentioned at least one decoder includes a first decoder. Specifically, the action generation module 604 is configured to:
[0184] Based on the above-mentioned fused vector and the pose information generated before the current moment through the above-mentioned first decoder, determine the pose information to be generated at the current moment, so as to obtain the action sequence of the virtual object.
[0185] In some embodiments, the action generation module 604 is specifically configured to:
[0186] Based on the above-mentioned fused vector and the pose information generated before the current moment through the above-mentioned first decoder, determine the target motion parameters required between the pose information before the current moment and the pose information at the current moment;
[0187] Based on the above-mentioned target motion parameters and the information representation of the pose information at the current moment, determine the pose information to be generated at the current moment, so as to obtain the action sequence of the virtual object.
[0188] In some embodiments, the action generation module 604 is specifically configured to:
[0189] Based on the above-mentioned fused vector, the weight indication information corresponding to the fused vector, and the pose information generated before the current moment through the above-mentioned first decoder, determine the target motion parameters required between the pose information before the current moment and the pose information at the current moment;
[0190] The above-mentioned weight indication information is the attention weight generated for each feature vector in the above-mentioned fused vector based on the attention mechanism corresponding to the above-mentioned first decoder.
[0191] In some embodiments, the at least one decoder includes a second decoder, and the action generation module 604 is specifically configured to:
[0192] Through the second decoder, based on the spatial distribution characteristics corresponding to the fused vector, adjust the action sequence generated by the first decoder to obtain an adjusted action sequence.
[0193] In some embodiments, the action generation device further includes an adjustment module, and the adjustment module is specifically configured to:
[0194] Calculate the pose adjustment information of the to-be-predicted pose information of the bone at multiple moments through the second kinematic equation;
[0195] The information processing module 601 is specifically configured to:
[0196] Based on the initial bone pose information, as well as the information representation and pose adjustment information corresponding to the to-be-predicted pose information at multiple moments, obtain the bone information vector of the bone.
[0197] In some embodiments, the music feature vector includes a music attribute feature vector and a music emotion feature vector, and the feature extraction module 602 is specifically configured to:
[0198] Extract the attribute features of the music file to obtain the music attribute feature vector of the music file;
[0199] Based on the music file and / or the music attribute feature vector, perform emotion feature extraction to obtain the music emotion feature vector of the music file.
[0200] In some embodiments, the feature extraction module 602 is specifically configured to:
[0201] Extract the spectral features of the music signal in the music file to obtain the music spectral feature vector of the music file;
[0202] Perform frequency domain conversion on the music signal to obtain the converted frequency domain signal;
[0203] Extract the frequency domain features of the frequency domain signal to obtain the music frequency domain feature vector of the music file;
[0204] Based on the music spectral feature vector and the music frequency domain feature vector, obtain the music attribute feature vector of the music file.
[0205] In some embodiments, the feature extraction module 602 is specifically configured to:
[0206] Perform emotion recognition based on the music attribute feature vector to obtain the initial emotion feature vector of the music file;
[0207] Extract text from the above music file to obtain music text data;
[0208] Perform semantic analysis on the above music text data to obtain semantic feature vectors;
[0209] Based on the above semantic feature vectors and the above initial emotion feature vectors, obtain the music emotion feature vectors of the above music file.
[0210] In some embodiments, the above encoding module 603 is specifically configured to:
[0211] Concatenate the above bone information vector and the above music feature vector to obtain a concatenated vector;
[0212] Encode the above concatenated vector based on at least one encoder to obtain an encoding result;
[0213] Based on the above encoding result, obtain a fused vector.
[0214] In some embodiments, the above at least one encoder includes a first encoder, and the above encoding module 603 is specifically configured to:
[0215] Perform attention processing on the above concatenated vector through the above first encoder based on the attention mechanism to obtain a first encoding result;
[0216] The above attention processing is to encode the above concatenated vector based on the attention feature information of the above concatenated vector, and the above attention feature information is used to represent the correlation degree between different information segments in the above concatenated vector.
[0217] In some embodiments, the above at least one encoder includes a second encoder, and the above encoding module 603 is specifically configured to:
[0218] Perform compression processing on the above concatenated vector through the above second encoder to obtain a second encoding result;
[0219] The above compression processing is to encode the above concatenated vector based on the spatial distribution feature of the above concatenated vector.
[0220] In some embodiments, the above action generation device further includes an action feedback module, and the above action feedback module is specifically configured to:
[0221] Send the action sequence of the above virtual object to the target user;
[0222] Receive the action feedback information of the above target user on the action sequence of the above virtual object;
[0223] Based on the reinforcement learning model, the above action feedback information is converted into a corresponding reward mechanism, and the action sequence of the above virtual object is updated based on the above reward mechanism.
[0224] In some embodiments, the above action generation device further includes a format encoding module, and the format encoding module is specifically configured to:
[0225] Encode the action sequence of the above virtual object in a preset format through a preset software toolkit to obtain an action file that conforms to the above preset format.
[0226] It can be seen from the above that by obtaining the skeleton file of the virtual object, processing the skeleton information of the above skeleton file to obtain the skeleton information vector of the above virtual object; obtaining the music file, extracting the features of the above music file to obtain the music feature vector of the above music file; encoding and processing the above skeleton information vector and the above music feature vector through at least one encoder to obtain a fused vector; obtaining the action sequence of the above virtual object through at least one decoder based on the above fused vector, so as to automatically generate the action sequence of the virtual object through the skeleton information vector of the skeleton and the music feature vector of the music, so as to improve the action generation efficiency of the virtual object.
[0227] Correspondingly, an embodiment of the present application further provides an electronic device. The electronic device may be a terminal, and the terminal may be a smart phone, a tablet computer, a notebook computer, a touch screen, a game console, a personal computer (PC, Personal Computer), a personal digital assistant (Personal Digital Assistant, PDA), and other terminal devices. Or, the electronic device may be a server.
[0228] As Figure 7 shown, Figure 7 is a schematic structural diagram of the electronic device provided by an embodiment of the present application. The electronic device 700 includes a processor 701 with one or more processing cores, a memory 702 with one or more computer-readable storage media, and a computer program stored on the memory 702 and executable on the processor. Among them, the processor 701 is electrically connected to the memory 702. Those skilled in the art can understand that the structural diagram of the electronic device shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or arrange different components.
[0229] The processor 701 is the control center of the electronic device 700, connecting various parts of the entire electronic device 700 through various interfaces and circuits. By running or loading software programs and / or units stored in the memory 702, and calling the data stored in the memory 702, it executes various functions of the electronic device 700 and processes data, thereby monitoring the entire electronic device 700. The processor 701 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0230] In the embodiments of the present application, the processor 701 in the electronic device 700 will load the instructions corresponding to the processes of one or more application programs into the memory 702 according to the following steps, and the processor 701 will run the application programs stored in the memory 702 to implement various functions, such as:
[0231] Obtain the skeleton file of the virtual object, process the skeleton information of the above skeleton file, and obtain the skeleton information vector of the above virtual object;
[0232] Obtain the music file, extract the features of the above music file, and obtain the music feature vector of the above music file;
[0233] Encode and process the above skeleton information vector and the above music feature vector through at least one encoder to obtain a fused vector;
[0234] Obtain the action sequence of the above virtual object through at least one decoder based on the above fused vector.
[0235] Thus, the electronic device 700 provided in this embodiment can bring the following technical effects: improving the action generation efficiency of the virtual object.
[0236] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.
[0237] Optionally, as Figure 7 shown, the electronic device 700 further includes: a touch display screen 703, a radio frequency circuit 704, an audio circuit 705, an input unit 706, and a power supply 707. Among them, the processor 701 is electrically connected to the touch display screen 703, the radio frequency circuit 704, the audio circuit 705, the input unit 706, and the power supply 707 respectively. Those skilled in the art can understand that Figure 7 the structure of the electronic device shown in
[0238] The touch display screen 703 can be used to display a graphical user interface and receive operation instructions generated by a user's interaction with the graphical user interface. The touch display screen 703 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like. The touch panel can be used to collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 701, and can also receive commands sent by the processor 701 and execute them. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 701 to determine the type of touch event. Subsequently, the processor 701 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present application, the touch panel and the display panel can be integrated into the touch display screen 703 to implement input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to implement input and output functions. That is, the touch display screen 703 can also be used as part of the input unit 706 to implement the input function.
[0239] The radio frequency circuit 704 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other electronic devices through wireless communication, and transmit and receive signals with the network device or other electronic devices.
[0240] The audio circuit 705 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 705 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 705 and then converted into audio data. After the audio data is output to the processor 701 for processing, it is transmitted through the radio frequency circuit 704 to, for example, another electronic device, or the audio data is output to the memory 702 for further processing. The audio circuit 705 may also include an earphone jack to provide communication between an external earphone and the electronic device.
[0241] The input unit 706 can be used to receive input numerical, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0242] The power supply 707 is used to supply power to each component of the electronic device 700. Optionally, the power supply 707 can be logically connected to the processor 701 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 707 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0243] Although Figure 7 not shown in the figure, the electronic device 700 may further include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0244] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0245] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0246] To this end, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute any action generation method provided by the embodiments of the present application. The computer programs can execute the steps of the following action generation method:
[0247] Obtain the skeleton file of the virtual object, process the skeleton information of the above skeleton file to obtain the skeleton information vector of the above virtual object;
[0248] Obtain the music file, extract features from the above music file to obtain the music feature vector of the above music file;
[0249] Encode and process the above skeleton information vector and the above music feature vector through at least one encoder to obtain a fused vector;
[0250] Obtain the action sequence of the above virtual object through at least one decoder based on the above fused vector.
[0251] It can be seen that the computer program can be loaded by the processor to execute any action generation method provided by the embodiments of the present application, thereby bringing the following technical effects: improving the action generation efficiency of virtual objects.
[0252] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.
[0253] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0254] Since the computer program stored in the computer-readable storage medium can execute any action generation method provided by the embodiments of the present application, the beneficial effects achievable by any action generation method provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0255] According to one aspect of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the methods provided in the various optional implementation manners in the above embodiments.
[0256] In the above embodiments of the action generation device, computer-readable storage medium, electronic device, and computer program product, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the beneficial effects that can be brought by the above-described action generation device, computer-readable storage medium, computer program product, electronic device, and their corresponding units can refer to the description of the action generation method in the above embodiments and will not be elaborated here specifically.
[0257] The above has introduced in detail an action generation method, device, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An action generation method, characterized in that: The method comprises: Acquire a skeleton file of a virtual object, perform skeleton information processing on the skeleton file, and obtain a skeleton information vector of the virtual object; Acquire a music file, perform feature extraction on the music file, and obtain a music feature vector of the music file; Encoding the skeleton information vector and the music feature vector by at least one encoder to obtain a fused vector; The action sequence of the virtual object is obtained based on the fused vector by at least one decoder.
2. The action generation method according to claim 1, characterized in that: The skeleton file includes initial skeleton pose information of the skeleton of the virtual object at an initial moment, the action sequence includes pose information of the skeleton of the virtual object at multiple moments, and the acquiring of the skeleton file of the virtual object and performing skeleton information processing on the skeleton file to obtain the skeleton information vector of the virtual object include: By associating the motion features of the posture information of the bones before and after the motion in the first kinematic equation, on the basis of the initial bone posture information, calculating the information representation of the posture information to be predicted of the bones at multiple moments, wherein the information representation of the posture information to be predicted at each moment is obtained based on the information representation of the posture information to be predicted before the corresponding moment; Based on the initial skeleton posture information and the information representation of the posture information to be predicted at the multiple moments, a skeleton information vector of the virtual object is obtained.
3. The action generation method according to claim 2, characterized in that: The motion feature association is based on the motion parameter representation, and the motion feature association of the posture information of the bones before and after the movement in the first kinematic equation is used to calculate the information representation of the predicted posture information of the bones at multiple moments based on the initial bone posture information, including: By using the correlation between the motion parameters provided by the first kinematic equation and the pose information of the skeleton before and after the motion, and based on the initial skeleton pose information, calculating the information representation of the to-be-predicted pose information of the skeleton at multiple moments; The information representation of the position and posture information to be predicted at each moment is obtained based on the information representation of the position and posture information to be predicted before the corresponding moment, including: The information representation of the to-be-predicted pose information at each moment is obtained based on the motion parameters at the corresponding moment and the information representation of the to-be-predicted pose information before the corresponding moment.
4. The action generation method according to claim 3, characterized in that: The at least one decoder includes a first decoder, and obtaining the action sequence of the virtual object based on the fused vector by the at least one decoder includes: The first decoder determines the pose information that needs to be generated at the current moment based on the fused vector and the pose information generated before the current moment, so as to obtain the action sequence of the virtual object.
5. The action generation method according to claim 4, characterized in that: The determining, by the first decoder based on the fused vector and the pose information generated before the current moment, the pose information to be generated at the current moment to obtain the action sequence of the virtual object includes: Determine, by the first decoder based on the fused vector and the pose information generated before the current moment, the target motion parameters required between the pose information before the current moment and the pose information at the current moment; Based on the target motion parameters and the information representation of the posture information at the current moment, the posture information that needs to be generated at the current moment is determined to obtain the action sequence of the virtual object.
6. The action generation method according to claim 5, characterized in that: The determining, by the first decoder based on the fused vector and the pose information generated before the current moment, the target motion parameters required between the pose information before the current moment and the pose information at the current moment includes: Determine, by the first decoder, target motion parameters required between the pose information before the current moment and the pose information at the current moment based on the fused vector, the weight indication information corresponding to the fused vector, and the pose information generated before the current moment; The weight indication information is the attention weight generated for each feature vector in the fused vector based on the attention mechanism corresponding to the first decoder.
7. The action generation method according to claim 4, characterized in that: The at least one decoder includes a second decoder, and obtaining the action sequence of the virtual object based on the fused vector by the at least one decoder includes: The action sequence generated by the first decoder is adjusted through the second decoder based on the spatial distribution characteristics corresponding to the fused vector to obtain an adjusted action sequence.
8. The action generation method according to claim 2, characterized in that: Also includes: Calculating the posture adjustment information of the to-be-predicted posture information of the skeleton at multiple moments through the second kinematic equation; The step of obtaining the skeleton information vector of the virtual object based on the initial skeleton posture information and the information representation of the posture information to be predicted at the multiple moments includes: Based on the initial skeleton posture information, and the information representation and posture adjustment information corresponding to the posture information to be predicted at the multiple moments, a skeleton information vector of the skeleton is obtained.
9. The action generation method according to claim 1, characterized in that: The music feature vector includes a music attribute feature vector and a music emotion feature vector. The feature extraction of the music file to obtain the music feature vector of the music file includes: Extracting attribute features from the music file to obtain a music attribute feature vector of the music file; Based on the music file and / or the music attribute feature vector, emotional feature extraction is performed to obtain the music emotional feature vector of the music file.
10. The action generation method according to claim 9, characterized in that: The extracting of attribute features from the music file to obtain a music attribute feature vector of the music file includes: Extracting spectrum features of the music signal in the music file to obtain a music spectrum feature vector of the music file; Performing frequency domain conversion on the music signal to obtain a converted frequency domain signal; Extracting frequency domain features from the frequency domain signal to obtain a music frequency domain feature vector of the music file; Based on the music spectrum feature vector and the music frequency domain feature vector, a music attribute feature vector of the music file is obtained.
11. The action generation method according to claim 9, characterized in that: The step of extracting emotional features based on the music file and / or the music attribute feature vector to obtain the music emotional feature vector of the music file includes: Perform emotion recognition based on the music attribute feature vector to obtain an initial emotion feature vector of the music file; Extracting text from the music file to obtain music text data; Performing semantic analysis on the music text data to obtain a semantic feature vector; A music emotion feature vector of the music file is obtained based on the semantic feature vector and the initial emotion feature vector.
12. The action generation method according to claim 1, characterized in that: The encoding process of the skeleton information vector and the music feature vector by at least one encoder to obtain a fused vector includes: Splicing the skeleton information vector and the music feature vector to obtain a splicing vector; Encoding the concatenated vector based on at least one encoder to obtain an encoding result; Based on the encoding result, a fused vector is obtained.
13. The action generation method according to claim 12, characterized in that: The at least one encoder includes a first encoder, and encoding the splicing vector based on the at least one encoder to obtain an encoding result includes: By using the first encoder, performing attention processing on the concatenated vector based on an attention mechanism to obtain a first encoding result; The attention processing is to encode the splicing vector based on attention feature information of the splicing vector, and the attention feature information is used to characterize the degree of association between different information segments in the splicing vector.
14. The action generation method according to claim 9, characterized in that: The at least one encoder includes a second encoder, and encoding the splicing vector based on the at least one encoder to obtain an encoding result includes: The concatenated vector is compressed by the second encoder to obtain a second encoding result; The compression process is to encode the splicing vector based on the spatial distribution characteristics of the splicing vector.
15. The action generation method according to claim 1, characterized in that: After obtaining the action sequence of the virtual object, the method further includes: Sending the action sequence of the virtual object to a target user; Receiving action feedback information of the target user on the action sequence of the virtual object; Based on the reinforcement learning model, the action feedback information is converted into a corresponding reward mechanism, and the action sequence of the virtual object is updated based on the reward mechanism.
16. The motion generation method according to any one of claims 1 to 15, characterized in that: After obtaining the action sequence of the virtual object, the method further includes: The action sequence of the virtual object is encoded in a preset format through a preset software toolkit to obtain an action file that complies with the preset format.
17. An action generating device, characterized in that: The device comprises: An information processing module, used for acquiring a skeleton file of a virtual object, performing skeleton information processing on the skeleton file, and obtaining a skeleton information vector of the virtual object; A feature extraction module is used to obtain a music file, perform feature extraction on the music file, and obtain a music feature vector of the music file; An encoding module, used for encoding the skeleton information vector and the music feature vector through at least one encoder to obtain a fused vector; An action generation module is used to obtain an action sequence of the virtual object based on the fused vector through at least one decoder.
18. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the action generation method according to any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the action generation method according to any one of claims 1 to 16.
Citation Information
Cited By
Method for displaying animation effect and electronic equipment
CN121523773A