Method and device for generating violoncello playing action driven by music audio and electronic equipment

Through the cello performance action generation method driven by music audio, the diffused self-attention module and interactive contact loss function are used to solve the problem of insufficient modeling of the interaction between the hands and the musical instrument in the prior art, and the high-quality generation of the whole-body cello performance action and the restoration of complex interactions are achieved.

CN120564671APending Publication Date: 2025-08-29TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510644489.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing instrumental action generation methods cannot accurately model the fine interaction between the hands and the instrument, lack unified modeling of the whole body performance movements, and have limited adaptability to the instrumental performance style and rhythm changes.

Method used

The cello performance action generation method driven by music audio is used to train the action generation model through the diffusion self-attention module, and the whole-body cello performance action is optimized and generated using the interactive contact loss function, including hand interaction contact loss and bow interaction contact loss function, to construct a sample set of whole-body cello performance action and perform reverse kinematic processing.

Benefits of technology

The whole-body cello performance movements are generated, the rationality and authenticity of the generation results are improved, the performer's performance intentions and complex interactions can be delicately restored, and the fine-grained action expression is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564671A_ABST
    Figure CN120564671A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for generating a cello playing action driven by a music audio, and electronic equipment. The method comprises the following steps: acquiring a target music audio generated by a to-be-performed cello playing action; generating a whole-body cello playing action sequence according to the target music audio based on a pre-trained action generation model; the action generation model is obtained by training and optimizing according to a whole-body cello playing action sample set and a corresponding music audio sample set based on a diffusion self-attention module, and a loss function used in the training process comprises an interactive contact loss function. According to the method, the music audio is used as an input mode for generating the playing action of the whole-body cello, so that the playing intention of a player can be restored more finely; the whole-body cello playing actions are generated through the action generation model, the performance is excellent in the aspects of generating fine-grained actions and restoring complex interaction, generation of the whole-body cello playing actions is achieved, and the rationality and authenticity of the generation result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of motion generation technology, and in particular to a method, device and electronic equipment for generating cello playing motions driven by music audio. Background Art

[0002] With the rapid development of generative artificial intelligence (AI) technology, significant breakthroughs have been achieved in motion generation tasks. In recent years, a new generation of generative models, represented by generative adversarial networks, variational autoencoders, and the Transformer architecture, have demonstrated strong modeling capabilities in cross-modal tasks. Existing studies have successfully applied these techniques to motion generation tasks in various scenarios, such as dance movement generation based on audio input, demonstrating the potential of generative models for extracting semantic information from unstructured audio signals and mapping them into complex human movements.

[0003] The current technologies for generating musical instrument playing movements can be roughly divided into two main paradigms: the first uses supervised learning methods and relies on pre-collected motion capture datasets for training; the second uses reinforcement learning methods to generate playing movements that meet physical constraints through a physical simulation environment.

[0004] On the one hand, supervised learning-based methods typically require a large amount of aligned audio-action paired data as a training foundation. However, existing research has mostly focused on the rough generation of torso movements, struggling to accurately capture the complex and subtle interactions between the hands and the instrument. This results in significant deficiencies in expressiveness and realism. Furthermore, the difficulty in obtaining high-quality annotated data limits the scalability and generalization capabilities of such methods.

[0005] On the other hand, reinforcement learning-based methods typically rely on symbolic musical representations and require training and generation within a physical simulation environment. While these methods can ensure the physical plausibility of generated movements to a certain extent, their heavy reliance on simulation systems and prior rule-setting limits their applicability in more general application scenarios. Furthermore, these methods often focus on generating localized hand movements and lack a holistic modeling of the player's full-body posture and interaction with the instrument, compromising the naturalness and immersiveness of the generated movements.

[0006] In summary, the existing technology still has the following problems in generating musical instrument playing movements: first, it is difficult to accurately model the fine interaction between the hands and the instrument; second, there is a lack of unified modeling of the whole body playing movements; third, the ability to adapt to changes in musical instrument playing style and rhythm is limited.

[0007] Therefore, how to solve the problem that existing instrument playing movement generation methods cannot directly generate full-body instrument playing movements is an important issue that needs to be urgently solved in the field of movement generation. Summary of the Invention

[0008] The present invention provides a method, device and electronic device for generating cello playing movements driven by music audio, which are used to overcome the defect that existing methods for generating musical instrument playing movements cannot directly generate full-body musical instrument playing movements, realize the generation of full-body cello playing movements, and perform well in generating fine-grained movements and restoring complex interactions.

[0009] On the one hand, the present invention provides a music audio-driven cello playing action generation method, comprising: obtaining target music audio for cello playing action generation to be performed; generating a full-body cello playing action sequence according to the target music audio based on a pre-trained action generation model; wherein the action generation model is obtained by training and optimizing a full-body cello playing action sample set and a corresponding music audio sample set based on a diffuse self-attention module, and the loss function used in the training process includes an interactive contact loss function.

[0010] Furthermore, constructing the whole-body cello playing motion sample set specifically includes: obtaining existing cello playing motion capture data; restoring the bridge of the cello in the cello playing motion capture data to an arched cello bridge, and aligning the tailpiece position of the cello in the cello playing motion capture data; performing inverse kinematics processing on the human body in the cello playing motion capture data to obtain the whole-body cello playing motion sample set; wherein each whole-body cello playing motion sample in the whole-body cello playing motion sample set includes a 6D rotation space vector and a bow direction unit vector, and the 6D rotation space vector includes a body joint rotation angle and a finger joint rotation angle.

[0011] Furthermore, training the action generation model specifically includes: extracting features from each music audio sample in the music audio sample set and encoding it to obtain a corresponding music coding representation sample; in each round of training, taking the music coding representation sample, the whole-body cello playing action sample with added noise, and the time step as input, and taking the predicted whole-body cello playing action as output, iteratively optimizing the action generation model through the interactive contact loss function to obtain a pre-trained action generation model.

[0012] Furthermore, training the action generation model also includes: determining the hand-string contact distance, bow-string contact distance and bow-changing action corresponding to the predicted whole-body cello playing action; scoring the predicted whole-body cello playing action according to the hand-string contact distance, the bow-string contact distance and the bow-changing action to obtain a scoring value; and guiding the action generation model to generate the whole-body cello playing action according to the scoring value.

[0013] Furthermore, the interaction contact loss function includes a hand interaction contact loss function and a bow interaction contact loss function; wherein the hand interaction contact loss function is as follows: ; The bow interaction contact loss function is as follows: ; in, The finger that plays the note is the finger that plays the note. Indicates that the finger is not the one that plays the note. represents the predicted distance from the fingertip to the contact point, Indicates the actual distance from the fingertip to the contact point. is the indicator function, when the fundamental frequency is detected will be activated when represents the distance between the playing string and the generated bow, represents the predicted distance between the bow end point and the played string, Indicates the actual distance between the end of the bow and the string being played.

[0014] Furthermore, the diffuse self-attention module includes a normalization layer, a self-attention layer, a cross-attention layer, a scaling layer, a scaling-offset layer, and a feedforward layer.

[0015] Furthermore, the pre-trained action generation model generates a full-body cello playing action sequence according to the target music audio, including: extracting audio features in the target music audio, and encoding the audio features to obtain a music coding representation; inputting the music coding representation into the pre-trained action generation model to obtain the output of the full-body cello playing action sequence.

[0016] In a second aspect, the present invention also provides a music audio-driven cello playing action generation device, comprising: a target music audio acquisition module, used to acquire the target music audio for cello playing action generation to be performed; a whole-body cello playing action generation module, used to generate a whole-body cello playing action sequence based on the target music audio based on a pre-trained action generation model; wherein, the action generation model is obtained by training and optimizing the whole-body cello playing action sample set and the corresponding music audio sample set based on the diffuse self-attention module, and the loss function used in the training process includes the interactive contact loss function.

[0017] In a third aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating cello playing movements driven by music audio as described above is implemented.

[0018] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a cello playing action generation method driven by music audio as described in any one of the above.

[0019] The present invention provides a method for generating cello playing movements driven by music audio. The method obtains target music audio for cello playing movements to be generated, and generates a full-body cello playing movement sequence based on the target music audio based on a pre-trained movement generation model. The movement generation model is obtained by training and optimizing a full-body cello playing movement sample set and a corresponding music audio sample set based on a diffuse self-attention module. The loss function used in the training process includes an interactive contact loss function. By using music audio as the input modality for generating full-body cello playing movements, the method can more delicately restore the performer's performance intention and provide the possibility of generating stylized playing movements. At the same time, the method generates full-body cello playing movements through the trained movement generation model, and performs well in generating fine-grained movements and restoring complex interactions. It not only realizes the generation of full-body cello playing movements, but also improves the rationality and authenticity of the generated results. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 4 is a flow chart of a method for generating cello playing movements driven by music audio provided in an embodiment of the present invention.

[0022] Figure 2 Schematic diagram of the bridge of an arched cello provided in an embodiment of the present invention.

[0023] Figure 3 3 is a schematic diagram of the overall training optimization of the action generation model provided by an embodiment of the present invention.

[0024] Figure 4 2 is a schematic structural diagram of a device for generating cello playing movements driven by music audio according to an embodiment of the present invention.

[0025] Figure 5 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0027] It's important to note that existing performance movement generation methods based on supervised learning mostly focus on roughly generating torso movements, making it difficult to capture the complex and subtle interactions between the hands and the instrument. Furthermore, performance movement generation methods based on reinforcement learning typically rely on symbolic musical representations and require training and generation in a physical simulation environment, limiting their applicability to more general applications. Furthermore, the generated movements are limited to hand movements and lack full-body performance interaction modeling.

[0028] In view of this, the present invention proposes a method for generating cello playing movements driven by music audio, specifically, Figure 1 A flow chart of a method for generating cello playing movements driven by music audio provided in an embodiment of the present invention is shown.

[0029] like Figure 1 As shown, the method includes: S110, obtaining target music audio for cello playing action generation to be performed; S120, generating a full-body cello playing action sequence according to the target music audio based on a pre-trained action generation model; wherein, the action generation model is obtained by training and optimizing based on a full-body cello playing action sample set and a corresponding music audio sample set based on a diffuse self-attention module, and the loss function used in the training process includes an interactive contact loss function.

[0030] The following will describe steps S110 - S120 and related steps in detail.

[0031] S110, obtaining target music audio generated by the cello playing action to be performed.

[0032] It is easy to understand that the target music audio in this step refers to music for cello performance, including but not limited to solo pieces, concertos, chamber music, and cello parts in symphony orchestras. Examples of such music include solo pieces such as Bach's unaccompanied cello suites, concertos such as Schumann's cello concertos, chamber music such as Beethoven's cello sonatas, and cello parts in symphonies such as the finale of Beethoven's Ninth Symphony.

[0033] The target music audio can be obtained from a variety of online resources (such as online music libraries and platforms, social media and communities), or from educational institutions and school resources, without specific limitation here.

[0034] It's important to note that the target music audio in this step is highly accessible thanks to abundant online resources. Furthermore, users can access and understand the audio content without requiring specialized musical knowledge. Generating cello performance movements based on this audio can depict fine-grained hand and bow movements and restore the complex interaction between the performer and the instrument. This allows for a more nuanced reproduction of the performer's performance intent, paving the way for the generation of stylized performance movements.

[0035] After obtaining the target music audio to be generated by the cello playing action in step S110, step S120 is further executed.

[0036] S120, based on the pre-trained action generation model, generate a full-body cello playing action sequence according to the target music audio; wherein, the action generation model is obtained by training and optimizing based on the diffuse self-attention module according to the full-body cello playing action sample set and the corresponding music audio sample set, and the loss function used in the training process includes the interactive contact loss function.

[0037] It is easy to understand that in this embodiment, a motion generation model is pre-trained, which is specifically used to generate a full-body cello playing motion sequence corresponding to the target music audio.

[0038] Specifically, the action generation model is built on a diffusion self-attention module. This module combines a diffusion model with a self-attention mechanism to improve the performance of the action generation model when processing complex data structures. Diffusion models are a type of generative model that generates new data by gradually adding noise to the data and then learning to reverse this process. The self-attention mechanism allows the model to attend to information at other locations within the sequence when processing sequential data.

[0039] In the task of generating full-body cello playing movements, the diffuse self-attention module can help the movement generation model capture the dependencies between different parts to improve the authenticity and detail expression of the generated movements.

[0040] It should be noted that the motion generation model in this embodiment is pre-trained. Firstly, based on existing cello performance motion capture data, this embodiment introduces dynamic constraints and further processes the data to obtain a baseline dataset for generating three-dimensional full-body cello performance motions, namely, a full-body cello performance motion sample set. Secondly, this embodiment proposes interactive contact loss functions for the training process of the motion generation model, including a hand interaction loss function and a bow interaction loss function. Building on the geometric loss function in traditional motion generation networks, this further constrains cello performance motions, effectively improving the rationality and authenticity of the generated results.

[0041] The construction of a whole-body cello playing movement sample set and the training process of the movement generation model will be elaborated in detail in the following embodiments.

[0042] When generating actual cello playing movements, the target music audio is first preprocessed (such as feature extraction and encoding). Then, the encoded representation is input into a pre-trained movement generation model to obtain the output full-body cello playing movement sequence.

[0043] Among them, the whole-body cello playing movement sequence not only includes the human body's torso movements during the performance, but also includes delicate hand movements and complex interaction information between the human body and the cello.

[0044] It should also be noted that the target music audio obtained in step S110 is an audio sequence. Correspondingly, the full-body cello playing movements generated by the movement generation model are also movement sequences, and the full-body cello playing movements (including the movement samples described later) are all in the form of image frames.

[0045] In this embodiment, by obtaining the target music audio for cello playing movements to be generated, and based on a pre-trained movement generation model, a full-body cello playing movement sequence is generated according to the target music audio; wherein, the movement generation model is obtained by training and optimizing based on a diffuse self-attention module according to a full-body cello playing movement sample set and a corresponding music audio sample set, and the loss function used in the training process includes an interactive contact loss function. By using music audio as the input modality for generating full-body cello playing movements, this method can more delicately restore the performer's performance intention and provide the possibility of generating stylized playing movements; at the same time, by using the trained movement generation model to generate full-body cello playing movements, it performs well in generating fine-grained movements and restoring complex interactions, not only realizing the generation of full-body cello playing movements, but also improving the rationality and authenticity of the generated results.

[0046] On the basis of the above embodiments, the process of generating a full-body cello playing action sequence according to the target music audio by the action generation model will be described in detail below.

[0047] Based on a pre-trained action generation model, a full-body cello playing action sequence is generated according to the target music audio, including: extracting audio features from the target music audio, encoding the audio features, and obtaining a music coding representation; inputting the music coding representation into the pre-trained action generation model to obtain an output full-body cello playing action sequence.

[0048] It is easy to understand that first, the target music audio is preprocessed. Specifically, features are first extracted from the target music audio, and then the extracted audio features are encoded to obtain a music encoding representation. For example, in a specific embodiment, a Jukebox model with frozen parameters can be used to extract audio features from the target music audio, and then a self-attention encoder is used to encode the extracted audio features to obtain a music encoding representation.

[0049] The Jukebox model, a neural network-based music generation system developed by OpenAI, is capable of generating raw audio music based on given labels (such as artist, genre, or lyrics). The Jukebox model is unique in that it directly generates raw audio files, rather than symbolic music representations like MIDI. This means the Jukebox model can capture subtle changes in sonic quality and details in music, resulting in more realistic and complex musical works.

[0050] The self-attention encoder is a key component in deep learning architectures, widely used in fields such as natural language processing, computer vision, and audio processing. It uses a self-attention mechanism to capture the relationships between elements in an input sequence, enabling efficient representation of sequential data. The self-attention encoder consists of an input embedding layer, a positional encoding layer, a self-attention layer, a feedforward neural network layer, and residual connections and normalization layers.

[0051] After obtaining the music encoding representation corresponding to the target music audio, the music encoding representation is input into a pre-trained motion generation model to generate a full-body cello performance motion sequence. This full-body cello performance motion sequence not only includes the human torso movements during the performance, but also includes detailed hand movements and complex interactions between the human body and the cello.

[0052] In this embodiment, audio features from the target music audio are extracted and encoded to obtain a music encoding representation. This representation is then input into a pre-trained motion generation model to produce a full-body cello performance movement sequence. By using music audio as the input modality for generating full-body cello performance movements, this method can more delicately restore the performer's performance intentions, paving the way for generating stylized performance movements. Furthermore, by using a trained motion generation model to generate full-body cello performance movements, this method excels in generating fine-grained movements and restoring complex interactions. This not only enables the generation of full-body cello performance movements, but also improves the rationality and authenticity of the generated results.

[0053] On the basis of the above embodiment, the process of constructing a whole-body cello playing movement sample set will be described in detail below.

[0054] Constructing a whole-body cello playing motion sample set specifically includes: obtaining existing cello playing motion capture data; restoring the cello bridge in the cello playing motion capture data to an arched cello bridge, and aligning the position of the cello tailpiece in the cello playing motion capture data; performing inverse kinematics processing on the human body in the cello playing motion capture data to obtain a whole-body cello playing motion sample set; wherein each whole-body cello playing motion sample in the whole-body cello playing motion sample set includes a 6D rotation space vector and a bow direction unit vector, and the 6D rotation space vector includes a body joint rotation angle and a finger joint rotation angle.

[0055] It's easy to understand that existing cello performance motion capture data includes performers of varying heights and genders, as well as cellos of varying shapes and positions. To ensure consistency in the generated cello performance motions, this embodiment normalizes the existing cello performance motion capture data. This is like having the same performer replay all the music pieces on the same cello.

[0056] Specifically, the normalization processing of existing cello performance motion capture data includes the normalization of the cello, the normalization of the human body (performer), and the correction of the interaction between the human body and the cello.

[0057] To normalize the cello, this example uses a manually annotated cello as the shared cello and restores the cello bridge in the cello performance motion capture data to an arched cello bridge to better match the bridge of a real cello. This adjustment allows the generated cello performance to theoretically play the two middle strings without clipping.

[0058] Accordingly, Figure 2The schematic diagram of the bridge of the arched cello provided by the embodiment of the present invention is shown. Figure 2 In the data, the “motion capture dataset” refers to the existing cello performance motion capture data. Figure 2 Above: Position the starting point of the bow (the base of the arch) midway between the PIP and DIP joints of the middle and ring fingers and thumb (highlighted in red). Figure 2 Bottom: As shown in (b), this embodiment reconstructs an arched cello bridge, unlike the flat bridge in the mocap dataset and closely resembling the actual instrument shown in (c). This prevents the performer from accidentally touching adjacent strings when playing the two middle strings, thus avoiding the potential clipping demonstrated in (a). The red dots in (a) and (b) indicate the points of contact between the bow and the strings during playing.

[0059] The cello in the cello performance motion capture data is then aligned with the cello's tailpiece position across all frames. Specifically, the Kabsch algorithm can be used to efficiently calculate the optimal rotation matrix to align the cello in each frame with the shared cello, while the performer's full body key points are also rotated along with the cello.

[0060] In human body normalization, this embodiment uses VPoser to perform inverse kinematics processing based on the SMPL-X model in two stages. In the first stage, the human body is subjected to initial inverse kinematics processing to fit the average body shape, global orientation, and global displacement of the human body across all frames. In the second stage, the average body shape from the first stage is used to refine the inverse kinematics processing, prioritizing accurate fitting of the wrist rather than fitting key points of the elbow and shoulder. This strategy can utilize the global wrist rotation provided in the cello performance motion capture data while preserving the local rotation of the hand joints.

[0061] Furthermore, after the cello and human body normalization, the interaction between the human body and the cello needs to be corrected to make the playing movements more natural and realistic.

[0062] Based on the above data processing, this embodiment can obtain a total of about 7000 seconds of full-body cello playing motion data represented by 6D rotation, that is, a full-body cello playing motion sample set. Among them, the body consists of 21 joints (excluding the pelvis), and each hand consists of 15 joints, and the rotation space vector is obtained. ,in The direction of the bow can be represented by the unit vector It represents the unit vector of the bow direction. Figure 2As shown in , the starting point of the bow (also called the bow root) is located between the middle finger, ring finger, and thumb of the left hand. Therefore, when the bow length is fixed, its end point (called the bow tip) can be determined by the bow root and the unit direction vector. Therefore, each full-body cello playing action sample in the full-body cello playing action sample set can be expressed as ,in .

[0063] In this embodiment, by introducing dynamic constraints on the basis of existing cello performance motion capture data and further processing the data, a baseline data set for generating three-dimensional cello performance motions, that is, a full-body cello performance motion sample set, is obtained for training and optimizing the motion generation model. The full-body cello performance motions are generated by the trained motion generation model, which performs well in generating fine-grained motions and restoring complex interactions. It not only realizes the generation of full-body cello performance motions, but also improves the rationality and authenticity of the generated results.

[0064] On the basis of the above embodiments, the training optimization process of the action generation model will be further described in detail below.

[0065] Training the action generation model specifically includes: extracting features from each music audio sample in the music audio sample set and encoding it to obtain a corresponding music encoding representation sample; in each round of training, the music encoding representation sample, the whole-body cello playing action sample with added noise, and the time step are used as input, and the predicted whole-body cello playing action is used as output. The action generation model is iteratively optimized through the interactive contact loss function to obtain a pre-trained action generation model.

[0066] It is easy to understand that after constructing a full-body cello playing movement sample set, the movement generation model is trained and optimized in combination with the corresponding music audio sample set.

[0067] Figure 3 The figure shows the overall training optimization diagram of the action generation model provided by the embodiment of the present invention. Figure 3 In the figure, the left side is the preprocessing part, the middle "diffused self-attention module" is the action generation model, and the right side is the predicted output of the training process.

[0068] like Figure 3 As shown in Figure 1, during training, the input consists of three parts: the noisy action (i.e., a noisy full-body cello playing action sample), the music audio sample, and the time step. The music audio sample needs to be preprocessed before being input into the action generation model.

[0069] Regarding the preprocessing process of music audio samples, a Jukebox model with frozen parameters can be used to extract audio features from the music audio samples, and then a self-attention encoder is used to encode the audio features to obtain music encoding representation samples corresponding to the music audio samples.

[0070] Subsequently, in each round of training, the music encoding representation samples, the whole-body cello playing movement samples with added noise, and the time steps are used as the input of the training model, and the predicted whole-body cello playing movement is used as the output of the training model. The action generation model is iteratively optimized through the interactive contact loss function, and finally a trained action generation model is obtained.

[0071] according to Figure 3 It can be seen that the action generation model in this embodiment is constructed based on the diffuse self-attention module, wherein the diffuse self-attention module includes a normalization layer, a self-attention layer, a cross-attention layer, a scaling layer, a scaling-offset layer, and a feedforward layer.

[0072] It is worth mentioning that the action generation model provided in this embodiment introduces a cross-attention layer and combines it with a scaling layer and a scaling-offset layer. The introduction of the cross-attention layer enables the action generation model to focus not only on the relationship between elements within a single sequence, but also on the interaction between two different sequences. The introduction of the scaling layer and the scaling-offset layer, combined with zero initialization, enables the model to begin training as an identity function, thereby improving training efficiency and enhancing model performance.

[0073] according to Figure 3 It can also be seen that in the diffuse self-attention module, there are residual connections between different network layers, which can effectively solve the problems of gradient disappearance and gradient explosion in deep network training.

[0074] About the interactive contact loss function. Figure 3 On the right, the orange solid line represents the "contact" in the interactive contact loss function; the gray dotted line represents the "interaction" relationship between the non-playing fingers and the contact point, and the "interaction" relationship between the two ends of the bow and the played string. The red dots represent the contact positions of the hands, and the blue highlighted strings represent the played strings. As far as the hands are concerned, the fingers that play the notes should strive to contact the position specified by the audio, while the other fingers should maintain an appropriate spatial relationship with the contact position. This embodiment uses the hand interaction release loss function for constraints. Similarly, the bow must contact the starting string while maintaining an appropriate distance relationship between the two ends of the bow and the string. This embodiment uses the bow interaction loss function for constraints.

[0075] Specifically, the hand interaction contact loss function is as follows (1), and the bow interaction contact loss function is as follows (2).

[0076] (1).

[0077] (2).

[0078] In formulas (1) and (2), The finger that plays the note is the finger that plays the note. Indicates that the finger is not the one that plays the note. represents the predicted distance from the fingertip to the contact point, Indicates the actual distance from the fingertip to the contact point. is an indicator function. When the pitch (i.e. fundamental frequency) is detected ) will be activated, represents the distance between the playing string and the generated bow, represents the predicted distance between the bow end point and the played string, Indicates the actual distance between the end of the bow and the string being played.

[0079] It should be noted that, in addition to the above-mentioned interaction contact loss function, the action generation model can also simultaneously apply other commonly used loss functions, such as joint position loss function and joint velocity loss function, which are not specifically limited here.

[0080] Furthermore, to better measure the quality of the predicted full-body cello playing movements output by the model during training, this embodiment also designs evaluation metrics for the predicted full-body cello playing movements. Specifically, for the predicted full-body cello playing movements output by the model during training, the hand-string contact distance, bow-string contact distance, and bow-changing motion corresponding to the predicted full-body cello playing movements are determined. Then, based on the hand-string contact distance, bow-string contact distance, and bow-changing motion, the predicted full-body cello playing movements are scored to obtain a score value. Finally, the score value is used to guide the movement generation model in generating the full-body cello playing movements.

[0081] The hand-string contact distance refers to the deviation between the fingertips of the left hand playing the note and the cello trigger position for the current pitch. While a cello passage offers a variety of plausible playing motions, since most notes can be played on different strings using different techniques, this embodiment uses the trigger position closest to the player's finger playing the note as the "intended" position for the generated motion. The shorter the hand-string contact distance, the more accurate and appropriate the generated predicted full-body cello playing motion.

[0082] The bow-string contact distance reflects the deviation between the bow and the string being played. Once the trigger position is determined, this deviation can be uniquely identified. The smaller the bow-string contact distance, the more accurately the interaction between the bow and the string vibration can be reproduced.

[0083] While there are no strict rules for bow changes, there are specific moments that are more suitable for bow changes based on musical rhythm and phrases. Bow Change F1 examines whether the generated movements are consistent with these musically motivated bow changes, reflecting the ability of the movement generation model to drive bow changes based on musical audio.

[0084] It is worth mentioning that the hand-string contact distance, bow-string contact distance and bow-changing movements are all based on the physics of sound production and specific field knowledge of cello, targeting key performance elements while improving the accuracy of the spatial relationship between the performer and the cello.

[0085] It should be noted that the measurement of the quality of the predicted whole-body cello playing movements can also be applied to the actual reasoning output of the whole-body cello playing movement sequence, which will not be elaborated here.

[0086] In this embodiment, features are extracted and encoded for each music audio sample in the music audio sample set to obtain corresponding music coding representation samples. In each round of training, the music coding representation samples, the whole-body cello playing action samples with added noise, and the time steps are used as input, and the predicted whole-body cello playing action is used as output. The action generation model is iteratively optimized through the interactive contact loss function to obtain a pre-trained action generation model, and then the whole-body cello playing action is generated through the trained action generation model. It performs well in generating fine-grained actions and restoring complex interactions, which not only realizes the generation of whole-body cello playing actions, but also improves the rationality and authenticity of the generated results.

[0087] In some other embodiments, a 10% full-body cello performance motion sample set and the corresponding music audio sample set are used as a test set. By analyzing the direction of the bow root's movement relative to the bridge, the bow attack in the ground truth and generated motion is detected. In this embodiment, a tolerance δ of 3 frames (0.1 seconds) is applied: if the predicted bow change is within the actual bow change of If the distance between the bow and the string falls within the range, it is considered a true positive. Furthermore, to further assess the consistency between bowing techniques and human performance, this example calculates the cosine similarity of the relative distance between the bow and the string in the time dimension. When the lower half of the bow plays the string, the relative position is negative, while when the upper half of the bow plays the string, the relative position is positive. Specific test results can be found in the table below.

[0088] In the table above, ICL refers to the Interaction Contact Loss function, HICL refers to the Hand Interaction Contact Loss function, and BICL refers to the Bow Interaction Contact Loss function. The lower the hand-string contact distance and bow-string distance, the higher the bowing F1 score and bowing cosine similarity, and the better the model prediction performance.

[0089] As can be seen from the table above, without configuring the interactive contact loss function, the model performs poorly in hand-string contact distance and bow-string contact distance, and the bowing F1 score and cosine similarity are average.

[0090] When only the hand interaction loss function is configured, the hand-string contact distance and bow-string contact distance are significantly reduced, indicating that the model has improved in this aspect. However, the bowing F1 score and cosine similarity decrease slightly, which may be because the hand interaction loss function mainly optimizes features related to contact distance.

[0091] When both the hand interaction contact loss function and the bow interaction contact loss function are configured, the hand-string contact distance is further reduced to 15.60. Although it is slightly higher than the case with only the hand interaction contact loss function, it is still better than the case without the interaction contact loss function; the bow-string contact distance is greatly reduced to 5.40, showing a significant improvement; the bowing F1 score is improved to 0.4721, indicating that the model has significantly improved in the classification or detection of bowing movements; the bowing cosine similarity reaches 0.7515, showing that the model performs best in capturing the similarity of bowing movements.

[0092] As can be seen, the introduction of the hand-to-bow interaction loss function significantly improves the model's performance in many aspects, especially in bow-string contact distance, bowing F1 score, and cosine similarity. This shows that the effective combination of these two loss functions helps the model better learn and understand complex bowing movements and their related features.

[0093] Finally, it is worth mentioning that the music audio-driven cello playing action generation method provided by the embodiment of the present invention realizes the generation of whole-body movements, compared with previous musical instrument playing action generation methods that can only generate partial torso or two-hand movements, and performs well in generating fine-grained movements and restoring complex interactions. The present invention further proposes hand interaction contact loss and bow interaction contact loss to maintain the authenticity of the interaction between the performer and the instrument. In addition, in order to better evaluate the generated movements, the present invention also introduces special indicators for string performance, including hand-string contact distance, bow-string contact distance and bowing score. In addition, the present invention also provides a benchmark dataset for musical instrument playing action generation, which is a dataset derived from the motion capture dataset after reprocessing and normalization and is specifically designed for action generation. The present invention opens up a variety of promising avenues for future work, provides new insights and inspiration for the research field, and is expected to promote further development in the fields of animation generation, music education, and interactive art creation.

[0094] Corresponding to the methods for generating cello playing movements driven by music audio described in the above embodiments, the present invention also proposes a device for generating cello playing movements driven by music audio.

[0095] Specifically, Figure 4 A schematic structural diagram of a device for generating cello playing movements driven by music audio provided in an embodiment of the present invention is shown.

[0096] like Figure 4 As shown, the device includes: a target music audio acquisition module 410, which is used to obtain the target music audio for cello playing action generation to be performed; a whole-body cello playing action generation module 420, which is used to generate a whole-body cello playing action sequence according to the target music audio based on a pre-trained action generation model; wherein, the action generation model is obtained by training and optimizing according to a whole-body cello playing action sample set and a corresponding music audio sample set based on a diffuse self-attention module, and the loss function used in the training process includes an interactive contact loss function.

[0097] In this embodiment, the target music audio for cello playing movements to be generated is obtained by the target music audio acquisition module 410, and the whole-body cello playing movement generation module 420 generates a whole-body cello playing movement sequence based on the target music audio based on the pre-trained movement generation model; wherein, the movement generation model is obtained by training and optimizing the whole-body cello playing movement sample set and the corresponding music audio sample set based on the diffuse self-attention module, and the loss function used in the training process includes the interaction contact loss function. By using music audio as the input modality for generating whole-body cello playing movements, the device can restore the performer's performance intention more delicately, providing the possibility of generating stylized playing movements; at the same time, by generating whole-body cello playing movements through the trained movement generation model, it performs well in generating fine-grained movements and restoring complex interactions, not only realizing the generation of whole-body cello playing movements, but also improving the rationality and authenticity of the generated results.

[0098] It should be noted that the music audio driven cello playing action generation device provided in the embodiment of the present invention and the music audio driven cello playing action generation method described in the above embodiments can be referred to each other, and will not be repeated here.

[0099] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communications bus 540. The processor 510 may call the logic instructions in the memory 530 to execute a music audio-driven cello playing action generation method, which includes: obtaining target music audio for cello playing action generation; generating a full-body cello playing action sequence based on the target music audio based on a pre-trained action generation model; wherein the action generation model is obtained by training and optimizing a full-body cello playing action sample set and a corresponding music audio sample set based on a diffuse self-attention module, and the loss function used in the training process includes an interactive contact loss function.

[0100] Furthermore, the logic instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0101] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the music audio-driven cello playing action generation method provided by the above-mentioned methods, the method comprising: obtaining the target music audio for cello playing action generation to be performed; based on a pre-trained action generation model, generating a full-body cello playing action sequence according to the target music audio; wherein, the action generation model is obtained by training and optimizing based on a diffuse self-attention module according to a full-body cello playing action sample set and a corresponding music audio sample set, and the loss function used in the training process includes an interactive contact loss function.

[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0103] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for generating cello playing movements driven by music audio, characterized in that: include: Obtain target music audio generated by the cello playing action to be performed; Based on the pre-trained action generation model, a full-body cello playing action sequence is generated according to the target music audio; Among them, the action generation model is obtained by training and optimizing based on the diffuse self-attention module according to the full-body cello performance action sample set and the corresponding music audio sample set. The loss function used in the training process includes the interactive contact loss function.

2. The method for generating cello playing movements driven by music audio according to claim 1, wherein: Constructing the whole-body cello playing action sample set specifically includes: Obtain existing cello performance motion capture data; Restoring the cello bridge in the cello performance motion capture data to an arched cello bridge, and aligning the tailpiece positions of the cello in the cello performance motion capture data; Performing inverse kinematics processing on the human body in the cello playing motion capture data to obtain the whole-body cello playing motion sample set; Each whole-body cello playing action sample in the whole-body cello playing action sample set includes a 6D rotation space vector and a bow direction unit vector, and the 6D rotation space vector includes a body joint rotation angle and a finger joint rotation angle.

3. The method for generating cello playing movements driven by music audio according to claim 1, wherein: Training the action generation model specifically includes: Extracting features from each music audio sample in the music audio sample set and encoding the sample to obtain a corresponding music encoding representation sample; In each round of training, the music encoding representation sample, the whole-body cello playing movement sample with added noise, and the time step are used as input, and the predicted whole-body cello playing movement is used as output. The movement generation model is iteratively optimized through the interactive contact loss function to obtain a pre-trained movement generation model.

4. The method for generating cello playing movements driven by music audio according to claim 1, wherein: Training the action generation model further includes: determining a hand-string contact distance, a bow-string contact distance, and a bow-changing action corresponding to the predicted full-body cello playing action; scoring the predicted full-body cello playing action according to the hand-string contact distance, the bow-string contact distance, and the bow-changing action to obtain a scoring value; According to the scoring value, the action generation model is guided to generate the whole-body cello playing action.

5. The method for generating cello playing movements driven by music audio according to claim 1, wherein: The interactive contact loss function includes a hand interactive contact loss function and a bow interactive contact loss function; wherein, The hand interaction contact loss function is as follows: ; The bow interaction contact loss function is as follows: ; in, The finger that plays the note is the finger that plays the note. Indicates that the finger is not the one that plays the note. represents the predicted distance from the fingertip to the contact point, Indicates the actual distance from the fingertip to the contact point. Is the indicator function, when the fundamental frequency is detected will be activated when represents the distance between the playing string and the generated bow, represents the predicted distance between the bow end point and the played string, Indicates the actual distance between the end of the bow and the string being played.

6. The method for generating cello playing movements driven by music audio according to claim 1, characterized in that: The diffuse self-attention module includes a normalization layer, a self-attention layer, a cross-attention layer, a scaling layer, a scaling-offset layer, and a feedforward layer.

7. The method for generating cello playing movements driven by music audio according to any one of claims 1 to 6, characterized in that: The pre-trained action generation model generates a full-body cello playing action sequence according to the target music audio, including: Extracting audio features from the target music audio, and encoding the audio features to obtain a music encoding representation; The music encoding representation is input into a pre-trained action generation model to obtain the output of the full-body cello playing action sequence.

8. A cello playing action generation device driven by music audio, characterized in that: include: A target music audio acquisition module is used to acquire target music audio generated by the cello performance action to be performed; A full-body cello playing movement generation module is used to generate a full-body cello playing movement sequence according to the target music audio based on a pre-trained movement generation model; Among them, the action generation model is obtained by training and optimizing based on the diffuse self-attention module according to the full-body cello performance action sample set and the corresponding music audio sample set. The loss function used in the training process includes the interactive contact loss function.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for generating cello playing movements driven by music audio as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating cello playing movements driven by music audio as claimed in any one of claims 1 to 7 is implemented.