A resource synthesis method, device, equipment and storage medium

By obtaining the characteristic information of music resources and using preset models to generate dance resources, the problem of high cost in the synthesis of virtual human dance moves is solved, and efficient automatic synthesis of dance and music is achieved.

CN114756706BActive Publication Date: 2025-07-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210359938.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-07-08
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

The prior art requires a lot of labor, time and economic costs in the synthesis of virtual human dance moves, and the synthesis efficiency of dance and music is inefficient.

Method used

By obtaining the characteristic information of the music resources, using the preset model to generate corresponding dance resources, combining the action completion model to process discontinuous movements, and achieving automatic synthesis of dance and music.

Benefits of technology

There is no need to manually design dance moves and collect resources, and the synthetic resources are generated quickly and accurately, improving the synthesis efficiency of dance and music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114756706B_ABST
    Figure CN114756706B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a resource synthesis method, apparatus, device, and storage medium, which relates to the field of computer technologies and can improve the synthesis efficiency of dance and music. The resource synthesis method includes: obtaining music resources to be synthesized; the music resources to be synthesized include at least one music segment; determining at least one piece of feature information corresponding one-to-one to the at least one music segment; inputting the at least one piece of feature information into a preset model to obtain at least one dance resource corresponding one-to-one to the at least one piece of feature information; the preset model is trained based on a plurality of sample resources to meet a first preset condition and is used to determine the dance resource corresponding to the input feature information; a sample resource pair includes a music sample resource and a dance sample resource; generating a synthesized resource according to the at least one dance resource and the at least one music segment; the synthesized resource includes the music resources to be synthesized and the dance resources corresponding to the music resources to be synthesized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a resource synthesis method, apparatus, device, and storage medium. Background Art

[0002] Digital virtual humans with lifelike behaviors are an important part of technologies such as the metaverse, games, virtual live broadcasts, and virtual reality. In virtual human live broadcasts, if virtual humans can have talents similar to real humans, such as automatically choreographing to given music, it will bring great attraction to virtual humans.

[0003] Currently, in most scenarios of virtual human dance movements, it is often necessary for motion capture actors who can dance to choreograph according to the music, then perform motion capture through optical capture devices, and finally manually repair the motion capture data before it can be applied to virtual humans. Although this method results in accurate and vivid dance movements, it consumes a large amount of labor costs, time costs, and economic costs, reducing the synthesis efficiency of dance and music. Summary of the Invention

[0004] The present disclosure provides a resource synthesis method, apparatus, device, and storage medium, which can improve the synthesis efficiency of dance and music.

[0005] The technical solution of the embodiment of the present disclosure is as follows:

[0006] According to the first aspect of the embodiment of the present disclosure, a resource synthesis method is provided, and this method can be applied to an electronic device. The method may include:

[0007] Obtain music resources to be synthesized; the music resources to be synthesized include at least one music segment;

[0008] Determine at least one piece of feature information corresponding one-to-one to the at least one music segment;

[0009] Input the at least one piece of feature information into a preset model to obtain at least one dance resource corresponding one-to-one to the at least one piece of feature information; the preset model is trained based on multiple sample resources to meet a first preset condition and is used to determine the dance resource corresponding to the input feature information; a sample resource pair includes a music sample resource and a dance sample resource;

[0010] Generate a synthesis resource based on the at least one dance resource and the at least one music segment; the synthesis resource includes the music resources to be synthesized and the dance resources corresponding to the music resources to be synthesized.

[0011] Optionally, the feature information includes: music type feature information and music rhythm feature information;

[0012] Determine at least one piece of feature information that corresponds one-to-one with at least one music segment, including:

[0013] Input at least one music segment into a music genre feature extraction model to obtain at least one music genre feature information that corresponds one-to-one with the at least one music segment; the music genre feature extraction model is trained based on a first sample resource; the first sample resource is a sample resource that includes music genre features among multiple sample resource pairs;

[0014] Input at least one music segment into a music rhythm feature extraction model to obtain at least one music rhythm feature information that corresponds one-to-one with the at least one music segment; the music rhythm feature extraction model is trained based on a second sample resource; the second sample resource is a sample resource that includes music rhythm features among multiple sample resource pairs.

[0015] Optionally, the at least one music segment includes: a first music segment and a second music segment consecutive to the first music segment; the at least one dance resource includes: a first dance resource corresponding to the first music segment and a second dance resource corresponding to the second music segment; the first dance resource includes: an ending action divided according to the order of the dance actions of the first dance resource; the second dance resource includes: a starting action divided according to the order of the dance actions of the second dance resource;

[0016] When the ending action and the starting action are non-consecutive dance actions, generate a composite resource according to the at least one dance resource and the at least one music segment, including:

[0017] Input the ending action and the starting action into an action completion model to obtain a first action corresponding to the ending action and a second action corresponding to the starting action; the action completion model is trained based on a third sample resource; the third sample resource is a sample resource that includes dance action features among multiple sample resource pairs; the first action and the second action are consecutive dance actions;

[0018] Update the ending action to the first action to obtain an updated first dance resource;

[0019] Update the starting action to the second action to obtain an updated second dance resource;

[0020] Generate a composite resource according to the first music segment, the second music segment, the updated first dance resource, and the updated second dance resource.

[0021] Optionally, the resource synthesis method further includes:

[0022] Obtain multiple pairs of sample resources; one pair of sample resources includes multiple segments of sample resource pairs; one segment of a sample resource pair includes a segment of a music sample resource and a segment of a dance sample resource;

[0023] Determine at least one clustering result of the multiple segments of sample resource pairs according to the multiple segments of sample resource pairs and a clustering algorithm; one clustering result corresponds to one dance resource;

[0024] Use at least one clustering result to train a preset hidden Markov model until a first preset condition is met to obtain a preset model.

[0025] Optionally, after obtaining multiple pairs of sample resources, it further includes:

[0026] Use the segments of music sample resources in the multiple segments of sample resource pairs to train a preset classification task model until a second preset condition is met to obtain a music type feature extraction model.

[0027] Optionally, after obtaining multiple pairs of sample resources, it further includes:

[0028] Obtain rhythm features corresponding to the music beats from the segments of music sample resources in each segment of the sample resource pairs;

[0029] Obtain motion features corresponding to the music beats from the segments of dance sample resources in each segment of the sample resource pairs;

[0030] Use the rhythm features and the motion features to train a preset feature extraction model until a third preset condition is met to obtain a music rhythm feature extraction model.

[0031] Optionally, the multiple segments of sample resource pairs include: a first segment of sample resource pair and a second segment of sample resource pair; the first segment of sample resource pair includes: a first segment of dance sample resource; the second segment of sample resource pair includes: a second segment of dance sample resource;

[0032] After obtaining multiple pairs of sample resources, it further includes:

[0033] Obtain a first sub-segment and a second sub-segment; the first sub-segment is at least one sub-segment sorted from the back to the front in the order of dance movements in the first segment of dance sample resource; the second sub-segment is at least one sub-segment sorted from the front to the back in the order of dance movements in the second segment of dance sample resource;

[0034] Train an action completion model according to the first sub-segment, the second sub-segment and a preset algorithm.

[0035] According to a second aspect of the embodiments of the present disclosure, a resource synthesis device is provided, which can be applied to an electronic device. The device may include: an acquisition unit, a processing unit, and a generation unit;

[0036] The acquisition unit is configured to acquire music resources to be synthesized; the music resources to be synthesized include at least one music segment;

[0037] The processing unit is configured to determine at least one piece of feature information corresponding one-to-one to the at least one music segment;

[0038] The processing unit is further configured to input the at least one piece of feature information into a preset model to obtain at least one dance resource corresponding one-to-one to the at least one piece of feature information; the preset model is trained based on multiple sample resources to meet a first preset condition and is used to determine the dance resource corresponding to the input feature information; a sample resource pair includes a music sample resource and a dance sample resource;

[0039] The generation unit is configured to generate a synthesized resource according to the at least one dance resource and the at least one music segment; the synthesized resource includes the music resources to be synthesized and the dance resource corresponding to the music resources to be synthesized.

[0040] Optionally, the feature information includes: music type feature information and music rhythm feature information;

[0041] The processing unit is specifically configured to:

[0042] Input the at least one music segment into a music type feature extraction model to obtain at least one piece of music type feature information corresponding one-to-one to the at least one music segment; the music type feature extraction model is trained based on a first sample resource; the first sample resource is a sample resource including music type features among the multiple sample resource pairs;

[0043] Input the at least one music segment into a music rhythm feature extraction model to obtain at least one piece of music rhythm feature information corresponding one-to-one to the at least one music segment; the music rhythm feature extraction model is trained based on a second sample resource; the second sample resource is a sample resource including music rhythm features among the multiple sample resource pairs.

[0044] Optionally, the at least one music segment includes: a first music segment and a second music segment consecutive to the first music segment; the at least one dance resource includes: a first dance resource corresponding to the first music segment and a second dance resource corresponding to the second music segment; the first dance resource includes: an ending action divided according to the sequence of dance actions of the first dance resource; the second dance resource includes: a starting action divided according to the sequence of dance actions of the second dance resource;

[0045] When the ending action and the starting action are non - consecutive dance actions, the generating unit is specifically configured to:

[0046] Input the ending action and the starting action into the action completion model to obtain a first action corresponding to the ending action and a second action corresponding to the starting action; the action completion model is trained according to the third sample resource; the third sample resource is the sample resource including dance action features in multiple sample resource pairs; the first action and the second action are consecutive dance actions;

[0047] Update the ending action to the first action to obtain the updated first dance resource;

[0048] Update the starting action to the second action to obtain the updated second dance resource;

[0049] Generate a composite resource according to the first music segment, the second music segment, the updated first dance resource, and the updated second dance resource.

[0050] Optionally, the obtaining unit is further configured to obtain multiple sample resource pairs; one sample resource pair includes multiple sample resource pair segments; one sample resource pair segment includes a music sample resource segment and a dance sample resource segment;

[0051] The processing unit is further configured to determine at least one clustering result of the multiple sample resource pair segments according to the multiple sample resource pair segments and a clustering algorithm; one clustering result corresponds to one dance resource;

[0052] The processing unit is further configured to train a preset hidden Markov model with at least one clustering result until a first preset condition is met to obtain a preset model.

[0053] Optionally, the processing unit is further configured to train a preset classification task model with the music sample resource segments in the multiple sample resource pair segments until a second preset condition is met to obtain a music type feature extraction model.

[0054] Optionally, the processing unit is further configured to obtain rhythm features corresponding to music beats from the music sample resource segments of each sample resource pair segment in the multiple sample resource pair segments;

[0055] Obtain action features corresponding to music beats from the dance sample resource segments of each sample resource pair segment;

[0056] Train a preset feature extraction model with the rhythm features and the action features until a third preset condition is met to obtain a music rhythm feature extraction model.

[0057] Optionally, multiple sample resource pairs of segments include: a first sample resource pair of segments and a second sample resource pair of segments; the first sample resource pair of segments includes: a first dance sample resource segment; the second sample resource pair of segments includes: a second dance sample resource segment;

[0058] The obtaining unit is further configured to obtain a first sub-segment and a second sub-segment; the first sub-segment is at least one sub-segment sorted from the back to the front in the order of dance movements in the first dance sample resource segment; the second sub-segment is at least one sub-segment sorted from the front to the back in the order of dance movements in the second dance sample resource segment;

[0059] The processing unit is further configured to train an action completion model according to the first sub-segment, the second sub-segment, and a preset algorithm.

[0060] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, which may include: a processor and a memory for storing processor-executable instructions; wherein, the processor is configured to execute the instructions to implement any one of the optional resource synthesis methods in the first aspect above.

[0061] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which instructions are stored. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute any one of the optional resource synthesis methods in the first aspect above.

[0062] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided. The computer program product includes computer instructions, and when the computer instructions run on an electronic device, the electronic device executes the resource synthesis method as described in any one of the optional implementation manners in the first aspect.

[0063] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.

[0064] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0065] Based on any of the above aspects, in the present disclosure, when obtaining the music resources to be synthesized, the music resources to be synthesized can be divided into at least one music segment, and at least one feature information corresponding to the at least one music segment is determined. Subsequently, the at least one feature information is input into a preset model to obtain at least one dance resource corresponding to the at least one feature information, and a synthesized resource is generated according to the at least one dance resource and the at least one music segment. Since the preset model is trained based on multiple sample resources to meet the first preset condition and is used to determine the dance resource corresponding to the input feature information, and the synthesized resource includes the music resources to be synthesized and the dance resources corresponding to the music resources to be synthesized, therefore, the present disclosure can quickly and accurately determine the dance resources corresponding to the music resources to be synthesized through the preset model and generate the synthesized resource, without designing a set of dance movements for each piece of music separately, nor collecting dance resources through dancers for simulation, improving the synthesis efficiency of music resources and dance resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0067] Figure 1 The flowchart showing the process of a resource synthesis method provided by an embodiment of the present disclosure;

[0068] Figure 2 The flowchart showing the process of another resource synthesis method provided by an embodiment of the present disclosure;

[0069] Figure 3 The flowchart showing the process of another resource synthesis method provided by an embodiment of the present disclosure;

[0070] Figure 4 The flowchart showing the process of another resource synthesis method provided by an embodiment of the present disclosure;

[0071] Figure 5 The flowchart showing the process of another resource synthesis method provided by an embodiment of the present disclosure;

[0072] Figure 6 The flowchart showing the process of another resource synthesis method provided by an embodiment of the present disclosure;

[0073] Figure 7 The flowchart showing the process of another resource synthesis method provided by an embodiment of the present disclosure;

[0074] Figure 8 The structural diagram showing another resource synthesis device provided by an embodiment of the present disclosure;

[0075] Figure 9 Shows a schematic structural diagram of a terminal provided by an embodiment of the present disclosure;

[0076] Figure 10 Shows a schematic structural diagram of a server provided by an embodiment of the present disclosure. Detailed implementation manners

[0077] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0078] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0079] It should also be understood that the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, and / or components.

[0080] The data involved in the present disclosure can be data authorized by the user or fully authorized by all parties.

[0081] As described in the background art, the method of directly collecting dance data requires a large amount of labor cost, time cost, and economic cost, reducing the efficiency of dance generation. And directly predicting dance movements through a neural network is prone to generating unreasonable posture movements, requiring a large amount of post-processing work, reducing the synthesis efficiency of dance and music.

[0082] Based on this, embodiments of the present disclosure provide a resource synthesis method. When obtaining music resources to be synthesized, the music resources to be synthesized can be divided into at least one music segment, and at least one feature information corresponding to the at least one music segment is determined. Subsequently, the at least one feature information is input into a preset model to obtain at least one dance resource corresponding to the at least one feature information, and a synthesized resource is generated according to the at least one dance resource and the at least one music segment. Since the preset model is trained based on multiple sample resources to meet a first preset condition and is used to determine the dance resource corresponding to the input feature information, and the synthesized resource includes the music resources to be synthesized and the dance resources corresponding to the music resources to be synthesized, therefore, the present disclosure can quickly and accurately determine the dance resources corresponding to the music resources to be synthesized through the preset model and generate the synthesized resource, without designing a set of dance movements for each piece of music separately, nor collecting dance resources through dancers for simulation, thus improving the synthesis efficiency of music resources and dance resources.

[0083] The following provides an exemplary description of the resource synthesis method provided by embodiments of the present disclosure:

[0084] The resource synthesis method provided by the present disclosure can be applied to an electronic device.

[0085] In some embodiments, the electronic device can be a server, a terminal, or other electronic devices for resource synthesis, and the present disclosure does not limit this.

[0086] Among them, the server can be a single server, or can also be a server cluster composed of multiple servers. In some embodiments, the server cluster can also be a distributed cluster. The present disclosure also does not limit the specific implementation manner of the server.

[0087] The terminal can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as devices such as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, etc. that can install and use a content community application (such as Kuaishou). The present disclosure does not impose special restrictions on the specific form of the electronic device. It can perform human-computer interaction with the user through one or more methods such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device.

[0088] The following provides a detailed introduction to the resource synthesis method provided by embodiments of the present application with reference to the accompanying drawings.

[0089] As Figure 1 shown, when the resource synthesis method is applied to an electronic device, the resource synthesis method may include:

[0090] S101. The electronic device obtains the music resources to be synthesized.

[0091] Among them, the music resources to be synthesized include at least one music segment.

[0092] Optionally, the music resources to be synthesized may be a piece of music. Since this piece of music may include multiple music styles or multiple music rhythms, the electronic device may divide the music resources into at least one music segment according to the music measures.

[0093] Optionally, when the electronic device divides the music resources to be synthesized according to the music measures, it may find the music measures by detecting the beats per second of the music. Then, since the duration of each measure may be different, the electronic device may set each measure to a fixed duration to obtain at least one music segment.

[0094] S102. The electronic device determines at least one piece of feature information corresponding to at least one music segment one by one.

[0095] Optionally, the feature information may include music type feature information and music rhythm feature information.

[0096] Among them, the music type feature information is used to represent the music type of the music segment. The music type may also be referred to as the music style, such as jazz, rock, hip-hop, classical music, etc.

[0097] The music rhythm feature information is used to represent the music rhythm of the music segment. For example, syncopation rhythm, triplet rhythm, etc.

[0098] Optionally, when the electronic device determines at least one piece of feature information corresponding to at least one music segment one by one, it may determine at least one piece of feature information through a feature extraction model, or may also determine at least one piece of feature information through a specific feature extraction algorithm. The present disclosure does not limit this.

[0099] S103. The electronic device inputs at least one piece of feature information into a preset model to obtain at least one dance resource corresponding to at least one piece of feature information one by one.

[0100] Among them, the preset model is a model trained according to multiple sample resources to meet the first preset condition and used to determine the dance resource corresponding to the input feature information; a sample resource pair includes a music sample resource and a dance sample resource.

[0101] Specifically, since the preset model is trained based on multiple sample resources to meet a first preset condition and is used to determine the dance resources corresponding to the input feature information, the preset model may include multiple dance resources. In this case, the electronic device may input at least one feature information into the preset model, so as to obtain at least one dance resource corresponding to the at least one feature information one by one.

[0102] Optionally, the first preset condition may be training based on multiple sample resources until a convergence state, that is, the error value output by the model obtained in the last training is less than a first preset threshold, or may be training based on multiple sample resources for a second preset number of times, or may be other training end conditions, which are not limited in this disclosure.

[0103] Optionally, the preset model may be a Hidden Markov Model (HMM). A Hidden Markov Model is a statistical model used to describe a Markov process containing hidden unknown parameters.

[0104] When the preset model is a Hidden Markov Model, when the electronic device inputs at least one feature information into the preset model to obtain at least one dance resource corresponding to the at least one feature information one by one, it performs a search on a first-order Markov model.

[0105] The Hidden Markov Model may include Dance Action Units (CAUs) corresponding to multiple dance resources one by one. Each CAU is a state, and the transition probability between states (i.e., the similarity between dance resources) may use empirical values or the similarity of music style features. The emission probability of each state (i.e., the probability that the dance action in the next dance resource is the same as the dance action in the previous dance resource) is determined by the type of the dance resource corresponding to this state and the similarity of rhythm features. Optionally, when the number of dance resources in the preset model is huge, if the similarity between one feature information and each of the large number of dance resources is determined one by one, it is time-consuming and laborious, and the efficiency is low. In this case, the electronic device may adopt a beam search method to determine the dance resources corresponding to the feature information from a large number of dance resources, improving the search efficiency.

[0106] S104. The electronic device generates a synthetic resource according to at least one dance resource and at least one music segment.

[0107] Among them, the synthetic resource includes the music resource to be synthesized and the dance resource corresponding to the music resource to be synthesized.

[0108] Specifically, after determining at least one dance resource, the electronic device may perform one-to-one synthesis of at least one dance resource and at least one music segment, and then obtain the synthetic resource.

[0109] However, since the dance resources in the synthesized resources are not necessarily continuous, if directly synthesized, it may lead to unreasonable connection of the dance actions in the synthesized resources. Therefore, when generating the synthesized resources according to at least one dance resource and at least one music segment, the electronic device may smooth the discontinuous dance resources, so as to obtain a synthesized resource with reasonable and continuous dance actions.

[0110] Optionally, when the electronic device smooths the discontinuous dance resources, it may determine the dance actions with reasonable connection according to the action completion model, or may also determine the dance actions with reasonable connection through other completion algorithms. The present disclosure does not limit this.

[0111] The technical solutions provided in the above embodiments at least bring the following beneficial effects: As can be seen from S101-S104, when the electronic device obtains the music resources to be synthesized, it may divide the music resources to be synthesized into at least one music segment, and determine at least one feature information corresponding to the at least one music segment one by one. Subsequently, input the at least one feature information into the preset model to obtain at least one dance resource corresponding to the at least one feature information one by one, and generate the synthesized resources according to the at least one dance resource and the at least one music segment. Since the preset model is trained with multiple sample resources to meet the first preset condition and is used to determine the dance resource corresponding to the input feature information, and the synthesized resources include the music resources to be synthesized and the dance resources corresponding to the music resources to be synthesized, therefore, the present disclosure can quickly and accurately determine the dance resources corresponding to the music resources to be synthesized through the preset model, and generate the synthesized resources, without designing a set of dance actions for each piece of music separately, nor collecting the dance resources through dance personnel for simulation, thus improving the synthesis efficiency of the music resources and the dance resources.

[0112] In one embodiment, the feature information includes: music type feature information and music rhythm feature information. Combining Figure 1 , as Figure 2 shown, in the above S102, the method for the electronic device to determine at least one feature information corresponding to the at least one music segment one by one specifically includes:

[0113] S201. The electronic device inputs the at least one music segment into the music type feature extraction model to obtain at least one music type feature information corresponding to the at least one music segment one by one.

[0114] Among them, the music type feature extraction model is trained according to the first sample resources; the first sample resources are the sample resources including music type features in the multiple sample resource pairs.

[0115] Specifically, since the music genre feature extraction model is trained based on the first sample resources, and the first sample resources are the sample resources including music genre features among multiple pairs of sample resources, the music genre feature extraction model can include multiple music genres. In this case, the electronic device can input at least one piece of feature information into the music genre feature extraction model to determine the similarity between each piece of feature information and each music genre.

[0116] Next, for each piece of feature information, the electronic device can select, from the similarities with each music genre, the music genres corresponding to the similarities greater than a preset similarity threshold, and determine them as the music genre feature information corresponding to this piece of feature information.

[0117] Optionally, the preset similarity threshold can be set according to human experience. In practical applications, the electronic device can also select the music genre with the highest similarity and determine it as the music genre feature information corresponding to this piece of feature information.

[0118] S202. The electronic device inputs at least one music segment into the music rhythm feature extraction model to obtain at least one piece of music rhythm feature information corresponding one by one to the at least one music segment.

[0119] Among them, the music rhythm feature extraction model is trained based on the second sample resources; the second sample resources are the sample resources including music rhythm features among multiple pairs of sample resources.

[0120] Specifically, since the music rhythm feature extraction model is trained based on the second sample resources, and the second sample resources are the sample resources including music rhythm features among multiple pairs of sample resources, the music rhythm feature extraction model can include multiple music rhythms. In this case, the electronic device can input at least one piece of feature information into the music rhythm feature extraction model to determine the similarity between each piece of feature information and each music rhythm.

[0121] Next, for each piece of feature information, the electronic device can select, from the similarities with each music rhythm, the music rhythms corresponding to the similarities greater than a preset similarity threshold, and determine them as the music rhythm feature information corresponding to this piece of feature information.

[0122] Optionally, the preset similarity threshold can be set according to human experience. In practical applications, the electronic device can also select the music rhythm with the highest similarity and determine it as the music rhythm feature information corresponding to this piece of feature information.

[0123] It should be noted that the electronic device can execute S201 first and then S202; or execute S202 first and then S201; or execute S201 and S202 simultaneously. The present disclosure does not make any limitation on this.

[0124] The technical solutions provided by the above embodiments at least bring the following beneficial effects: As can be seen from S201-S202, the electronic device can quickly and accurately determine at least one music type feature information corresponding to at least one music segment through the music type feature extraction model, and quickly and accurately determine at least one music rhythm feature information corresponding to at least one music segment through the music rhythm feature extraction model, so as to facilitate determining the corresponding dance resources according to the feature information subsequently, thereby improving the synthesis efficiency of music resources and dance resources.

[0125] In one embodiment, at least one music segment includes: a first music segment and a second music segment consecutive to the first music segment; at least one dance resource includes: a first dance resource corresponding to the first music segment and a second dance resource corresponding to the second music segment; the first dance resource includes: an ending action divided according to the sequence of dance actions of the first dance resource; the second dance resource includes: a starting action divided according to the sequence of dance actions of the second dance resource.

[0126] Combined Figure 1 , such as Figure 3 shown, when the ending action and the starting action are non-consecutive dance actions, in the above S104, the method for the electronic device to generate a composite resource according to at least one dance resource and at least one music segment specifically includes:

[0127] S301. The electronic device inputs the ending action and the starting action into an action completion model to obtain a first action corresponding to the ending action and a second action corresponding to the starting action.

[0128] Among them, the action completion model is trained according to a third sample resource; the third sample resource is a sample resource including dance action features among multiple sample resource pairs.

[0129] The related technology usually establishes a model for associating music and dance actions through a neural network, and then inputs the music to predict the dance actions. This method requires a large amount of music-dance pairing data to be input by relying on motion capture devices in the early stage to ensure that a relatively stable network model can be trained. In addition, the dance actions predicted by the model cannot guarantee that there will be no unreasonable dance actions, and the smoothness of the actions is difficult to guarantee.

[0130] In this application, the first music segment and the second music segment are any two consecutive music segments among at least one music segment. After inputting the feature information of the first music segment and the feature information of the second music segment into a preset model to obtain a first dance resource corresponding to the first music segment and a second dance resource corresponding to the second music segment, the electronic device can extract the dance actions in the first dance resource and the second dance resource.

[0131] Since the first music segment and the second music segment are consecutive music segments, the first dance resource and the second dance resource are also consecutive dance resources. After extracting the dance actions from the first dance resource and the second dance resource, the electronic device can determine whether the ending action of the first dance resource and the starting action of the second dance resource are consecutive dance actions.

[0132] When the ending action and the starting action are non-consecutive dance actions, in order to make the connection of the dance actions more reasonable, the electronic device can input the ending action and the starting action into the action completion model to obtain a first action corresponding to the ending action and a second action corresponding to the starting action.

[0133] Wherein, the first action and the second action are consecutive dance actions.

[0134] Exemplarily, it is assumed that the first dance resource is divided into dance action A and dance action B in the order of the dance actions, and the second dance resource includes dance action C and dance action D. Among them, dance action B is the ending action of the first dance resource, and dance action C is the starting action of the second dance resource.

[0135] When dance action B and dance action C are non-consecutive dance actions, in order to make the connection of the dance actions more reasonable, the electronic device can input dance action B and dance action C into the action completion model to obtain a first action corresponding to dance action B and a second action corresponding to dance action C.

[0136] Optionally, the ending action can be the last action of the first dance resource, or the last multiple actions (such as the last 2 actions, the last 3 actions, etc.). Correspondingly, the starting action can be the first action of the second dance resource, or the first few actions (such as the first 2 actions, the first 3 actions, etc.).

[0137] S302. The electronic device updates the ending action to the first action to obtain an updated first dance resource.

[0138] S303. The electronic device updates the starting action to the second action to obtain an updated second dance resource.

[0139] S304. The electronic device generates a composite resource according to the first music segment, the second music segment, the updated first dance resource, and the updated second dance resource.

[0140] Specifically, after updating the ending action to the first action to obtain the updated first dance resource and updating the starting action to the second action to obtain the updated second dance resource, the electronic device needs to synthesize the first music segment with the updated first dance resource. Correspondingly, the electronic device also needs to synthesize the second music segment with the updated second dance resource. In addition, for other dance resources, the electronic device can perform the same operation to generate a synthesized resource.

[0141] The technical solutions provided in the above embodiments at least bring the following beneficial effects: As can be seen from S301 - S304, the electronic device can quickly and accurately update discontinuous dance actions to continuous dance actions through the action completion model, so as to generate a synthesized resource including smooth and reasonably connected dance resources.

[0142] In one embodiment, as Figure 4 shown, the resource synthesis method further includes:

[0143] S401. The electronic device obtains multiple pairs of sample resources.

[0144] Specifically, in order to train the preset model, the electronic device can obtain a large number of pairs of sample resources.

[0145] Among them, one pair of sample resources includes multiple segments of sample resource pairs; one segment of sample resource pair includes one segment of music sample resource and one segment of dance sample resource.

[0146] Optionally, the pairs of sample resources can be pre - synthesized multimedia resources including music and dance. The multiple pairs of sample resources can include positive pairs of sample resources and negative pairs of sample resources.

[0147] Among them, the positive pair of sample resources can be a multimedia resource in which the type and rhythm of the music match the dance. Correspondingly, the negative pair of sample resources can be a multimedia resource in which the type and rhythm of the music do not match the dance.

[0148] After obtaining multiple pairs of sample resources, the electronic device can divide each pair of sample resources into multiple segments of sample resource pairs according to the method in S101, in which the electronic device divides the music resources to be synthesized according to the musical measures.

[0149] S402. The electronic device determines at least one clustering result of the multiple segments of sample resource pairs according to the multiple segments of sample resource pairs and the clustering algorithm.

[0150] Specifically, after obtaining multiple sample resource pairs and dividing each sample resource pair into multiple sample resource pair segments, in order to train a preset model, the electronic device can determine at least one clustering result of the multiple sample resource pair segments according to the multiple sample resource pair segments and a clustering algorithm, that is, cluster the multiple sample resource pairs according to the types of dance resources and music resources.

[0151] Among them, one clustering result corresponds to one dance resource.

[0152] Optionally, when the electronic device determines at least one clustering result of the multiple sample resource pair segments according to the multiple sample resource pair segments and a clustering algorithm, it can calculate the similarity between each sample resource pair segment and the clustering center in the frequency domain, and select the sample resource pair segments with large similarity differences, so as to determine at least one clustering result of the multiple sample resource pair segments.

[0153] Optionally, the electronic device can also determine at least one clustering result of the multiple sample resource pair segments through a clustering algorithm such as the K-Means clustering algorithm, and the present disclosure does not limit this.

[0154] Optionally, in order to facilitate determining at least one clustering result of the multiple sample resource pair segments, the electronic device can also extract the feature information of the multiple sample resource pair segments through a feature extraction model, and determine at least one clustering result of the multiple sample resource pair segments according to the feature information of the multiple sample resource pair segments and a clustering algorithm.

[0155] S403. The electronic device uses at least one clustering result to train a preset hidden Markov model until a first preset condition is met to obtain a preset model.

[0156] The hidden Markov model (HMM) is a statistical model that describes a Markov process with hidden unknown parameters.

[0157] The hidden Markov model includes at least one state corresponding one-to-one to at least one clustering result.

[0158] Among the above at least one state, the transition probability between states can use empirical values or be determined using the similarity between music type feature information. The emission probability of each state is determined by the type of the sample pair resource corresponding to the clustering result corresponding to the state and the similarity of the rhythm features.

[0159] The technical solutions provided by the above embodiments at least bring the following beneficial effects: As can be seen from S401 - S403, a specific implementation manner for an electronic device to train a preset model is given, so as to subsequently determine a dance resource corresponding to a music resource to be synthesized according to the preset model, thereby improving the synthesis efficiency of the music resource and the dance resource.

[0160] In one embodiment, in combination with Figure 4 , such as Figure 5 shown, after the above S401, the resource synthesis method further includes:

[0161] S501. The electronic device uses multiple sample resources to train a music sample resource segment in the segment for a preset classification task model until a second preset condition is met, so as to obtain a music type feature extraction model.

[0162] Specifically, when training to obtain the music type feature extraction model, since the music type is not related to the dance resource, therefore, the electronic device only needs to use multiple sample resources to train a music sample resource segment in the segment for the preset classification task model until the second preset condition is met, then the music type feature extraction model can be obtained.

[0163] The preset classification task model can be an initial model for common classification tasks in machine learning. Since the model parameters of the initial model are initial values, therefore, multiple sample resources are used to train a music sample resource segment in the segment for the preset classification task model, thereby adjusting the model parameters of the model, and then the music type feature extraction model is obtained.

[0164] Optionally, the preset classification task model can be a binary classification model, a multi - category classification model, etc.

[0165] Optionally, the second preset condition can be to train the preset classification task model to a convergent state, that is, the error value output by the model obtained in the last training is less than the second preset threshold, or it can be to train the preset classification task model to the second preset number of times, or it can be other training end conditions, and the present disclosure does not limit this.

[0166] Optionally, the first preset threshold and the second preset threshold can be the same or different; the first preset number of times and the second preset number of times can be the same or different, and the present disclosure does not limit this.

[0167] The technical solutions provided by the above embodiments at least bring the following beneficial effects: As can be seen from S501, a specific implementation manner for an electronic device to train a music type feature extraction model is given, so as to subsequently determine music type feature information according to the music type feature extraction model, so as to subsequently determine the corresponding dance resource according to the feature information, thereby improving the synthesis efficiency of the music resource and the dance resource.

[0168] In one embodiment, in combination with Figure 5 , such as Figure 6 shown, after the above S401, the resource synthesis method further includes:

[0169] S601. The electronic device obtains rhythm features corresponding to the music beats from the music sample resource segments of each sample resource pair segment among a plurality of sample resource pair segments.

[0170] Optionally, when the electronic device obtains the rhythm features corresponding to the music beats, it can detect the beats per second of the music to find the musical bars, and then obtain the rhythm features corresponding to the music beats according to the musical bars.

[0171] S602. The electronic device obtains motion features corresponding to the music beats from the dance sample resource segments of each sample resource pair segment.

[0172] Optionally, the motion features mainly include the motion rhythm information of the joints of the dance objects in the dance resources.

[0173] Optionally, when the electronic device obtains the motion features corresponding to the music beats, it can detect whether there are motion beat points at these positions according to the rhythm positions in the musical bars (for example, there are 8 rhythm positions in each bar), so as to obtain the motion features corresponding to the music beats.

[0174] S603. The electronic device uses the rhythm features and the motion features to train a preset feature extraction model until it meets the third preset condition to obtain a music rhythm feature extraction model.

[0175] The preset feature extraction model can be an initial model commonly used for feature extraction in machine learning. Since the model parameters of the initial model are initial values, the preset feature extraction model is trained with the rhythm features and the motion features to adjust the model parameters of the model, and then a music rhythm feature extraction model is obtained.

[0176] Optionally, the third preset condition can be to train the preset feature extraction model to a convergence state, that is, the error value output by the model obtained in the last training is less than the third preset threshold, or to train the preset feature extraction model to the third preset number of times, or other training end conditions, which are not limited in this disclosure.

[0177] Optionally, the first preset threshold, the second preset threshold, and the third preset threshold can be the same or different; the first preset number of times, the second preset number of times, and the third preset number of times can be the same or different, which are not limited in this disclosure.

[0178] The technical solutions provided by the above embodiments have at least the following beneficial effects: As can be seen from S601 - S603, a specific implementation method for an electronic device to train a music rhythm feature extraction model is given, so as to subsequently determine music rhythm feature information according to the music rhythm feature extraction model, and further determine corresponding dance resources according to the feature information, thereby improving the synthesis efficiency of music resources and dance resources.

[0179] In one embodiment, the multiple sample resource pairs of segments include: a first sample resource pair of segments and a second sample resource pair of segments; the first sample resource pair of segments includes: a first dance sample resource segment; the second sample resource pair of segments includes: a second dance sample resource segment.

[0180] Combined with Figure 6 , as Figure 7 shown, after S401 above, the resource synthesis method further includes:

[0181] S701. The electronic device obtains a first sub - segment and a second sub - segment.

[0182] Specifically, when training the motion completion model, the electronic device can train the first sample resource pair of segments and the second sample resource pair of segments in the multiple sample resource pairs of segments.

[0183] The first sample resource pair of segments and the second sample resource pair of segments are any two sample resource pairs of segments in the multiple sample resource pairs of segments.

[0184] Since each sample resource pair of segments includes a dance sample resource segment, the electronic device can obtain a first sub - segment and a second sub - segment.

[0185] Wherein, the first sub - segment is at least one sub - segment sorted from back to front in the first dance sample resource segment according to the dance movement sequence; the second sub - segment is at least one sub - segment sorted from front to back in the second dance sample resource segment according to the dance movement sequence.

[0186] Optionally, the number of at least one sub - segment can be determined according to the segment length of the dance sample resource segment. For example, the number of at least one sub - segment can be 1 / 2, 1 / 4, etc. of the segment length of the dance sample resource segment.

[0187] Exemplarily, the first dance sample resource segment is sorted from back to front according to the dance movement sequence and divided into sub - segment 1, sub - segment 2, and sub - segment 3, and the second dance sample resource segment is sorted from front to back according to the dance movement sequence and divided into sub - segment 4, sub - segment 5, and sub - segment 6.

[0188] The preset electronic device needs to obtain the connection point of the first dance sample resource segment and the second dance sample resource segment, 2 sub-segments of the first dance sample resource segment, and 2 sub-segments of the second dance sample resource segment. In this case, the electronic device can obtain the first sub-segment as sub-segment 1 and sub-segment 2. Correspondingly, the electronic device can obtain the second sub-segment as sub-segment 4 and sub-segment 5.

[0189] S702. The electronic device trains an action completion model according to the first sub-segment, the second sub-segment, and a preset algorithm.

[0190] Specifically, after obtaining the first sub-segment and the second sub-segment, the electronic device can extract the dance actions in the first sub-segment and the second sub-segment, and determine whether the ending action of the first sub-segment and the starting action of the second sub-segment are continuous dance actions.

[0191] Optionally, the preset algorithm can be a self-supervised learning algorithm.

[0192] When the ending action of the first sub-segment and the starting action of the second sub-segment are continuous dance actions, the electronic device can train the two sub-segments with continuous dance actions at the ending and starting actions according to the self-supervised learning algorithm.

[0193] Correspondingly, when the ending action of the first sub-segment and the starting action of the second sub-segment are non-continuous dance actions, the electronic device can train other dance sample resource segments except the first sub-segment and the second sub-segment according to the self-supervised learning algorithm, and does not supervise the first sub-segment and the second sub-segment.

[0194] Exemplarily, the electronic device can splice two dance sample resource segments with a length of T (i.e., the first dance sample resource segment and the second dance sample resource segment) together, cut off a part S on the left and right of the splicing point (i.e., the first sub-segment and the second sub-segment), and make a prediction.

[0195] If the two sub-segments (i.e., the first sub-segment and the second sub-segment) are continuous, then a complete prediction is made for the entire 2T sequence (i.e., the first dance sample resource segment and the second dance sample resource segment).

[0196] If the two segments (i.e., the first sub-segment and the second sub-segment) are non-continuous, then during the prediction sequence supervision, random non-supervision is performed in the length from S / 4 to S / 2 on the left and right of the splicing point. In this way, an action completion model can be trained.

[0197] The technical solutions provided by the above embodiments at least bring the following beneficial effects: As can be seen from S701 - S702, a specific implementation method for training an action completion model by an electronic device is given, so as to subsequently update discontinuous dance actions into continuous dance actions according to the action completion model, thereby generating a synthetic resource including dance resources with smooth and reasonable connections.

[0198] It can be understood that in actual implementation, the terminal / server described in the embodiments of the present disclosure may include one or more hardware structures and / or software modules for implementing the foregoing corresponding resource synthesis method, and these execution hardware structures and / or software modules may constitute an electronic device. Those skilled in the art should easily realize that, combining the algorithm steps of each example described in the embodiments disclosed herein, the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0199] Based on such an understanding, the embodiments of the present disclosure also correspondingly provide a resource synthesis device, which can be applied to an electronic device. Figure 8 The structural schematic diagram of the resource synthesis device provided by the embodiments of the present disclosure is shown. As Figure 8 shown, the resource synthesis device may include:

[0200] An acquisition unit 801, a processing unit 802, and a generation unit 803;

[0201] The acquisition unit 801 is used to acquire the music resource to be synthesized; the music resource to be synthesized includes at least one music segment;

[0202] The processing unit 802 is used to determine at least one piece of feature information corresponding to at least one music segment one by one;

[0203] The processing unit 802 is further used to input at least one piece of feature information into a preset model to obtain at least one dance resource corresponding to at least one piece of feature information one by one; the preset model is a model trained based on multiple sample resources to meet a first preset condition and is used to determine the dance resource corresponding to the input feature information; a sample resource pair includes a music sample resource and a dance sample resource;

[0204] The generation unit 803 is used to generate a synthetic resource according to at least one dance resource and at least one music segment; the synthetic resource includes the music resource to be synthesized and the dance resource corresponding to the music resource to be synthesized.

[0205] Optionally, the feature information includes: music type feature information and music rhythm feature information;

[0206] The processing unit 802 is specifically configured to:

[0207] Input at least one music segment into the music type feature extraction model to obtain at least one music type feature information corresponding one-to-one to the at least one music segment; the music type feature extraction model is trained according to the first sample resource; the first sample resource is the sample resource including music type features among multiple sample resource pairs;

[0208] Input at least one music segment into the music rhythm feature extraction model to obtain at least one music rhythm feature information corresponding one-to-one to the at least one music segment; the music rhythm feature extraction model is trained according to the second sample resource; the second sample resource is the sample resource including music rhythm features among multiple sample resource pairs.

[0209] Optionally, the at least one music segment includes: a first music segment and a second music segment consecutive to the first music segment; the at least one dance resource includes: a first dance resource corresponding to the first music segment and a second dance resource corresponding to the second music segment; the first dance resource includes: an ending action divided according to the sequence of dance actions of the first dance resource; the second dance resource includes: a starting action divided according to the sequence of dance actions of the second dance resource;

[0210] When the ending action and the starting action are non-consecutive dance actions, the generating unit 803 is specifically configured to:

[0211] Input the ending action and the starting action into the action completion model to obtain a first action corresponding to the ending action and a second action corresponding to the starting action; the action completion model is trained according to the third sample resource; the third sample resource is the sample resource including dance action features among multiple sample resource pairs; the first action and the second action are consecutive dance actions;

[0212] Update the ending action to the first action to obtain an updated first dance resource;

[0213] Update the starting action to the second action to obtain an updated second dance resource;

[0214] Generate a synthesized resource according to the first music segment, the second music segment, the updated first dance resource, and the updated second dance resource.

[0215] Optionally, the obtaining unit 801 is further configured to obtain a plurality of sample resource pairs; one sample resource pair includes a plurality of sample resource pair segments; one sample resource pair segment includes a music sample resource segment and a dance sample resource segment;

[0216] The processing unit 802 is further configured to determine at least one clustering result of the plurality of sample resource pair segments according to the plurality of sample resource pair segments and a clustering algorithm; one clustering result corresponds to one dance resource;

[0217] The processing unit 802 is further configured to train a preset hidden Markov model with at least one clustering result until a first preset condition is met to obtain a preset model.

[0218] Optionally, the processing unit 802 is further configured to train a preset classification task model with the music sample resource segments in the plurality of sample resource pair segments until a second preset condition is met to obtain a music type feature extraction model.

[0219] Optionally, the processing unit 802 is further configured to obtain rhythm features corresponding to the music beats from the music sample resource segments of each sample resource pair segment in the plurality of sample resource pair segments;

[0220] Obtain motion features corresponding to the music beats from the dance sample resource segments of each sample resource pair segment;

[0221] Train a preset feature extraction model with the rhythm features and the motion features until a third preset condition is met to obtain a music rhythm feature extraction model.

[0222] Optionally, the plurality of sample resource pair segments include: a first sample resource pair segment and a second sample resource pair segment; the first sample resource pair segment includes: a first dance sample resource segment; the second sample resource pair segment includes: a second dance sample resource segment;

[0223] The obtaining unit 801 is further configured to obtain a first sub-segment and a second sub-segment; the first sub-segment is at least one sub-segment sorted from the back to the front in the order of dance movements in the first dance sample resource segment; the second sub-segment is at least one sub-segment sorted from the front to the back in the order of dance movements in the second dance sample resource segment;

[0224] The processing unit 802 is further configured to train an action completion model according to the first sub-segment, the second sub-segment and a preset algorithm.

[0225] As described above, the embodiments of the present disclosure may divide the functional modules of an electronic device according to the above method examples. Among them, the above integrated modules may be implemented in the form of hardware or in the form of software functional modules. In addition, it should be noted that the division of modules in the embodiments of the present disclosure is illustrative, and is only a logical function division. In actual implementation, there may be other division methods. For example, each functional module may be corresponding to each function, or two or more functions may be integrated into one processing module.

[0226] Regarding the resource synthesis device in the above embodiments, the specific ways in which each module performs operations and the beneficial effects thereof have been described in detail in the foregoing method embodiments, and will not be elaborated here.

[0227] The embodiments of the present disclosure also provide a terminal, which may be a user terminal such as a mobile phone or a computer. Figure 9 The structural schematic diagram of the terminal provided by the embodiments of the present disclosure is shown. The terminal may be that the resource synthesis device may include at least one processor 61, a communication bus 62, a memory 63, and at least one communication interface 64.

[0228] The processor 61 may be a central processing unit (CPU), a microprocessing unit, an ASIC, or one or more integrated circuits for controlling the execution of the program of the present disclosure solution. As an example, in combination with Figure 8 , the functions implemented by the processing unit 802 in the electronic device are the same as those implemented by the processor 61 in Figure 9 .

[0229] The communication bus 62 may include a path for transmitting information between the above components.

[0230] The communication interface 64 uses any device such as a transceiver for communicating with other devices or communication networks, such as a server, Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. As an example,

[0231] The memory 63 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited to this. The memory can exist independently and be connected to the processing unit through a bus. The memory can also be integrated with the processing unit.

[0232] Among them, the memory 63 is used to store the application program code for executing the solution of the present disclosure and is controlled by the processor 61 for execution. The processor 61 is used to execute the application program code stored in the memory 63, thereby implementing the functions in the method of the present disclosure.

[0233] In a specific implementation, as an embodiment, the processor 61 may include one or more CPUs, such as Figure 9 CPU0 and CPU1 in

[0234] In a specific implementation, as an embodiment, the terminal may include multiple processors, such as Figure 9 the processor 61 and the processor 65 in

[0235] Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). In a specific implementation, as an embodiment, the terminal may further include an input device 66 and an output device 67. The input device 66 and the output device 67 communicate and can accept user input in various ways. For example, the input device 66 can be a mouse, a keyboard, a touch screen device, or a sensing device, etc. The output device 67 communicates with the processor 61 and can display information in various ways. For example, the output device 61 can be a liquid crystal display (LCD), a light emitting diode (LED) display device, etc.

[0236] Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the terminal, and it may include more or fewer components than those shown in the figure, or combine some components, or adopt a different component arrangement.

[0237] The embodiments of the present disclosure also provide a server. Figure 10 The schematic structural diagram of the server provided by the embodiments of the present disclosure is shown. This server may be a resource synthesis device. This server may vary greatly due to different configurations or performances, and may include one or more processors 71 and one or more memories 72. Among them, at least one instruction is stored in the memory 72, and at least one instruction is loaded and executed by the processor 71 to implement the resource synthesis method provided by each of the above method embodiments. Of course, this server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. This server may also include other components for implementing the functions of the device, which will not be elaborated here.

[0238] The present disclosure also provides a computer-readable storage medium including instructions. When the instructions in the computer-readable storage medium are executed by the processor of a computer device, the computer can execute the resource synthesis method provided by the above-mentioned embodiments. For example, the computer-readable storage medium may be a memory 63 including instructions, and the above instructions may be executed by the processor 61 of the terminal to complete the above method. Another example, the computer-readable storage medium may be a memory 72 including instructions, and the above instructions may be executed by the processor 71 of the server to complete the above method. Optionally, the computer-readable storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc.

[0239] The present disclosure also provides a computer program product. The computer program product includes computer instructions. When the computer instructions run on an electronic device, the electronic device is enabled to execute the above-mentioned Figures 1-7 resource synthesis method shown in any of the drawings.

[0240] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present disclosure. The present disclosure aims to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0241] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A resource synthesis method, characterized in that, Including: Obtain music resources to be synthesized; the music resources to be synthesized include at least one music segment; Determine at least one feature information corresponding one-to-one to the at least one music segment; The feature information includes: music type feature information and music rhythm feature information; the determining at least one feature information corresponding one-to-one to the at least one music segment includes: inputting the at least one music segment into a music type feature extraction model to obtain at least one music type feature information corresponding one-to-one to the at least one music segment, and inputting the at least one music segment into a music rhythm feature extraction model to obtain at least one music rhythm feature information corresponding one-to-one to the at least one music segment; Input the at least one feature information into a preset model to obtain at least one dance resource corresponding one-to-one to the at least one feature information; the preset model is a model trained according to multiple sample resources to meet a first preset condition for determining the dance resource corresponding to the input feature information; a sample resource pair includes a music sample resource and a dance sample resource; Generate a synthesized resource according to the at least one dance resource and the at least one music segment; the synthesized resource includes the music resources to be synthesized and the dance resource corresponding to the music resources to be synthesized.

2. The resource synthesis method according to claim 1, wherein The music type feature extraction model is trained according to a first sample resource; the first sample resource is a sample resource including music type features among the multiple sample resource pairs; The music rhythm feature extraction model is trained according to a second sample resource; The second sample resource is a sample resource including music rhythm features among the multiple sample resource pairs.

3. The resource synthesis method according to claim 1, wherein The at least one music segment includes: a first music segment and a second music segment consecutive to the first music segment; the at least one dance resource includes: a first dance resource corresponding to the first music segment and a second dance resource corresponding to the second music segment; the first dance resource includes: an ending action divided according to the sequence of dance actions of the first dance resource; the second dance resource includes: a starting action divided according to the sequence of dance actions of the second dance resource; When the ending action and the starting action are non-consecutive dance actions, the generating a synthesized resource according to the at least one dance resource and the at least one music segment includes: Input the ending action and the starting action into an action completion model to obtain a first action corresponding to the ending action and a second action corresponding to the starting action; the action completion model is trained according to a third sample resource; the third sample resource is a sample resource including dance action features among the multiple sample resource pairs; the first action and the second action are consecutive dance actions; Update the ending action to the first action to obtain an updated first dance resource; Update the starting action to the second action to obtain an updated second dance resource; Generate the composite resource according to the first music segment, the second music segment, the updated first dance resource, and the updated second dance resource.

4. The resource synthesis method according to claim 1, wherein Further comprising: Obtain the multiple sample resource pairs; one sample resource pair includes multiple sample resource pair segments; One sample resource pair segment includes one music sample resource segment and one dance sample resource segment; Determine at least one clustering result of the multiple sample resource pair segments according to the multiple sample resource pair segments and a clustering algorithm; one clustering result corresponds to one dance resource; Train a preset hidden Markov model with the at least one clustering result until a first preset condition is met to obtain the preset model.

5. The resource synthesis method according to claim 4, wherein After obtaining the multiple sample resource pairs, further comprising: Train a preset classification task model with the music sample resource segments in the multiple sample resource pair segments until a second preset condition is met to obtain a music type feature extraction model.

6. The resource synthesis method according to claim 4, characterized in that After obtaining the multiple sample resource pairs, further comprising: Obtain rhythm features corresponding to music beats from the music sample resource segments of each sample resource pair segment in the multiple sample resource pair segments; Obtain action features corresponding to the music beats from the dance sample resource segments of each sample resource pair segment; Train a preset feature extraction model with the rhythm features and the action features until a third preset condition is met to obtain a music rhythm feature extraction model.

7. The resource synthesis method according to claim 4, wherein The multiple sample resource pair segments include: a first sample resource pair segment and a second sample resource pair segment; the first sample resource pair segment includes: a first dance sample resource segment; the second sample resource pair segment includes: a second dance sample resource segment; After obtaining the multiple sample resource pairs, further comprising: Obtain a first sub-segment and a second sub-segment; the first sub-segment is at least one sub-segment sorted from back to front in the order of dance actions in the first dance sample resource segment; the second sub-segment is at least one sub-segment sorted from front to back in the order of dance actions in the second dance sample resource segment; Train an action completion model according to the first sub-segment, the second sub-segment, and a preset algorithm.

8. A resource synthesis device, characterized in that, Comprising: An acquisition unit, a processing unit, and a generation unit; The acquisition unit is configured to acquire a music resource to be synthesized; the music resource to be synthesized includes at least one music segment; The processing unit is configured to determine at least one feature information corresponding one-to-one to the at least one music segment; The feature information includes: music type feature information and music rhythm feature information; specifically, the processing unit is configured to: input the at least one music segment into the music type feature extraction model to obtain at least one music type feature information corresponding one-to-one to the at least one music segment; input the at least one music segment into the music rhythm feature extraction model to obtain at least one music rhythm feature information corresponding one-to-one to the at least one music segment; The processing unit is further configured to input the at least one feature information into a preset model to obtain at least one dance resource corresponding to the at least one feature information one by one; the preset model is a model trained according to a plurality of sample resources to meet a first preset condition and is used to determine the dance resource corresponding to the input feature information; a pair of sample resources includes a music sample resource and a dance sample resource. The generating unit is configured to generate a synthesized resource according to the at least one dance resource and the at least one music segment; the synthesized resource includes the music resource to be synthesized and the dance resource corresponding to the music resource to be synthesized.

9. The resource synthesis device according to claim 8, wherein, The music type feature extraction model is trained according to a first sample resource; the first sample resource is a sample resource including music type features among the plurality of pairs of sample resources. The music rhythm feature extraction model is trained according to a second sample resource. The second sample resource is a sample resource including music rhythm features among the plurality of pairs of sample resources.

10. The resource synthesis device according to claim 8, wherein, The at least one music segment includes: a first music segment and a second music segment consecutive to the first music segment; the at least one dance resource includes: a first dance resource corresponding to the first music segment and a second dance resource corresponding to the second music segment; the first dance resource includes: an ending action divided according to the sequence of dance actions of the first dance resource; the second dance resource includes: a starting action divided according to the sequence of dance actions of the second dance resource. When the ending action and the starting action are non - consecutive dance actions, the generating unit is specifically configured to: Input the ending action and the starting action into an action completion model to obtain a first action corresponding to the ending action and a second action corresponding to the starting action; the action completion model is trained according to a third sample resource; the third sample resource is a sample resource including dance action features among the plurality of pairs of sample resources; the first action and the second action are consecutive dance actions. Update the ending action to the first action to obtain an updated first dance resource. Update the starting action to the second action to obtain an updated second dance resource. Generate the synthesized resource according to the first music segment, the second music segment, the updated first dance resource, and the updated second dance resource.

11. The resource synthesis device according to claim 8, wherein The obtaining unit is further configured to obtain the plurality of pairs of sample resources; a pair of sample resources includes a plurality of segments of sample resources; a segment of sample resources includes a music sample resource segment and a dance sample resource segment. The processing unit is further configured to determine at least one clustering result of the plurality of segments of sample resources according to the plurality of segments of sample resources and a clustering algorithm; one clustering result corresponds to one dance resource. The processing unit is further configured to use the at least one clustering result to train a preset Hidden Markov Model until the first preset condition is satisfied, so as to obtain the preset model.

12. The resource synthesis device according to claim 11, wherein The processing unit is further configured to use the multiple sample resource pairs in the music sample resource segments of the segment to train a preset classification task model until the second preset condition is satisfied, so as to obtain a music type feature extraction model.

13. The resource synthesis device according to claim 11, wherein The processing unit is further configured to obtain rhythm features corresponding to the music beats from the music sample resource segments of each sample resource pair segment in the multiple sample resource pairs; obtain action features corresponding to the music beats from the dance sample resource segments of each sample resource pair segment; use the rhythm features and the action features to train a preset feature extraction model until the third preset condition is satisfied, so as to obtain a music rhythm feature extraction model.

14. The resource synthesis device according to claim 11, wherein The multiple sample resource pairs include: a first sample resource pair segment and a second sample resource pair segment; the first sample resource pair segment includes: a first dance sample resource segment; the second sample resource pair segment includes: a second dance sample resource segment; The obtaining unit is further configured to obtain a first sub-segment and a second sub-segment; the first sub-segment is at least one sub-segment sorted from back to front in the order of dance actions in the first dance sample resource segment; the second sub-segment is at least one sub-segment sorted from front to back in the order of dance actions in the second dance sample resource segment; The processing unit is further configured to train an action completion model according to the first sub-segment, the second sub-segment and a preset algorithm.

15. An electronic device, characterized in that, The electronic device includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the resource synthesis method according to any one of claims 1-7.

16. A computer-readable storage medium having instructions stored thereon, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the resource synthesis method according to any one of claims 1-7.

17. A computer program product comprising instructions, characterized in that, When the instructions run on the electronic device, the electronic device executes the resource synthesis method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Dance motion synthesis method, device and equipment and storage medium

    CN110992449A

  • Display device and music recommendation method

    CN112788376A