Method and apparatus for generating music, electronic device and program product

By analyzing the elements of input music and generating music using existing audio materials, the problem of high computing resources and low sound quality in existing technologies is solved, and efficient and logically coherent new music generation is achieved.

CN120708566APending Publication Date: 2025-09-26BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410295211.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing music generation methods require high computing resources when directly generating music through deep neural networks. The generated music has low sound quality and is logically incoherent, making it difficult to create high-quality new music related to the input music.

Method used

By analyzing the musical elements of the input music, utilizing existing audio materials and a pool of randomly mixed audio tracks, combined with music information retrieval models and sound source separation technology, high-quality, logically coherent new music related to the input music is generated.

Benefits of technology

The generated music is of high quality and strong logic, and can adapt to a variety of music types and styles, which improves generation efficiency and reduces computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708566A_ABST
    Figure CN120708566A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for generating music, electronic equipment and a program product. The method includes determining music elements of the input music, wherein the music elements of the input music include at least one of a structure, a beat, a genre, a music tone, a harmony, and an emotion. The method further includes determining a set of track groups of the output music based on the music elements of the input music, where the track groups are combinations of tracks of the music. The method further includes generating an output music corresponding to the input music based on the set of track groups of the output music. According to the embodiment of the invention, the works related to the input music can be generated, the music quality and logic of the generated output music are ensured, and brand new music can be created by combining and recombining different audio track groups. According to the method for generating the output music by directly analyzing and utilizing the input music and the audio track group set related to the input music, the requirement of the whole process for computing resources is also reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to the field of computer technology, and more particularly, to a method, apparatus, electronic device, and program product for generating music. Background Art

[0002] Automatic music generation technology has now penetrated multiple fields, not only occupying a place in music creation and music recommendation, but also playing an important role in fields such as dance art, greatly enriching the musical experience of modern life. In terms of music creation, automatic music generation technology can quickly generate diverse musical ideas through advanced algorithms and models, helping creators easily explore and experiment with different musical styles and structures. Creators are no longer limited by traditional creative methods and can use this technology to create creative music works in a more efficient and flexible way.

[0003] In the field of music recommendation, automatic music generation technology has also shown great potential. In addition, in terms of music processing, this technology can also be used for post-processing such as noise reduction and reverberation to improve the quality and effect of music. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, and program product for generating music.

[0005] According to a first disclosed aspect, a method for generating music is provided. The method includes determining musical elements of input music, wherein the musical elements of the input music include at least one of structure, tempo, genre, instrumentation timbre, harmony, and emotion. The method also includes determining a set of track groups for output music based on the musical elements of the input music, wherein a track group is a combination of musical tracks. Furthermore, the method also includes generating output music corresponding to the input music based on the set of track groups for the output music.

[0006] In a second aspect of the disclosure, a device for generating music is provided. The device includes a music element determination module configured to determine the music elements of input music, where the music elements of the input music include at least one of structure, tempo, genre, instrumentation timbre, harmony, and emotion. The device also includes a track group set determination module configured to determine a track group set for output music based on the music elements of the input music, where a track group is a combination of music tracks. Furthermore, the device also includes an output music generation module configured to generate output music corresponding to the input music based on the track group set for output music.

[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory coupled to the processor, wherein the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device executes the method according to the first aspect.

[0008] In a fourth aspect of the present disclosure, a computer program product is provided, wherein a computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to the first aspect.

[0009] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0011] Figure 1 A schematic diagram illustrating an example environment in which some embodiments of the present disclosure may be implemented;

[0012] Figure 2 A flowchart illustrating a method for generating music according to some embodiments of the present disclosure is shown;

[0013] Figure 3 A schematic diagram illustrating a process for generating music according to some embodiments of the present disclosure is shown;

[0014] Figure 4A Schematic diagrams illustrating track groups, track group sets, and track group set pools according to some embodiments of the present disclosure;

[0015] Figure 4B A schematic diagram showing a set of audio track groups and mixed audio tracks according to some embodiments of the present disclosure is shown;

[0016] Figure 4C A schematic diagram illustrating a pool of mixed audio tracks according to some embodiments of the present disclosure is shown;

[0017] Figure 5 A schematic diagram illustrating a module for generating music according to some embodiments of the present disclosure is shown;

[0018] Figure 6 A block diagram illustrating an apparatus for generating music according to some embodiments of the present disclosure; and

[0019] Figure 7A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0020] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION

[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, upon receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0027] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects, unless explicitly stated otherwise. Other explicit and implicit definitions may also be included below.

[0028] In the field of music generation, a related approach is to use deep neural networks to directly generate music, for example by generating an audio spectrogram and then reconstructing a waveform from it, or directly generating an audio waveform. Another related approach also uses deep neural networks for music generation, but first generates a symbolic representation (such as Midi, sheet music), and then uses other modules to render audio waveforms from these symbolic representations. Both methods fail to directly use existing audio to create new music, and both methods require a large amount of computing resources to complete the task of generating music. In addition, the sound quality of the music generated by these two methods is not high, and there is no guarantee that the musical logic of the generated music is coherent.

[0029] In an embodiment of the present disclosure, by analyzing the musical elements of the input music (such as musical structure, etc.) and directly utilizing existing audio materials related to the input music to create new music that is related to the input music in terms of musical elements, the music related to the input music generated in this way is coherent and logical high-quality music, and this method does not require additional training, which greatly improves efficiency, especially when generating unseen music types (such as unknown instruments, styles or rhythm patterns).

[0030] Figure 1 1 is a schematic diagram of an example environment 100 in which some embodiments of the present disclosure may be implemented. The schematic diagram of the example environment 100 is for illustrative purposes only and is not intended to limit the present invention. Figure 1 As shown, an output music can be obtained based on the input music. At 110, music input, music related to the music to be output can be input into the system, and this music can be pure music. For example, if you want to generate pure music with piano sound, you can input pure music of piano into the system.

[0031] Continue to refer Figure 1At 120, music analysis is performed on the input music. After receiving the input music, music analysis can be performed on the input music in order to generate output music related to the input music. In some embodiments, the input music can be analyzed using a music information retrieval model (MIR) to obtain musical elements of the input music, such as genre 121, structure 122, beat 123, instrumentation 124, harmony 125, emotion 126, or other musical elements such as tempo and theme.

[0032] Continue to refer Figure 1 After obtaining the music element information of the input music, the randomly mixed audio track pool 132 can be called to determine the track group set of the output music at 130. The tracks, track groups and track group sets in the track group are all audio materials of the music. Each track group in the track group set is an independent component of the original song and can be processed separately without affecting other elements. For example, a song may be divided into a vocal track group, a guitar track group, a drum track group, etc. In some embodiments, the randomly mixed audio track pool is obtained from the track group set. In some embodiments, the audio tracks in the randomly mixed audio track pool are different. In some embodiments, the track groups in the track group set pool are of the same style and are harmoniously combined. In some embodiments, these track group set pools are pre-defined and can be obtained from music producers or through sound source separation technology.

[0033] Continue to refer Figure 1 , at 130, a set of track groups for output music is determined. In some embodiments, the set of track groups for output music can be determined in the randomly mixed audio track pool based on the musical elements of the input music. For example, the set of track groups for output music can be determined in the randomly mixed audio track pool 132 based on the genre 121 of the input music. In some embodiments, music analysis can be performed on the track groups in the randomly mixed audio track pool to obtain the musical elements of each track group, and the set of track groups for output music can be determined in the randomly mixed audio track pool based on the musical elements of each track group and the musical elements of the input music. At 142, output music related to the input music can be obtained. For example, when the genre 121 of the input music is selected as a reference to generate output music, the genre of the output music is related to the genre of the input music.

[0034] In the embodiments of the present disclosure, an in-depth analysis is performed on the musical elements of the input music, including but not limited to music genre 121, structure 122, music beat 123, orchestration 124, harmony 125, emotion 126, etc. By fully understanding and utilizing these elements and incorporating existing randomly mixed audio tracks related to the input music, i.e., audio materials, new music is created that is closely connected to the input music and full of logic. This music generation method based on music element analysis and the use of existing audio materials not only generates high-quality and logical music, but also demonstrates its strong adaptability and efficiency when faced with a variety of music types and styles.

[0035] It should be understood that the architecture and functions in the example environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. The embodiments of the present disclosure may also be applied to other environments with different structures and / or functions.

[0036] The following will be combined Figures 2 to 7 The process according to the embodiment of the present disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.

[0037] Figure 2 A flowchart of a method 200 for generating music according to some embodiments of the present disclosure is shown. In box 202, the musical elements of the input music are determined, wherein the musical elements of the input music include at least one of structure, beat, genre, and instrumentation timbre. In some embodiments, the input music can be analyzed by multiple MIR models or a unified multi-task basic MIR model to obtain the musical elements of the input music, such as genre 121, structure 122, beat 123 and instrumentation 124, harmony 125, emotion 126 or other musical elements such as beat, speed, theme, etc. Alternatively, the musical elements of the input music can also be obtained by predetermined rules.

[0038] At block 204, a track group set for the output music is determined based on the musical elements of the input music, where a track group is a combination of tracks of the music. In some embodiments, the track group set 130 for the output music can be determined based on the genre 121 of the input music. In some embodiments, the track group set 130 for the output music can be determined based on the instrumentation 124 of the input music. In some embodiments, a randomly mixed audio track pool can be used to determine the track group set 130 for the input music. In some embodiments, the randomly mixed audio track pool is obtained from a pre-constructed track group set.

[0039] At block 206, output music corresponding to the input music is generated based on the set of output music track groups. In some embodiments, the output music is similar to the input music in relevant musical elements, such as genre 121 or structure 122, tempo 123, instrumentation 124, harmony 125, and emotion 126.

[0040] In the embodiments of the present disclosure, various key elements of the input music are analyzed, including but not limited to its unique musical structure, beat, genre, instrumentation timbre, etc. By understanding and applying these elements, it is possible to cleverly combine existing audio materials to determine a set of audio track groups that are closely related to the input music, thereby creating new music that is related to the original input music and full of internal logic. This music generation method that combines the analysis of musical elements with the use of existing audio materials ensures that the generated music is not only of excellent quality and rigorous logic, but also demonstrates its excellent adaptability and efficiency when dealing with a variety of music types and styles.

[0041] Figure 3 FIG2 is a schematic diagram of a process 300 for generating music according to some embodiments of the present disclosure. Figure 3 , music is input at 301. In some embodiments, the input music is pure music. At 302, the input music is analyzed. In some embodiments, the music elements of the input music can be analyzed by multiple MIR models. These music elements can be information such as the structure of the music, the time point when the sound starts in the music, the beat position of the music, the tonality estimation of the music, the chords of the music, the characteristics of the music genre, and the characteristics of the instrument factors. In some embodiments, the music elements of the input music can also be analyzed by a unified multi-task basic MIR model. The Music Information Retrieval Model (MIR) is a technology for automatically analyzing and extracting various information and features in music. By analyzing the input music and extracting various music elements through MIR technology, the understanding of the input music can be enhanced.

[0042] like Figure 3 As shown, a random mixed audio track pool can be obtained in the track group pool 304. In some embodiments, these track group pools can be obtained from music producers or can be obtained using sound source separation technology (MSS). Figure 4A , there are many track group sets in the track group set pool 410A, such as track group set 402A, track group set 404A and track group set 406A. Figure 4AIn track group set 402A, there are a large number of track groups, such as track group 412A, track group 414A and track group 426A. These track groups are of the same style and are harmoniously combined. Track groups are used to separate the different elements of a song (such as vocals, instruments, etc.) into different tracks, so as to facilitate more sophisticated mixing or editing. For example, a song may be divided into a vocal track group, a guitar track group, a drum track group, etc. Each track group is an independent component of the original song and can be processed separately without affecting other elements. In some embodiments, these track groups 412A, track group 414A and track group 426A can be arranged in a certain order in the track group set according to length or instrument attributes. In some embodiments, the track groups in the track group set can be divided into track group segments of different lengths based on the structural information of the track groups extracted by the MIR model.

[0043] Continue to refer Figure 3 , at 305, the track groups within the track group set are randomly mixed to obtain random mixed audio tracks. A mixed audio track refers to an independent audio stream or track that can contain a mixture of one or more instruments or sound elements. For example, a song may include a mixed track containing all instruments and vocals, as well as separate tracks for guitar, drums, bass, etc. Each track is typically an independent audio file that can be edited and processed in audio editing software. A song may be split into multiple mixed audio tracks, and each audio track may be further split into multiple track groups. Figure 4B As shown, mixed audio track 404B can be mixed by track group 412B, track group 414B and track group 416B.In certain embodiments, the track group in the track group set pool is the track group with audio content that has been screened.For example, the activation ratio of the track group can be calculated by audio event detection (AED), using the track group with an activation ratio greater than a threshold value.In this way, only the track group with outstanding sound content can be randomly mixed, thereby the quantity of the mixed audio track can be reduced and the same random mixed audio track can be avoided in the same random mixed audio track pool.In certain embodiments, the track group can be randomly selected according to the beat tracking result, i.e., the beat position, and the beat synchronization is utilized to mix the track group.By combining different track groups, more abundant and complex music effects can be created, enhancing the expressiveness and the appeal of music.

[0044] Continue to refer Figure 3 At 306, a pool of randomly mixed audio tracks is determined. Figure 4CAs shown, there are a large number of mixed audio tracks in the mixed audio track pool 410C, and these mixed audio tracks are arranged according to track group sets, for example, according to track group set 402C, track group set 404C and track group set 406C.

[0045] At 303, track set selection processing is performed. After the track set selection process, a selected track set can be obtained at 307. In some embodiments, a pool of randomly mixed audio tracks is called based on the musical elements of the input music, and a track set corresponding to the output music is selected from it. For example, this can be mood, theme, or timbre. By considering various musical elements such as genre, mood, theme, harmony, and emotion, more personalized music can be generated for the user.

[0046] In some embodiments, the set of track groups corresponding to the output music can be determined based on the genre of the input music. For example, a K-nearest neighbor search (KNN) is performed in the genre feature space to select a suitable set of track groups for the input music, the input music is divided into multiple music segments according to the music structure, the genre feature vectors of the multiple music segments are extracted, and the genre features of the mixed audio tracks in the random mixed audio track pool 410C are also extracted. KNN is used to recall multiple combinations in the random mixed track audio track pool 410C. For example, if there are N music segments, N combinations are recalled, each of which has K mixed audio tracks. The mixed audio tracks in these combinations are all related to the music segments, and these mixed audio tracks are different from each other. Using KNN to search in the genre feature vector space can more accurately find a set of track groups similar to the input music, thereby improving matching accuracy.

[0047] Next, these N combinations are voted on and ranked, and the track set with the most recalled mixed audio tracks is selected as the track set corresponding to the output music. For example, if Track Set A recalls 19 mixed audio tracks related to a music clip, and Track Set B recalls 20 mixed audio tracks related to a music clip, and Track Set B is ranked higher than Track Set A, then Track Set B is selected as the track set corresponding to the input music. Through this voting and ranking process, the most suitable track set can be selected as candidate material, thereby optimizing generation efficiency.

[0048] In some embodiments, the tempo and starting point information of each music segment may also be considered during the ranking of these N combinations. For example, a comparison is performed to search for the optimal starting point density match for each music segment under different tempo (or BPM) multiple combinations. During this matching phase, a penalty factor is calculated for different time stretch ratios to penalize the selection of large time stretch ratios. This penalty factor is applied to the ranking process, with the larger the penalty factor, the lower the ranking. This is because choosing a large time stretch ratio may cause the music to sound unnatural or distorted. Time stretching is a technique that changes the playback speed of music without changing its pitch. For example, if a song has a BPM of 120, time stretching can make it sound like a song with a BPM of 100, but the pitch of each note remains unchanged. For example, audio within a certain BPM range (e.g., a stretch ratio between 0.7 and 1.3) can be stretched and beat-matched without changing the pitch. Starting point density refers to the number of starting points within each measure. BPM is a key characteristic of music, determining its tempo. By considering musical elements such as mode and rhythm, more harmonious and consistent output music can be generated.

[0049] Continue to refer Figure 3 At 309, the structure of the output music is planned. After the track set is selected at 307, the structure of the output music is determined at 309. In some embodiments, the structure of the output music can be determined based on the music elements obtained according to the MIR model and the selected track set, that is, the number of bars corresponding to each part of the output music and the music segment should be calculated. In some embodiments, the number of music segments of the output music is the same as or close to the number of music segments of the input music. In some embodiments, if the optimal time stretch ratio has been calculated in the track set selection process 303, the number of bars can be determined based on the optimal time stretch ratio. If the optimal time stretch ratio has not been calculated, the speed ratio between the selected track set and the input music can be directly used to calculate the number of bars required for each music segment (portion) of the output music. This makes it easy to adjust the speed of the output music to match the rhythm of the input music, thereby outputting music with the same music duration as the input music. In some embodiments, during this process, the starting time point of the output music and the duration of the output music audio content can also be calculated, thereby paving the way for the subsequent generation of the output music. By preparing information such as the start timestamp, audio content duration, and audio duration, necessary input can be provided for subsequent processing, thereby simplifying the processing flow and improving overall efficiency.

[0050] Continue to refer Figure 3At 308, a subset of the randomly mixed audio track pool is selected, where the mixed audio tracks in this subset are composed of the mixed audio tracks within the selected track group set 303. In some embodiments, this subset is a subset of the randomly mixed audio track pool, where the mixed audio tracks in this subset are the set of mixed audio tracks corresponding to the selected track group set 307.

[0051] Continue to refer Figure 3 At 310, a random mix of audio tracks is selected from the selected track set for use in the music arrangement. This means that the mixed audio track that best matches the music arrangement 311 is again selected from the already selected subset pool. In some embodiments, the mixed audio track that best matches the music arrangement can be selected based on the musical elements of the input music. For example, instrument timbre can be used to find the track set that best matches each short segment of the input music for the output music in the instrument timbre feature space. For example, a KNN search can be performed in the timbre feature space to find the mixed audio track that best matches the input music. In some embodiments, during the KNN search, the length of the mixed audio track and the number of bars required for each short segment of the music can be compared to ensure temporal matching. In some embodiments, the harmony or melody of the mixed audio track pool can be compared to that of the input music, and the mixed audio track pool with a closer harmony or melody can be selected to ensure that the output music is harmonically or melodically consistent with the input music. In some embodiments, if a recalled mixed audio track does not meet the aforementioned constraints, a penalty factor can be calculated, and the ranking of the mixed audio track pool can be determined based on the penalty factor during the sorting process. For example, if the selection of a mixed audio track results in a shorter mixed audio content, a higher penalty factor can be given to reduce this deviation from the constraint in the subsequent selection process. This is because if the length is too short (less than 4 bars), increasing the length by looping and splicing will create a strong sense of mechanical repetition, which should be avoided. When the mixed audio is longer, you can cut off audio clips to make the length consistent. Similarly, if the selected mixed audio track does not match the paragraph type or timbre characteristics, a corresponding penalty factor can also be given.

[0052] Continue to refer Figure 3, at 312, audio is generated and the generated audio is post-processed. After the music arrangement 311, the corresponding track group will be combined according to the arrangement structure of each music segment, that is, the required number of bars for each music segment will be considered to combine the corresponding track group. When the audio length of a certain music segment of the output music is not enough to meet the required number of bars, the audio length can be made to meet the requirements by looping the partial audio of this music segment. In some embodiments, the audio can be guaranteed to have a certain length by copying a certain segment of the audio and appending it to the end of this audio.

[0053] In some embodiments, after the audio of each small segment of each output music is generated and expanded, the audio can be post-processed to ensure the continuity of the output audio. In some embodiments, the audio can be filled with blanks. For example, a specific background sound or silence can be added at the beginning of the audio. In some embodiments, the audio can be time-stretched to change the playback speed of the audio without changing the pitch of the audio, so as to ensure that the output audio matches the rhythm of the input music. In some embodiments, the length of the audio can also be trimmed. If the length of the generated audio exceeds the required length, it can be pruned to ensure that it meets the requirements. In some embodiments, the audio can also be dynamically compressed, that is, the difference between the maximum and minimum volume in the audio is reduced.

[0054] Continue to refer Figure 3 , at 313, music related to the input music is output. Through the above process, a complete, coherent and high-quality music piece can be generated. The music piece is composed of multiple audio parts, each of which is carefully selected and post-processed so that the output music is related to the input music.

[0055] Figure 4A Schematic diagram of a track group, a track group set, and a track group set pool 400A is shown in some embodiments of the present disclosure. Figure 4A Track group set pool 410A is composed of multiple track group sets, such as track group set 402A, track group set 404A, and track group set 406A. In some embodiments, the track group set pool is composed of multiple track group sets.

[0056] Continue to refer Figure 4A Track group set 402A is composed of multiple track groups such as track group 412A, track group 414A, and track group 416A. In some embodiments, a track group set includes at least one track group.

[0057] Figure 4B FIG2 shows a schematic diagram of a track group set and a mixed audio track 400B according to some embodiments of the present disclosure. Figure 4BMixed audio track 404B is composed of a mix of multiple track groups, including track group 412B, track group 414B, and track group 416B. In some embodiments, mixed audio track 404B is composed of a random mix of at least two track groups. In some embodiments, the multiple track groups that make up mixed audio track 404B are pre-processed. For example, audio event detection (AED) technology can be applied to the track groups, and the activation ratio of the track groups can be calculated. Only track groups with activation ratios greater than a threshold are used for random mixing. The advantage of this is that only track groups with significant content are randomly mixed, thereby reducing the number of mixed audio tracks and avoiding the appearance of similar mixed audio tracks in the track group collection pool. In some embodiments, track groups can be randomly selected and beat-synchronized based on beat tracking results before being mixed together. Beat tracking is a technique in music analysis that is used to automatically detect beats or rhythms in music. Beat tracking can obtain the exact position and time of each beat in the music. Beat synchronization is the process of ensuring that the selected tracks match the beat of the original music. By adjusting the speed of the tracks, time stretching or other methods, they can be perfectly aligned with the beat of the original music. By applying AED technology, it is ensured that only groups of tracks with significant content are used for random mixing, thereby improving the quality of the final generated music.

[0058] Continue to refer Figure 4B In the track group set 410B, there are multiple mixed audio tracks such as mixed audio track 402B, mixed audio track 404B and mixed audio track 406B arranged in sequence. In some embodiments, these mixed audio tracks are different.

[0059] Figure 4C Schematic diagram of a mixed audio track pool 400C according to some embodiments of the present disclosure is shown. Figure 4C The mixed audio track pool 410C is composed of multiple mixed audio tracks, which are organized and arranged in a certain order, for example, according to the category of the audio track group set. Figure 4C The same is true for the mixed audio tracks arranged according to the categories of track group set 402C, track group set 404C and track group set 406C.

[0060] In this way, different groups of tracks can be quickly accessed and combined, thereby speeding up the music creation process. The ability to randomly mix different groups of tracks can make each generated mixed audio track unique, increasing the diversity and freshness of the music.

[0061] Figure 5 FIG2 shows a schematic diagram of a module for generating music according to some embodiments of the present disclosure. Figure 5Music understanding 510, selecting a track set 520, and music generation 530 are three different functional components of music generation. Music understanding 510 primarily involves music analysis 512, which can be used to extract musical elements 514, such as the genre 502, structure 506, and other musical features or elements 504. In some embodiments, the input music can be divided into multiple musical segments based on the structure of the input music at 516, so that each musical segment can be input into a music information retrieval model to obtain the musical elements of each segment.

[0062] Continue to refer Figure 5 , a pool of randomly mixed audio tracks may be called at 520. In some embodiments, the track groups within the track group set may be randomly mixed to obtain randomly mixed audio tracks. Figure 4B As shown, mixed audio track 404B can be mixed by track group 412B, track group 414B and track group 416B. Figure 4C As described above, the mixed audio track pool 410C contains a large number of mixed audio tracks, which are arranged according to track group sets, such as track group set 402C, track group set 404C, and track group set 406C. In some embodiments, the track groups in the track group set pool are screened track groups containing audio content. In some embodiments, track groups can be randomly selected based on beat tracking results, i.e., beat positions, and mixed using beat synchronization. By combining different track groups, richer and more complex musical effects can be created, enhancing the expressiveness and appeal of the music. At 520, by calling the randomly mixed audio track pool and referencing the musical elements of the input music obtained based on the music understanding 520, a track group set related to or corresponding to the output music can be obtained. For example, relevant track group sets can be first determined from the already constructed randomly mixed audio track pool based on the genre of the input music. Subsequently, a set of randomly mixed audio tracks most suitable for music arrangement can be determined from the determined track group sets based on the timbre characteristics of the input music.

[0063] Continue to refer Figure 5 At 530, music can be generated. In the music generation 530, it mainly involves planning the music structure of the output music 532 and arranging the output music 534 according to the secondary selected mixed audio track. Figure 3 As shown, after selecting the track group set at 307, the music structure planning 532 of the output music will be performed at 309. The music structure planning can ensure that the output music finally generated is consistent with the input music in structure and rhythm.

[0064] In some embodiments, the structure of the output music can be determined based on the music elements obtained according to the MIR model and the selected track group set, that is, how many bars each part of the output music and each part corresponding to the music segment should be divided into. In some embodiments, if the best time stretch ratio has been calculated in the track group set selection process 303, the number of bars can be determined based on the best time stretch ratio. If the best time stretch ratio has not been calculated, the speed ratio between the selected track group set and the input music can be directly used to calculate how many bars each part of the output music needs to generate. This can facilitate adjusting the speed of the output music to match the rhythm of the input music, thereby allowing the output music to have the same music duration as the input music. In some embodiments, in this process, the starting time point of the output music and the duration of the output music audio content can also be calculated, thereby paving the way for subsequent generation of the output music. By preparing information such as the start timestamp, the audio content duration, and the audio duration, necessary inputs can be provided for subsequent processing, thereby simplifying the processing flow and improving overall efficiency.

[0065] Continue to refer Figure 5 At 534, a random mix of audio tracks is selected from the selected track group set for use in the music arrangement. Figure 3 As described above, after the track set selection process 303, a selected track set can be obtained at 307. Next, a random mix of audio tracks is selected from the selected track set for use in the music arrangement, i.e., the mixed audio track that best suits the music arrangement 534 is again selected from the selected subset. In some embodiments, the mixed audio track that best suits the music arrangement can be selected based on the musical elements of the input music. For example, the timbre of an instrument can be used to find the track set that best matches each short segment of the input music for the output music in the instrument timbre feature space. For example, a KNN search can be performed in the timbre feature space to find the mixed audio track that best matches the input music. In some embodiments, during the KNN search, the length of the mixed audio track and the number of bars required for each short segment of the music can be compared to ensure temporal matching. In some embodiments, the harmony or melody of the mixed audio track pool is compared with that of the input music, and the mixed audio track pool with the closest harmony or melody is selected. This ensures that the output music is harmonically or melodically consistent with the input music.

[0066] Decoupling the music generation and music understanding modules can improve the flexibility and scalability of the system. This design can enable the system to better generate music of various musical styles and genres while ensuring the quality and efficiency of the generated music.

[0067] Figure 6FIG. 6 is a block diagram of an apparatus 600 for generating music according to some embodiments of the present disclosure. Figure 6 As shown, apparatus 600 includes a music element determination module 602, configured to determine the music elements of input music, wherein the music elements of the input music include at least one of structure, tempo, genre, instrumentation timbre, harmony, and emotion. Apparatus 600 also includes a track group set determination module 604, configured to determine a track group set for output music based on the music elements of the input music, wherein a track group is a combination of music tracks. Apparatus 600 also includes an output music generation module 606, configured to generate output music corresponding to the input music based on the track group set for output music.

[0068] Figure 7 FIG2 shows a block diagram of an electronic device 700 according to some embodiments of the present disclosure. The device 700 may be a device or apparatus described in the embodiments of the present disclosure. Figure 7 As shown, the device 700 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or computer program instructions loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The CPU / GPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 707. An input / output (I / O) interface 705 is also connected to the bus 704. Although not shown in FIG. Figure 7 As shown in FIG, device 700 may further include a co-processor.

[0069] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0070] The various methods or processes described above may be performed by the CPU / GPU 701. For example, in some embodiments, the methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU / GPU 701, one or more steps or actions in the methods or processes described above may be performed.

[0071] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.

[0072] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0073] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0074] The computer program instructions for performing the disclosed operation can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or the object code written in any combination of one or more programming languages, programming languages ​​include object-oriented programming languages, and conventional procedural programming languages.Computer-readable program instructions can be performed completely on a user's computer, partially on a user's computer, performed as an independent software package, partly on a user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (such as utilizing an internet service provider to connect by the internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to carry out personalized customization electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLA), this electronic circuit can perform computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0075] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0076] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0077] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a special hardware-based system that performs the prescribed function or action, or can be implemented by a combination of special hardware and computer instructions.

[0078] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements to existing technologies, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

[0079] Some example implementations of the present disclosure are listed below.

[0080] Example 1. A method for generating music, comprising:

[0081] Determining musical elements of input music, wherein the musical elements of the input music include at least one of structure, tempo, genre, instrumentation, harmony, and emotion;

[0082] Determining a set of track groups of output music based on the music elements of the input music, wherein the track group is a combination of tracks of music; and

[0083] Based on the track group set of output music, the output music corresponding to the input music is generated.

[0084] Example 2. The method of example 1, wherein determining the musical elements of the input music comprises:

[0085] The music elements of the input music are determined by using multiple music information retrieval models or a unified multi-task basic music information retrieval model.

[0086] Example 3. The method of any one of Examples 1-2, wherein determining a set of track groups for output music based on musical elements of the input music comprises:

[0087] Invoke a pool of randomly mixed audio tracks; and

[0088] Based on the randomly mixed audio track pool, a set of audio track groups of the output music is determined.

[0089] Example 4. The method of any of Examples 1-3, wherein invoking the pool of randomly mixed audio tracks comprises:

[0090] Based on the structure of the track groups in the track group set, the track groups in the track group set are divided into a plurality of track group sections, the plurality of track group sets are combined into a track group set pool, each track group set in the plurality of track group sets includes one or more track groups, the track groups in the track groups have the same style and are harmoniously and uniformly combined; and

[0091] Based on the audio event detection, a plurality of valid track group bars are determined from the plurality of track group bars.

[0092] Example 5. The method of any one of Examples 1-4, further comprising:

[0093] Based on beat tracking results of the plurality of valid track group bars, generating a plurality of mixed audio tracks in the track group set by randomly mixing the plurality of valid track bars in the track group set, wherein the plurality of mixed audio tracks are different from each other; and

[0094] Based on the plurality of mixed audio tracks, the randomly mixed audio track pool is generated.

[0095] Example 6. The method of any one of Examples 1-5, further comprising recalling a plurality of combinations from the pool of randomly mixed audio tracks based on musical elements of the plurality of music segments and musical elements of the plurality of mixed audio tracks, the combinations comprising the plurality of mixed audio tracks, the recalled mixed audio tracks being from the plurality of track group sets, the recalled mixed audio tracks being associated with the music segments, and the music segments being divided according to the structure of the input music.

[0096] voting for the plurality of mixed audio tracks recalled from the plurality of combinations based on musical elements of the plurality of music segments, respectively;

[0097] Determining a ranking of the set of track groups based on the votes; and

[0098] Based on the sorting, a set of track groups for the output music is determined.

[0099] Example 7. The method of any of Examples 1-6, wherein determining a set of track groups for the output music based on the sorting comprises:

[0100] Determining a time stretching ratio between the plurality of music segments and the plurality of recalled mixed audio tracks in the track group set based on the beat and start node information of the music segments; and

[0101] Based on the time stretching ratio, a set of track groups for the output music is determined.

[0102] Example 8. The method of any one of Examples 1-7, further comprising:

[0103] Based on the musical elements of the input music and the determined track group set, the number of bars of each segment of the output music, the start time of the output music, the audio content duration of the output music, and the audio duration of the output music are determined.

[0104] Example 9. The method of any one of Examples 1-8, wherein determining the number of measures of each segment of the output music based on the musical elements of the input music and the determined set of track groups comprises:

[0105] determining the number of measures of each segment of the output music based on the time stretching ratio; or

[0106] The number of measures of each segment of the output music is determined based on a beat ratio between the input music and the determined set of track groups.

[0107] Example 10. The method of any one of Examples 1-9, wherein based on the determined set of track groups, a mixed audio track corresponding to the determined set of track groups is determined from the pool of randomly mixed audio tracks.

[0108] generating a subset of the pool of randomly mixed audio tracks based on the corresponding mixed audio tracks; and

[0109] Based on the subset, a combination of mixed audio tracks of the output music is determined.

[0110] Example 11. The method of any of Examples 1-10, wherein determining a combination of mixed audio tracks of the output music based on the subset comprises:

[0111] Recalling a combination of multiple mixed audio tracks from the subset based on the musical elements of the multiple music segments, respectively, wherein the multiple mixed audio tracks recalled in the combination are associated with the multiple music segments;

[0112] voting for the plurality of mixed audio tracks recalled from the combination based on musical elements of the plurality of music segments, respectively; and

[0113] A combination of mixed audio tracks of the output music is determined based on the ranking.

[0114] Example 12. The method of any of Examples 1-11, wherein determining a combination of mixed audio tracks of the output music based on the ranking comprises:

[0115] Based on the length of the recalled mixed audio track and the number of bars of the output music, a combination of the mixed audio tracks of the output music is determined.

[0116] Example 13. The method of any of Examples 1-12, wherein generating the output music based on a set of track groups of the output music comprises:

[0117] The output music is generated based on the number of bars of each segment of the output music and a combination of the determined mixed audio tracks of the output music.

[0118] Example 14. An apparatus for generating music, comprising:

[0119] a music element determination module configured to determine music elements of input music, wherein the music elements of the input music include at least one of structure, beat, genre, instrumentation timbre, harmony, and emotion;

[0120] a track group set determining module configured to determine a track group set of output music based on the music elements of the input music, wherein the track group is a combination of music tracks; and

[0121] The output music generation module is configured to generate the output music corresponding to the input music based on the track group set of the output music.

[0122] Example 15. The apparatus of any one of Example 14, wherein the music element determination module comprises:

[0123] The music elements of the input music are determined by using multiple music information retrieval models or a unified multi-task basic music information retrieval model.

[0124] Example 16. The apparatus of any of Examples 14-15, wherein the module for determining the set of audio track groups comprises:

[0125] Invoke a pool of randomly mixed audio tracks; and

[0126] Based on the randomly mixed audio track pool, a set of audio track groups of the output music is determined.

[0127] Example 17. The apparatus of any of Examples 14-16, wherein invoking the pool of randomly mixed audio tracks comprises:

[0128] Based on the structure of the track groups in the track group set, the track groups in the track group set are divided into a plurality of track group sections, the plurality of track group sets are combined into a track group set pool, each track group set in the plurality of track group sets includes one or more track groups, the track groups in the track groups have the same style and are harmoniously and uniformly combined; and

[0129] Based on the audio event detection, a plurality of valid track group bars are determined from the plurality of track group bars.

[0130] Example 18. The apparatus of any of Examples 14-17, further comprising:

[0131] Based on beat tracking results of the plurality of valid track group bars, generating a plurality of mixed audio tracks in the track group set by randomly mixing the plurality of valid track bars in the track group set, wherein the plurality of mixed audio tracks are different from each other; and

[0132] Based on the plurality of mixed audio tracks, the randomly mixed audio track pool is generated.

[0133] Example 19. The apparatus of any of Examples 14-18, further comprising:

[0134] Recalling a plurality of combinations from the randomly mixed audio track pool based on musical elements of the plurality of music segments and musical elements of the plurality of mixed audio tracks, respectively, the combinations comprising the plurality of mixed audio tracks, the recalled mixed audio tracks being from a plurality of the track group sets, the recalled mixed audio tracks being associated with the music segments, and the music segments being divided according to the structure of the input music;

[0135] voting for the plurality of mixed audio tracks recalled from the plurality of combinations based on musical elements of the plurality of music segments, respectively;

[0136] Determining a ranking of the set of track groups based on the votes; and

[0137] Based on the sorting, a set of track groups for the output music is determined.

[0138] Example 20. The apparatus of any of Examples 14-19, wherein determining, based on the ranking, a set of track groups for the output music comprises:

[0139] Determining a time stretching ratio between the plurality of music segments and the plurality of recalled mixed audio tracks in the track group set based on the beat and start node information of the music segments; and

[0140] Based on the time stretching ratio, a set of track groups for the output music is determined.

[0141] Example 21. The apparatus of any of Examples 14-20, further comprising:

[0142] Based on the musical elements of the input music and the determined track group set, the number of bars of each segment of the output music, the start time of the output music, the audio content duration of the output music, and the audio duration of the output music are determined.

[0143] Example 22. The apparatus of any of Examples 14-21, wherein determining the number of measures of each segment of the output music based on musical elements of the input music and the determined set of track groups comprises:

[0144] determining the number of measures of each segment of the output music based on the time stretching ratio; or

[0145] The number of measures of each segment of the output music is determined based on a beat ratio between the input music and the determined set of track groups.

[0146] Example 23. The apparatus of any of Examples 14-22, further comprising:

[0147] Based on the determined track group set, determining a mixed audio track corresponding to the determined track group set from the randomly mixed audio track pool;

[0148] generating a subset of the pool of randomly mixed audio tracks based on the corresponding mixed audio tracks; and

[0149] Based on the subset, a combination of mixed audio tracks of the output music is determined.

[0150] Example 24. The apparatus of any of Examples 14-23, wherein determining a combination of mixed audio tracks of the output music based on the subset comprises:

[0151] Recalling a combination of multiple mixed audio tracks from the subset based on the musical elements of the multiple music segments, respectively, wherein the multiple mixed audio tracks recalled in the combination are associated with the multiple music segments;

[0152] voting for the plurality of mixed audio tracks recalled from the combination based on musical elements of the plurality of music segments, respectively; and

[0153] A combination of mixed audio tracks of the output music is determined based on the ranking.

[0154] Example 25. The apparatus of any of Examples 14-24, wherein determining a combination of mixed audio tracks of the output music based on the ranking comprises:

[0155] Based on the length of the recalled mixed audio track and the number of bars of the output music, a combination of the mixed audio tracks of the output music is determined.

[0156] Example 26. The apparatus of any of Examples 14-25, wherein generating the output music based on a set of track groups of the output music comprises:

[0157] The output music is generated based on the number of bars of each segment of the output music and a combination of the determined mixed audio tracks of the output music.

[0158] Example 27. An electronic device comprising:

[0159] processor; and

[0160] A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs actions, the actions comprising:

[0161] Determining musical elements of input music, wherein the musical elements of the input music include at least one of structure, tempo, genre, instrumentation, harmony, and emotion;

[0162] Determining a set of track groups of output music based on the music elements of the input music, wherein the track group is a combination of tracks of music; and

[0163] Based on the track group set of output music, the output music corresponding to the input music is generated.

[0164] Example 28. The electronic device of Example 27, wherein determining the music elements of the input music comprises:

[0165] The music elements of the input music are determined by using multiple music information retrieval models or a unified multi-task basic music information retrieval model.

[0166] Example 29. The electronic device of any of Examples 27-28, wherein determining a set of track groups for output music based on musical elements of the input music comprises:

[0167] Invoke a pool of randomly mixed audio tracks; and

[0168] Based on the randomly mixed audio track pool, a set of audio track groups of the output music is determined.

[0169] Example 30. The electronic device of any of Examples 27-29, wherein invoking the randomly mixed pool of audio tracks comprises:

[0170] Based on the structure of the track groups in the track group set, the track groups in the track group set are divided into a plurality of track group sections, the plurality of track group sets are combined into a track group set pool, each track group set in the plurality of track group sets includes one or more track groups, the track groups in the track groups have the same style and are harmoniously and uniformly combined; and

[0171] Based on the audio event detection, a plurality of valid track group bars are determined from the plurality of track group bars.

[0172] Example 31. The electronic device of any of Examples 27-30, further comprising:

[0173] Based on beat tracking results of the plurality of valid track group bars, generating a plurality of mixed audio tracks in the track group set by randomly mixing the plurality of valid track bars in the track group set, wherein the plurality of mixed audio tracks are different from each other; and

[0174] Based on the plurality of mixed audio tracks, the randomly mixed audio track pool is generated.

[0175] Example 32. The electronic device of any of Examples 27-31, further comprising:

[0176] Recalling a plurality of combinations from the randomly mixed audio track pool based on musical elements of the plurality of music segments and musical elements of the plurality of mixed audio tracks, respectively, the combinations comprising the plurality of mixed audio tracks, the recalled mixed audio tracks being from a plurality of the track group sets, the recalled mixed audio tracks being associated with the music segments, and the music segments being divided according to the structure of the input music;

[0177] voting for the plurality of mixed audio tracks recalled from the plurality of combinations based on musical elements of the plurality of music segments, respectively;

[0178] Determining a ranking of the set of track groups based on the votes; and

[0179] Based on the sorting, a set of track groups for the output music is determined.

[0180] Example 33. The electronic device of any of Examples 27-32, wherein determining, based on the sorting, a set of track groups for the output music comprises:

[0181] Determining a time stretching ratio between the plurality of music segments and the plurality of recalled mixed audio tracks in the track group set based on the beat and start node information of the music segments; and

[0182] Based on the time stretching ratio, a set of track groups for the output music is determined.

[0183] Example 34. The electronic device of any of Examples 27-33, further comprising:

[0184] Based on the musical elements of the input music and the determined track group set, the number of bars of each segment of the output music, the start time of the output music, the audio content duration of the output music, and the audio duration of the output music are determined.

[0185] Example 35. The electronic device of any of Examples 27-34, wherein determining the number of bars of each segment of the output music based on the musical elements of the input music and the determined set of track groups comprises:

[0186] determining the number of measures of each segment of the output music based on the time stretching ratio; or

[0187] The number of measures of each segment of the output music is determined based on a beat ratio between the input music and the determined set of track groups.

[0188] Example 36. The electronic device of any of Examples 27-35, further comprising:

[0189] Based on the determined track group set, determining a mixed audio track corresponding to the determined track group set from the randomly mixed audio track pool;

[0190] generating a subset of the pool of randomly mixed audio tracks based on the corresponding mixed audio tracks; and

[0191] Based on the subset, a combination of mixed audio tracks of the output music is determined.

[0192] Example 37. The electronic device of any of Examples 27-36, wherein determining a combination of mixed audio tracks of the output music based on the subset comprises:

[0193] Recalling a combination of multiple mixed audio tracks from the subset based on the musical elements of the multiple music segments, respectively, wherein the multiple mixed audio tracks recalled in the combination are associated with the multiple music segments;

[0194] voting for the plurality of mixed audio tracks recalled from the combination based on musical elements of the plurality of music segments, respectively; and

[0195] A combination of mixed audio tracks of the output music is determined based on the ranking.

[0196] Example 38. The electronic device of any of Examples 27-37, wherein determining a combination of mixed audio tracks of the output music based on the ranking comprises:

[0197] Based on the length of the recalled mixed audio track and the number of bars of the output music, a combination of the mixed audio tracks of the output music is determined.

[0198] Example 39. The electronic device of any of Examples 27-38, wherein generating the output music based on the set of track groups of the output music comprises:

[0199] The output music is generated based on the number of bars of each segment of the output music and a combination of the determined mixed audio tracks of the output music.

[0200] Example 40. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of Examples 1 to 13.

[0201] Example 41. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any one of Examples 1 to 13.

[0202] Although the present disclosure has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating music, comprising: Determining musical elements of input music, wherein the musical elements of the input music include at least one of structure, tempo, genre, instrumentation, harmony, and emotion; Determining a set of audio track groups of output music based on the musical elements of the input music, wherein the audio track group is a combination of audio tracks of the music; as well as Based on the track group set of output music, the output music corresponding to the input music is generated.

2. The method according to claim 1, wherein determining the musical elements of the input music comprises: The music elements of the input music are determined by using multiple music information retrieval models or a unified multi-task basic music information retrieval model.

3. The method according to claim 1 , wherein determining a set of audio track groups of output music based on the musical elements of the input music comprises: Calls a pool of randomly mixed audio tracks; as well as Based on the randomly mixed audio track pool, a set of audio track groups of the output music is determined.

4. The method according to claim 3, wherein calling the randomly mixed audio track pool comprises: Based on the structure of the track groups in the track group set, the track groups in the track group set are divided into a plurality of track group sections, the plurality of track group sets are combined into a track group set pool, each track group set in the plurality of track group sets includes one or more track groups, the track groups in the track groups have the same style and are harmoniously and uniformly combined; and Based on the audio event detection, a plurality of valid track group bars are determined from the plurality of track group bars.

5. The method according to claim 4, further comprising: Based on beat tracking results of the plurality of valid track group bars, generating a plurality of mixed audio tracks in the track group set by randomly mixing the plurality of valid track bars in the track group set, wherein the plurality of mixed audio tracks are different from each other; as well as Based on the plurality of mixed audio tracks, the randomly mixed audio track pool is generated.

6. The method according to claim 5, further comprising: Recalling a plurality of combinations from the randomly mixed audio track pool based on musical elements of the plurality of music segments and musical elements of the plurality of mixed audio tracks, respectively, the combinations comprising the plurality of mixed audio tracks, the recalled mixed audio tracks being from a plurality of the track group sets, the recalled mixed audio tracks being associated with the music segments, and the music segments being divided according to the structure of the input music; voting for the plurality of mixed audio tracks recalled from the plurality of combinations based on musical elements of the plurality of music segments, respectively; Based on the votes, determining the ranking of the set of audio track groups; as well as Based on the sorting, a set of track groups for the output music is determined.

7. The method of claim 6 , wherein determining a set of track groups for outputting the music based on the sorting comprises: Determining a time stretching ratio between the plurality of music segments and the plurality of recalled mixed audio tracks in the track group set based on the beat and start node information of the music segments; and Based on the time stretching ratio, a set of track groups for the output music is determined.

8. The method according to claim 7, further comprising: Based on the musical elements of the input music and the determined track group set, the number of bars of each segment of the output music, the start time of the output music, the audio content duration of the output music, and the audio duration of the output music are determined.

9. The method according to claim 8, wherein determining the number of bars of each section of the output music based on the musical elements of the input music and the determined set of track groups comprises: determining the number of measures of each segment of the output music based on the time stretching ratio; or The number of measures of each segment of the output music is determined based on a beat ratio between the input music and the determined set of track groups.

10. The method according to claim 6, further comprising: Based on the determined track group set, determining a mixed audio track corresponding to the determined track group set from the randomly mixed audio track pool; generating a subset of the pool of randomly mixed audio tracks based on the corresponding mixed audio tracks; as well as Based on the subset, a combination of mixed audio tracks of the output music is determined.

11. The method of claim 10 , wherein determining a combination of mixed audio tracks of the output music based on the subset comprises: Recalling a combination of multiple mixed audio tracks from the subset based on the musical elements of the multiple music segments, respectively, wherein the multiple mixed audio tracks recalled in the combination are associated with the multiple music segments; voting for the plurality of mixed audio tracks recalled from the combination based on musical elements of the plurality of music segments, respectively; as well as A combination of mixed audio tracks of the output music is determined based on the ranking.

12. The method of claim 11 , wherein determining a combination of mixed audio tracks of the output music based on the ranking comprises: Based on the length of the recalled mixed audio track and the number of bars of the output music, a combination of the mixed audio tracks of the output music is determined.

13. The method of claim 12 , wherein generating the output music based on the set of track groups of the output music comprises: The output music is generated based on the number of bars of each segment of the output music and a combination of the determined mixed audio tracks of the output music.

14. An apparatus for generating music, comprising: a music element determination module configured to determine music elements of input music, wherein the music elements of the input music include at least one of structure, beat, genre, instrumentation timbre, harmony, and emotion; a track group determination module configured to determine a track group set of output music based on the music elements of the input music, wherein the track group is a combination of music tracks; as well as The output music generation module is configured to generate the output music corresponding to the input music based on the track group set of the output music.

15. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs the method according to any one of claims 1 to 13.

16. A computer program product comprising computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method according to any one of claims 1 to 13.