Target music generation method and device, terminal and storage medium

CN114299899BActive Publication Date: 2026-08-07特赞(上海)信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
特赞(上海)信息科技有限公司
Filing Date
2021-12-03
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本申请的主要目的在于提供一种视频的分幕节点预测方法、装置、终端及存储介质,以解决相关技术中对分割点进行预测存在准确度低的问题

Benefits of technology

[0038]本发明实施例提供了一种目标音乐的生成方法、装置、终端及存储介质,包括:先基于目标音频文件和初始模型,确定目标片段生成模型,然后基于目标片段生成模型和目标音频特征数据,得到多个音频片段,再从多个音频片段中选取一个音乐片段作为目标音频片段,最后基于目标音频片段、目标音频片段对应的类型和目标排列方式,生成目标音乐。本发明基于音频片段的音乐重组,能够使生产的AI音乐更加流畅,符合人类对音乐的听感需求。并且,可以根据需要生产不同时长的有版权音乐,能够为媒体创作者的生产提高效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299899B_ABST
    Figure CN114299899B_ABST
Patent Text Reader

Abstract

The application discloses a target music generation method and device, a terminal and a storage medium. The method comprises the following steps: determining a target segment generation model based on a target audio file and an initial model; obtaining a plurality of audio segments based on the target segment generation model and target audio feature data; selecting a music segment from the plurality of audio segments as a target audio segment; and generating target music based on the target audio segment, a type corresponding to the target audio segment and a target arrangement mode. The music reorganization based on the audio segment can make the produced AI music more fluent, meet the listening needs of human beings for music, and produce copyrighted music with different lengths according to needs, thereby improving the production efficiency of media creators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a method, apparatus, terminal, and storage medium for generating target music. Background Technology

[0002] In the media information age, the volume of media creation is increasing daily, with soundtracks being an essential part of the process. This means there is a significant demand for copyrighted music, and media creators often further edit the soundtracks to fit the media's length. Therefore, AI can leverage its versatility and efficiency to improve the productivity of copyrighted music creation.

[0003] Currently, AI music generation technologies on the market mainly use MIDI musical notes as temporal signals, drawing on the language generation ideas in the field of NLG (Nature Language Generate), and aim to enable AI models to learn the temporal distribution patterns, thereby generating new note segments.

[0004] However, the aforementioned music generation methods based on note levels suffer from poor user experience. Summary of the Invention

[0005] The main objective of this application is to provide a method, apparatus, terminal, and storage medium for predicting video segmentation nodes, in order to solve the problem of low accuracy in predicting segmentation points in related technologies.

[0006] To achieve the above objectives, firstly, this application provides a method for generating target music, comprising:

[0007] Based on the target audio file and the initial model, determine the target segment generation model;

[0008] Based on the target segment generation model and target audio feature data, multiple audio segments are obtained;

[0009] Select one music segment from multiple audio segments as the target audio segment;

[0010] The target music is generated based on the target audio segment, the type of the target audio segment, and the target arrangement.

[0011] In one possible implementation, based on the target audio file and the initial model, a target segment generation model is determined, including:

[0012] The target audio file is converted to a different format to obtain the Mel spectrogram corresponding to the target audio file.

[0013] The initial model was trained using Mel spectrograms to obtain the target fragment generation model.

[0014] In one possible implementation, multiple audio segments are obtained based on the target segment generation model and target audio feature data, including:

[0015] Determine the target audio feature data;

[0016] The target audio feature data is input into the target segment generation model to obtain multiple audio segments.

[0017] In one possible implementation, target music is generated based on the target audio segment, the type of the target audio segment, and the target arrangement, including:

[0018] The arrangement of the targets is determined based on the type of the target audio segments;

[0019] The target audio segments are arranged using a target arrangement method to generate the target music.

[0020] In one possible implementation, the target audio segment is a bass track audio segment;

[0021] The target audio segments are arranged using a target arrangement method to generate target music, including:

[0022] The bass track audio clip is continuously looped for a first preset duration to obtain the target music.

[0023] In one possible implementation, the target audio segment is a drum track audio segment, a chord track audio segment, or a melody track audio segment.

[0024] The target audio segments are arranged using a target arrangement method to generate target music, including:

[0025] The second preset duration is determined according to a preset probability;

[0026] The third preset duration is obtained by subtracting the first preset duration from the second preset duration.

[0027] The target music is obtained by continuously looping the audio clips of the drum track, chord track, or melody track within a third preset duration.

[0028] In one possible implementation, before determining the target segment generation model based on the target audio file and the initial model, the following steps are also included:

[0029] Select the target audio file type from different types of audio files;

[0030] Select a preset number of audio files of the target type as the target audio files.

[0031] Secondly, embodiments of the present invention provide a target music generation apparatus, comprising:

[0032] The target model determination module is used to determine the target segment generation model based on the target audio file and the initial model;

[0033] The initial segment determination module is used to generate multiple audio segments based on the target segment generation model and target audio feature data;

[0034] The target segment determination module is used to select one music segment as the target audio segment from multiple audio segments;

[0035] The target music generation module is used to generate target music based on the target audio segments, the type of the target audio segments, and the target arrangement.

[0036] Thirdly, embodiments of the present invention provide a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods for generating target music.

[0037] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-described methods for generating target music.

[0038] This invention provides a method, apparatus, terminal, and storage medium for generating target music, comprising: first, determining a target segment generation model based on a target audio file and an initial model; then, obtaining multiple audio segments based on the target segment generation model and target audio feature data; selecting one music segment from the multiple audio segments as the target audio segment; and finally, generating target music based on the target audio segment, its corresponding type, and a target arrangement. This invention, based on music recombination of audio segments, enables the produced AI music to be smoother and meets human listening preferences. Furthermore, it can produce copyrighted music of varying lengths as needed, improving production efficiency for media creators. Attached Figure Description

[0039] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings:

[0040] Figure 1 This is a flowchart illustrating the implementation of a method for generating target music according to an embodiment of the present invention;

[0041] Figure 2 This is a flowchart illustrating the implementation of training an initial model according to an embodiment of the present invention;

[0042] Figure 3 This is a schematic diagram of the structure of a target music generation device provided in an embodiment of the present invention;

[0043] Figure 4 This is a schematic diagram of the terminal provided in an embodiment of the present invention. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in sequences other than those illustrated or described herein.

[0046] It should be understood that in the various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0047] It should be understood that in this invention, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.

[0048] It should be understood that in this invention, "multiple" refers to two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, "and / or B" can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "Contains A, B, and C", "Contains A, B, and C" means that all three A, B, and C are contained; "Contains A, B, or C" means that one of A, B, and C is contained; "Contains A, B, and / or C" means that any one, two, or three of A, B, and C are contained.

[0049] It should be understood that in this invention, "B corresponding to A", "B corresponding to A", "A and B correspond", or "B and A correspond" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. Matching A and B is defined as a similarity between A and B that is greater than or equal to a preset threshold.

[0050] Depending on the context, "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection."

[0051] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0052] To make the objectives, technical solutions, and advantages of the present invention clearer, specific embodiments will be described below in conjunction with the accompanying drawings.

[0053] In one embodiment, such as Figure 1 As shown, a method for generating target music is provided, including the following steps:

[0054] Step S101: Based on the target audio file and the initial model, determine the target segment generation model;

[0055] Step S102: Based on the target segment generation model and target audio feature data, obtain multiple audio segments;

[0056] Step S103: Select one music segment from multiple audio segments as the target audio segment;

[0057] Step S104: Generate target music based on the target audio segment, the type of the target audio segment, and the target arrangement.

[0058] Specifically, the initial model is a deep learning model based on VAE (Variational Autoencoder). This model learned from a large number of Jazz Hiphop music tracks to obtain the target segment generation model. The initial model learned the distribution characteristics of pitch, timbre, and duration of different music tracks such as Drumtrack, Chord track, and Melody track, and included more than 10 instruments such as drums, guitar, piano, bass, horn, and violin, enabling it to generate different tracks.

[0059] This invention provides a method for generating target music, comprising: first, determining a target segment generation model based on a target audio file and an initial model; then, obtaining multiple audio segments based on the target segment generation model and target audio feature data; selecting one music segment from the multiple audio segments as the target audio segment; and finally, generating target music based on the target audio segment, its corresponding type, and a target arrangement. This invention, based on the music recombination of audio segments, enables the generated AI music to be smoother and meets human listening preferences. Furthermore, it can generate copyrighted music of varying lengths as needed, improving production efficiency for media creators.

[0060] In one embodiment, step S101 includes a process of determining the target audio file, that is, firstly selecting the target type audio file from different types of audio files, and then selecting a preset number of target type audio files as the target audio file.

[0061] Specifically, the audio file type is the track type in Table 1 below, namely Drum Track, Bass Track, Chord Track, Melody Track, etc., but it is not limited to the types in Table 1 below, and can also be other types.

[0062] From over 100 Jazz and Hip Hop songs, WAV audio files with a fixed BPM (Beat Per Minute) for different instrument types were collected as a single track. Each WAV file contains music for only one instrument. The data size for each instrument type is as follows:

[0063] Table 1. Music Track Types

[0064]

[0065]

[0066] In one embodiment, step S101 includes:

[0067] Step S201: Convert the format of the target audio file to obtain the Mel spectrogram corresponding to the target audio file;

[0068] Step S202: Train the initial model using Mel spectrograms to obtain the target segment generation model.

[0069] Combination Figure 2 The training of the initial model will be explained in detail. Specifically, for each track, we will train a VAE model. We will convert the WAV files of different tracks into Mel spectrograms and input them into the Audio VAE model. The Encoder part of the model learns the distribution characteristics of existing music segments from two aspects: pitch and timbre. The Decoder decodes and generates new segments, which will conform to the pitch and timbre distribution of the original WAV.

[0070] In one embodiment, step S102 includes:

[0071] Step S301: Determine the target audio feature data;

[0072] Step S302: Input the target audio feature data into the target segment generation model to obtain multiple audio segments.

[0073] Specifically, target audio feature data refers to data that includes two features: pitch and timbre. When the target audio feature data is input into the target segment generation model, multiple audio segments will be output, and then one audio segment will be randomly selected as the target audio segment.

[0074] In one embodiment, step S104 includes:

[0075] Step S401: Determine the target arrangement method based on the type corresponding to the target audio segment;

[0076] Step S402: Arrange the target audio segments using the target arrangement method to generate the target music.

[0077] Specifically, when the target audio segment is a bass track audio segment, the bass track audio segment is continuously looped within a first preset duration to obtain the target music; when the target audio segment is a drum track audio segment, chord track audio segment, or melody track audio segment, a second preset duration is determined according to a preset probability; the difference between the first preset duration and the second preset duration is taken to obtain a third preset duration; the drum track audio segment, chord track audio segment, or melody track audio segment is continuously looped within the third preset duration to obtain the target music. Here, the first preset duration refers to the overall duration of the target music, and the second preset duration refers to a period of time preceding the first preset duration, without specific limitations.

[0078] Furthermore, the process of generating corresponding target music from different types of audio segments is illustrated with specific embodiments:

[0079] Human-created Jazz Hiphop music has the following characteristics: a piece of music consists of the following segments: Intro - Verse - Build up - Drop / Chrous - Bridge - Verse - Build up - Drop / Chorus - Outro.

[0080] A musical segment consists of several parts: rhythm instruments (drum kit), orchestration (bass, piano, guitar, trumpet, etc.). The drum kit determines the rhythmic pattern of the music, while the orchestration forms the chord progression. Different musical styles will be paired with different orchestrations.

[0081] Generally, rhythm music first establishes the drum kit rhythm, which includes the bass drum, snare drum, and hi-hit cymbals. Then, the bass and other chords and instruments are laid out, and finally, the different tracks are merged, the volume is balanced, and the music is mixed.

[0082] This patent incorporates the above features. Unlike market methods that generate music at the note level, this patent generates music by generating bar segments and then arranging them in sequence. Therefore, once we set the duration for generating Jazz Hiphop music (i.e., the first preset duration), the segment arrangement follows these characteristics:

[0083] The bass track audio serves as the foundation of the overall music and will loop continuously from beginning to end;

[0084] The Drum Track audio will be delayed by several eight beats with a certain probability before entering;

[0085] The chord track audio will be resampled with a certain probability by filling in gaps within an eight-beat interval, and will enter with a certain probability by delaying for eight beats.

[0086] The Melody Track audio will be resampled with a certain probability by filling in gaps within an eight-beat interval, and will enter with a certain probability by delaying for eight beats.

[0087] It should be noted that the method for determining the second external structure model is similar to that for determining the first external structure model, and will not be elaborated here.

[0088] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0089] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.

[0090] Figure 3 The diagram illustrates the structure of a target music generation device according to an embodiment of the present invention. For ease of explanation, only the parts relevant to the embodiment of the present invention are shown. The target music generation device includes a target model determination module 31, an initial segment determination module 32, a target segment determination module 33, and a target music generation module 34, as detailed below:

[0091] The target model determination module 31 is used to determine the target segment generation model based on the target audio file and the initial model;

[0092] The initial segment determination module 32 is used to generate multiple audio segments based on the target segment generation model and target audio feature data;

[0093] The target segment determination module 33 is used to select a music segment as the target audio segment from multiple audio segments;

[0094] The target music generation module 34 is used to generate target music based on the target audio segment, the type of the target audio segment, and the target arrangement.

[0095] In one possible implementation, the target model determination module 31 includes:

[0096] The format conversion submodule is used to convert the format of the target audio file to obtain the Mel spectrogram corresponding to the target audio file;

[0097] The model training submodule is used to train the initial model using Mel spectrograms to obtain the target segment generation model.

[0098] In one possible implementation, the initial fragment determination module 32 includes:

[0099] The feature data determination submodule is used to determine the target audio feature data;

[0100] The initial audio determination submodule is used to input the target audio feature data into the target segment generation model to obtain multiple audio segments.

[0101] In one possible implementation, the target music generation module 34 includes:

[0102] The arrangement determination submodule is used to determine the target arrangement based on the type of the target audio segment;

[0103] The target music generation submodule is used to arrange target audio segments using the target arrangement method to generate target music.

[0104] In one possible implementation, the target audio segment is a bass track audio segment;

[0105] The target music generation submodule includes:

[0106] The first target music generation unit is used to continuously loop the bass track audio clip for a first preset duration to obtain the target music.

[0107] In one possible implementation, the target audio segment is a drum track audio segment, a chord track audio segment, or a melody track audio segment.

[0108] The target music generation submodule includes:

[0109] The first duration determination unit is used to determine the second preset duration according to a preset probability;

[0110] The second duration determination unit is used to calculate the difference between the first preset duration and the second preset duration to obtain the third preset duration.

[0111] The second target music generation unit is used to continuously loop the drum track audio clip, chord track audio clip, or melody track audio clip within a third preset duration to obtain the target music.

[0112] In one possible implementation, prior to the target model determination module 31, the following is also included:

[0113] The file selection submodule is used to select audio files of a target type from different types of audio files.

[0114] The target model determination submodule is used to select a preset number of audio files of the target type as target audio files.

[0115] Figure 4 This is a schematic diagram of a terminal provided in an embodiment of the present invention. Figure 4 As shown, the terminal 4 in this embodiment includes a processor 40, a memory 41, and a computer program 42 stored in the memory 41 and executable on the processor 40. When the processor 40 executes the computer program 42, it implements the steps in the above-described embodiments of the target music generation method, for example... Figure 1 Steps 101 to 104 are shown. Alternatively, when processor 40 executes computer program 42, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 3The functions of modules / units 31 to 34 shown.

[0116] The present invention also provides a readable storage medium storing a computer program, which, when executed by a processor, is used to implement the methods provided in the various embodiments described above.

[0117] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0118] The present invention also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the device to implement the methods provided in the various embodiments described above.

[0119] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.

[0120] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for generating target music, characterized in that, include: Based on the target audio file and the initial model, the target segment generation model is determined, including: The target audio file is converted to a different format to obtain the Mel spectrogram corresponding to the target audio file. The initial model is trained using the Mel spectrogram to obtain the target segment generation model, which has the ability to generate music for different tracks; Based on the target segment generation model and the target audio feature data, multiple audio segments are obtained, including: determining the target audio feature data, wherein the target audio feature data includes pitch features and timbre features; inputting the target audio feature data into the target segment generation model to obtain the multiple audio segments; Select one music segment from the plurality of audio segments as the target audio segment; Based on the target audio segment, the type corresponding to the target audio segment, and the target arrangement, generate the target music; The target audio segment is a drum track audio segment, a chord track audio segment, or a melody track audio segment; Arranging the target audio segments using the target arrangement method to generate the target music includes: When the target audio segment is a bass track audio segment, the bass track audio segment is continuously looped within a first preset duration to obtain the target music; a second preset duration is determined according to a preset probability, wherein the first preset duration refers to the overall duration of the target music, and the second preset duration refers to a period of time before the first preset duration; The third preset duration is obtained by subtracting the first preset duration from the second preset duration. The target music is obtained by continuously looping the drum track audio segment, the chord track audio segment, or the melody track audio segment within the third preset duration.

2. The method for generating target music as described in claim 1, characterized in that, The step of generating target music based on the target audio segment, the type corresponding to the target audio segment, and the target arrangement includes: The target arrangement is determined based on the type corresponding to the target audio segment; The target audio segments are arranged using the target arrangement method to generate the target music.

3. The method for generating target music as described in any one of claims 1-2, characterized in that, Before determining the target segment generation model based on the target audio file and the initial model, the process also includes: Select the target audio file type from different types of audio files; Select a preset number of audio files of the target type as the target audio files.

4. A device for generating target music, characterized in that, include: The target model determination module is used to determine the target segment generation model based on the target audio file and the initial model; The target model determination module is also used for: The target audio file is converted to a different format to obtain the Mel spectrogram corresponding to the target audio file. The initial model is trained using the Mel spectrogram to obtain the target segment generation model, which has the ability to generate music for different tracks; An initial segment determination module is used to obtain multiple audio segments based on the target segment generation model and target audio feature data, including: determining the target audio feature data, wherein the target audio feature data includes pitch features and timbre features; and inputting the target audio feature data into the target segment generation model to obtain the multiple audio segments. The target segment determination module is used to select one music segment as the target audio segment from the plurality of audio segments; The target music generation module is used to generate target music based on the target audio segment, the type corresponding to the target audio segment, and the target arrangement. The target music generation module is also used for: The target audio segment is a drum track audio segment, a chord track audio segment, or a melody track audio segment; The target music generation module is further configured to: when the target audio segment is a bass track audio segment, continuously loop the bass track audio segment within a first preset duration to obtain the target music; determine a second preset duration according to a preset probability, wherein the first preset duration refers to the overall duration of the target music, and the second preset duration refers to a period of time prior to the first preset duration; The third preset duration is obtained by subtracting the first preset duration from the second preset duration. The target music is obtained by continuously looping the drum track audio segment, the chord track audio segment, or the melody track audio segment within the third preset duration.

5. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for generating target music as described in any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for generating target music as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Music generation method and device based on generative adversarial network

    CN109346043A

  • Method and system for template based variant generation of hybrid ai generated song

    US20210118416A1