Controllable Generation Method, System, Electronic Device and Medium for Symbolic Music

By preprocessing the music content and chord conditional masking, a discrete diffusion model is formed, which solves the problem of insufficient harmony and rationality of multi-track and multi-instrument music generation content in the prior art, and achieves high-quality music generation.

CN119647529BActive Publication Date: 2025-06-24COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510167925.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-24
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The existing symbolic music generation technology lacks harmony and rationality in multi-track multi-instrument music generated content, resulting in poor quality of generated content and cannot meet users' needs for high-quality music content.

Method used

The music content and labeling information are preprocessed through the preset representation module, the data set is masked according to the chord condition rules, and the data is sent to the music discrete diffusion framework for repeated training to form a music discrete diffusion model. The music bands and requirements entered by the user are processed in the preamble and input into the model to generate the target music score.

Benefits of technology

It significantly improves the generation quality of multi-track symbol music, improves the harmony and rationality of music, and can generate high-quality music content to meet users' needs for high-quality music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119647529B_ABST
    Figure CN119647529B_ABST
Patent Text Reader

Abstract

The present invention provides a controllable generation method, system, electronic device and computer-readable storage medium for symbolic music. First, a preset representation module preprocesses the pre-acquired music content and annotation information to form an original data set, and performs masking processing on the original data set according to preset chord condition rules to form an original masked data pair. Then, the original masked data pair is fed into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model. Subsequently, the music frequency band and music requirements input by the user are preprocessed to form a music content representation and a condition matrix, and the music content representation and the condition matrix are input into the music discrete diffusion model, so that the music discrete diffusion model generates music to output a target music score. In this way, not only can the richness of controllable music generation be achieved, but also the reliability and rationality can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly, to a method for generating music, and more specifically, to a method, system, electronic device and medium for controllably generating symbolic music. Background Art

[0002] Traditional music creation requires producers to have professional music theory knowledge and rich production experience, with high professional thresholds and artistic accomplishment requirements. With the rapid development of artificial intelligence and multimedia technology, but with the gradual development of AI-generated content, symbolic music (specifically computer music in MIDI format) can also achieve intelligent composition, intelligent arrangement and other music creation assistance for humans, becoming a music production tool. However, the existing symbolic music generation technology is slightly lacking in the harmony and rationality of the generated content of multi-track multi-instrument music, resulting in poor quality of the generated content and unable to meet the needs of users for high-quality music content.

[0003] For example, Microsoft Pop-MAG proposed a multi-track music representation (Multi-Track MIDI Representation, MTRepresentation) to re-encode the token sequence into a two-dimensional sequence according to event attributes, providing an idea for data organization. However, in the face of multi-track and multi-instrument music, the model poorly learns the note distribution characteristics of various instruments, and the rationality of the generated music for instruments is poor.

[0004] Microsoft Pop-MAG is a classic accompaniment generation work based on the Transformer method. These classic generation models require the decoder network to pay attention to the correct conditional tokens or hidden layer feature vectors when predicting the next sentence. Obtaining high-quality samples depends on a large model scale, and as the sentence gets longer, the conditional constraints become weaker, which makes the autoregressive model mainly based on the decoder encounter bottlenecks in terms of generation quality and controllable generation problems.

[0005] SCHmUBERT is a discrete diffusion model for music generation, but it requires a trained classifier and an additional convolutional neural network, and it is unconditional accompaniment generation. Its public work can only process bass and drum accompaniments. There are still problems to be solved in terms of exerting the controllable generation ability of the probabilistic diffusion model.

[0006] Therefore, there is an urgent need for a controllable generation scheme for symbolic music that can improve the generation quality of multi-track symbolic music and enhance the harmony and rationality of multi-track music. Summary of the Invention

[0007] In view of the above problems, the purpose of the present invention is to provide a controllable generation method, system, electronic device and medium for symbolic music, so as to solve the technical problems that the existing symbolic music generation technology lacks harmony and rationality in the generated content of multi-track and multi-instrument music, resulting in poor quality of the generated content and inability to meet the user's demand for high-quality music content.

[0008] In a first aspect, an embodiment of the present application provides a controllable generation method for symbolic music. The controllable generation method for symbolic music includes:

[0009] Preprocessing the pre-acquired music content and annotation information through a preset representation module to form an original data set;

[0010] Performing masking processing on the original data set according to a preset chord condition rule to form an original masked data pair;

[0011] Feeding the original masked data pair into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model;

[0012] Performing preprocessing on the music frequency band and music requirements input by the user to form a music content representation and a condition matrix, and inputting the music content representation and the condition matrix into the music discrete diffusion model, so that the music discrete diffusion model performs music generation to output a target music score.

[0013] Optionally, the performing preprocessing on the music frequency band and music requirements input by the user to form a music content representation and a condition matrix includes:

[0014] Segmenting the music frequency band to form music segments;

[0015] Encoding each music segment to form a music content representation; obtaining the element information of the music segment, combining the element information with the music requirements to obtain music combination information, and encoding the music combination information to form a condition matrix.

[0016] Optionally, the music discrete diffusion model performs music generation to output a target music score, including:

[0017] Concatenating the music content representation and the condition matrix of each music segment to form a combined representation;

[0018] Performing continuous-domain word embedding on the combined representation using a preset three-dimensional word embedding space to form a combined word embedding vector;

[0019] Performing Markov transfer on the combined word embedding vector using a preset layer normalization module to obtain the noise initial state of the combined word embedding vector;

[0020] Dynamically add noise to the masked part in the initial state of the noise through a preset diffusion module to generate a hidden state vector;

[0021] Apply positional word embedding and time step word embedding to the hidden state vector to generate a diffusion hidden vector, and then perform an inverse operation on the diffusion hidden vector through a preset bidirectional encoder network to obtain a target hidden vector;

[0022] Perform class mapping on the target hidden vector through a preset multi-head output module to obtain a target music content representation matrix;

[0023] Decode the target music content representation matrix to generate a target music score frequency band;

[0024] Stitch all the target music score frequency bands in chronological order to form a target music score.

[0025] Optionally, load a music sampling file for the target music score through a preset audio library to generate target music.

[0026] Optionally, the preset representation module preprocesses the pre-acquired music content and annotation information to form an original data set, including:

[0027] Use the chords of the music content as the time line, automatically quantify the start time of the notes in units of 32nd notes, quantify the note duration in units of 16th notes, perform self-correction, and automatically delete empty tracks and empty notes to form a standardized music content, and use the annotation information to perform self-annotation on the standardized music content to form a music content set; wherein, each music event of the standardized music content is included in the music content set, and music attributes are annotated on the music event;

[0028] Perform matrix combination on all music events of a piece of music content to form a music content representation, and combine the music content representations of all music contents to form a music content representation data set;

[0029] Use the music content representation data set as the original data set.

[0030] Optionally, the masking process of the original data set according to the preset chord condition rules to form an original masked data pair includes:

[0031] Display the chord music events in the music content representation of the original data set, and mask the non-chord music events in the music content representation; wherein, the masked music events are marked with PAD.

[0032] Optionally, in the process of feeding the original masked data pair into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model,

[0033] the original masked data is used to train the music discrete framework through normalized embedding learning; wherein, the music discrete framework includes a diffusion model framework, a decoding model framework, and a word embedding module;

[0034] the norm loss of the diffusion model framework is calculated by a preset BERT model, and the attention mechanism of the diffusion model framework is data-related to the word embedding vectors involved in the word embedding module;

[0035] The objective function during the training of the music discrete diffusion framework is the sum of the loss functions of the diffusion model framework, the decoding model framework, and the word embedding module.

[0036] In a second aspect, an embodiment of the present application provides a controllable generation system for symbolic music, which implements the controllable generation method for symbolic music as described above. The system includes a music discrete diffusion model; wherein,

[0037] the music discrete diffusion model is learned and generated by repeatedly training a preset music discrete diffusion framework with the original masked data pair; wherein, obtaining the original masked data pair includes: preprocessing the pre-acquired music content and annotation information through a preset characterization module to form an original data set; masking the original data set according to a preset chord condition rule to form an original masked data pair;

[0038] The music discrete diffusion model is used to receive the music content representation and the condition matrix formed by preprocessing the user input music frequency band and music requirements, and perform music generation according to the music content representation and the condition matrix to output a target musical score.

[0039] In a third aspect, an embodiment of the present application provides an electronic device, which includes: a processor, a memory, and a system bus; the processor and the memory are connected through the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, and when the instructions are executed by the processor, the processor executes the steps in any optional controllable generation method for symbolic music in the first aspect.

[0040] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any optional controllable generation method for symbolic music in the first aspect.

[0041] As can be seen from the above technical solutions, the embodiments of the present application provide a method, a system, an electronic device, and a computer-readable storage medium for controllable generation of symbolic music. Compared with the prior art, the present application has the following beneficial effects:

[0042] Taking the chords of the music content as the timeline can significantly reduce the sequence length and retain the time and multi-track spatial information;

[0043] Masking the original data set according to the preset chord condition rules to form an original masked data pair, strengthening the restrictions of chords in composition, and improving the effectiveness of feature learning and generation of pop music;

[0044] When training and using the music discrete diffusion model, first splice the matrices, then perform discrete denoising, and finally perform decoding. After decoding, probability mapping of the hidden vector is also required, which can not only achieve the richness of controllable music generation, but also improve the reliability and rationality. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] By referring to the following description of the specification in conjunction with the drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more apparent and easier to understand. In the drawings:

[0046] Figure 1 is a flowchart of a method for controllable generation of symbolic music according to an embodiment of the present invention;

[0047] Figure 2 is a schematic diagram of the usage scenario of a method for controllable generation of symbolic music according to an embodiment of the present invention;

[0048] Figure 3 is a schematic diagram of a music discrete diffusion model of a method for controllable generation of symbolic music according to an embodiment of the present invention;

[0049] Figure 4 is a schematic diagram of a three-dimensional word embedding space of a method for controllable generation of symbolic music according to an embodiment of the present invention;

[0050] Figure 5 is a schematic diagram of a multi-head output module of a method for controllable generation of symbolic music according to an embodiment of the present invention;

[0051] Figure 6 is a logic block diagram of a system for controllable generation of symbolic music according to an embodiment of the present invention;

[0052] Figure 7 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The existing symbolic music generation technology lacks harmony and rationality in the generated content of multi-track and multi-instrument music, resulting in poor quality of the generated content and failing to meet the user's demand for high-quality music content.

[0054] To address the above problems, the present invention provides a controllable generation method, system, electronic device, and computer-readable storage medium for symbolic music. First, a preset representation module preprocesses the pre-acquired music content and annotation information to form an original data set. The original data set is then masked according to preset chord condition rules to form an original masked data pair. The original masked data pair is then fed into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model. Subsequently, the music frequency band and music requirements input by the user are preprocessed to form a music content representation and a condition matrix. The music content representation and the condition matrix are input into the music discrete diffusion model, enabling the music discrete diffusion model to generate music and output a target music score. In this way, using the chords of the music content as the time line can significantly reduce the sequence length and retain the time and multi-track spatial information. Masking the original data set according to the preset chord condition rules to form an original masked data pair strengthens the constraints of chords in composition and improves the effectiveness of feature learning and generation of pop music. When training and using the music discrete diffusion model, first splice the matrices, then perform discrete noise addition, and finally perform decoding. After decoding, probability mapping of the hidden vector is also required, which can not only achieve the richness of controllable music generation but also improve the reliability and rationality.

[0055] To enable those skilled in the art of this technology to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. The description of the following exemplary embodiments is actually only illustrative and in no way constitutes a limitation on the present invention and its application or use. Technologies and devices known to those of ordinary skill in the relevant fields may not be discussed in detail, but where appropriate, the said technologies and devices should be regarded as part of the specification. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.

[0056] Figure 1 The flowchart of a controllable generation method for symbolic music provided by an embodiment of this application is as Figure 1 shown, and the method includes:

[0057] S1: Preprocess the pre-acquired music content and annotation information through a preset representation module to form an original data set;

[0058] S2: Mask the original data set according to the preset chord condition rules to form an original masked data pair;

[0059] S3: Feed the original masked data pair into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model;

[0060] S4: Perform preprocessing on the music frequency band and music requirements input by the user to form a music content representation and a condition matrix, and input the music content representation and the condition matrix into the music discrete diffusion model, so that the music discrete diffusion model generates music to output a target music score.

[0061] Step S1 is a process of preprocessing the pre-acquired music content and annotation information through a preset representation module to form an original data set. This process is a basic step for training the subsequent music discrete diffusion model. Among them, the process of preprocessing the pre-acquired music content and annotation information through a preset representation module to form an original data set includes:

[0062] S11: Use the chords of the music content as the time line, automatically quantify the start time of the notes in units of 32nd notes, quantify the note duration in units of 16th notes, perform self-correction, and automatically delete empty tracks and empty notes to form a standardized music content, and use the annotation information to perform self-annotation on the standardized music content to form a music content set; among them, each music event of the standardized music content is included in the music content set, and music attributes are annotated on the music event;

[0063] S12: Perform matrix combination on all music events of a piece of music content to form a music content representation, and combine the music content representations of all music contents to form a music content representation data set;

[0064] S13: Use the music content representation data set as the original data set.

[0065] In a specific embodiment of the present invention, the miditoolkit tool of python can be used to read music content information from a Musical Instrument Digital Interface (MIDI) file (a music file with the extension.mid), and jsonl can be used to read the corresponding text annotation information of the music from the annotation file. The annotation information includes the style and emotional atmosphere (Mood, abbreviated as emotion) of the music. We extracted a total of about 6000 pairs of MIDI music and annotations.

[0066] Due to problems often occurring in the music content recorded in MIDI files, such as misalignment of note beats, incorrect editing of note durations, empty tracks, etc., and MIDI files usually having no instrument type markings. Therefore, in this specific embodiment, preprocessing is first performed on the MIDI file, for example: (1) Automatically quantize the start time of notes in units of 32nd notes; (2) Quantize note durations in units of 16th notes; (3) Automatically correct errors, delete empty tracks and empty notes; (4) Automatically annotate the instrument type according to the MIDI protocol, etc.

[0067] In symbolic music, notes are the most basic elements, which are jointly described by note attributes such as instrument type (ins), start time (time), pitch (pitch), and duration (dur). Therefore, in this embodiment, the music event e ij is used to represent the j-th attribute (Attribute) of the i-th note (Note). Chords are regarded as special notes, and chord types (such as major chords, augmented fourth chords) are classified as pitch events. And in this specific embodiment, music information other than chords and notes is collectively referred to as meta events, such as music tempo (Tempo), bar number markings, and track change markings, etc.

[0068] In this specific embodiment, a chord condition rule is adopted, which includes two parts: relative time encoding; a masking method based on chord conditions.

[0069] The relative time encoding of the chord condition rule is: using chords to calculate the relative time information of music content, that is, the chord progression is the timeline of the entire music content, rather than using meta information markings (bars, tracks) like SymphonyNet and other REMI works. In this embodiment, using the chords of the music content as the timeline method can significantly reduce the sequence length and retain time and multi-track space information. Explanation: (1) Bar number markings and track change markings are used to record musical scores and have little impact on auditory appreciation. (2) For modern pop music, all tracks follow a chord progression in auditory perception, and the entire music also follows the development of this chord progression. Strengthening the restrictions of chords in composition is very effective for the feature learning and generation of pop music.

[0070] Therefore, in this embodiment, the attribute event e ij is used to tokenize the music content, and other important meta events, such as music tempo, will be introduced as conditions into the training and generation processes like emotion and style labels.

[0071] The token sequence must conform to the correct encoding grammar structure. For example, according to the customized encoding scheme, the time value event should follow the pitch, otherwise the generated token sequence cannot be correctly decoded into a MIDI file. According to repeated experiments, using a language model to train a REMI-like token sequence representation without grammatical structure constraints is very unstable.

[0072] Therefore, in this embodiment, a structure-aligned multi-track multi-instrument representation is adopted, using I x J musical events e ij (i = 1, 2,..., I; j = 1,..., J) to form a matrix W = [e ij to represent the content of a piece of symbolic music. In the experiment, it is stipulated that the event e i1 = ins i , e i2 = time i, e i3 = pit i, e i4 = dur i represent the instrument attribute, start time attribute, pitch attribute, and time value attribute of the i-th note in sequence. Therefore, the music representation W can be expressed as:

[0073] Formula (1)

[0074] The detailed design of the music representation is as follows:

[0075] (1) ins i

[0076] ins i represents the melody, chord label, and instrument type of the i-th note. Notes belonging to the same instrument are merged into one track. Chords are extracted through the open-source algorithm of the Compound word transformer. In this embodiment, all 16 instrument categories are adopted according to the MIDI protocol, and the guitar is further divided into acoustic guitar and electric guitar, adding chord and label, melody label, and unknown label, for a total of 20 categories. All existing instruments and sounds are classified into these 20 categories. The more specific categories are shown in the following table;

[0077] Table 1 Instrument Type Table

[0078]

[0079] (2) time i

[0080] The start time of the i-th note (chord) relative to the previous chord, with the thirty-second note as the time unit, ranging from 0 to 64 (at most 2 bars under the 4 / 4 time signature).

[0081] (3)pit i

[0082] The pitch of the i-th note, ranging from C1 to B6, or the type of the i-th chord. It includes 12 root notes and 9 types.

[0083] (4)dur i

[0084] The duration of the i-th note, in sixteenth notes, ranging from 1 to 16 (at most 1 measure in 4 / 4 time signature). Chord events are filled with the "PAD" marker.

[0085] Chord progressions naturally form a set of time intervals. Based on the semi-permutation invariant of multi-track music, we arrange and flatten the note units within each chord time interval into a sequence. This flattening method combined with a bidirectional encoder model can naturally establish harmony between multiple tracks, and each instrument category can learn its note distribution characteristics.

[0086] Finally, according to the above method, the music content information in the MIDI file is encoded into a music content representation matrix, named Wm.

[0087] Step S2 is the process of masking the original data set according to the preset chord condition rules to form an original masked data pair. Among them, masking the original data set according to the preset chord condition rules to form an original masked data pair includes:

[0088] Display the chord music events in the music content representation of the original data set, and mask the non-chord music events in the music content representation; among them, the masked music events are marked with PAD.

[0089] For example, in a specific embodiment, for a specific task, a dynamic event-level masking strategy is adopted on W m to mask the music events e that need to be learned and predicted, and display other events. The events that need to be displayed are called conditional events and are displayed normally. This process is completed by a trained masked language model (MLM), or more precisely, a masked noise model is trained in the diffusion framework to reconstruct the masked event information. If there is no music event at a certain position, it is predicted as the PAD marker. ij In this embodiment, the above-mentioned chord condition rules are adopted: during the training process, on W

[0090] m ​The chord events in it are displayed, making the chord events conditional. When training the music discrete diffusion model and subsequently using the music discrete diffusion model, the masking strategy for text-to-symbol music generation is as shown in formula (2), W c fully displayed, W m The chord events in the matrix are displayed, and other events are masked.

[0091] Formula (2) is:

[0092] W c = [label Tempo_Andante jazz... form VERSE fragment_1] (fully displayed)

[0093] W m =

[0094] [ins chord , time1, pit1, PAD], (These 4 events in this line are chord-related events and are all displayed)

[0095] [ins melody , time2, pit2, dur2], (These 4 events in this line are all masked)

[0096] [ins3, time3, pit3, dur3], (These 4 events in this line are all masked) ...

[0097] [ins chord , time i , pit i , PAD], (These 4 events in this line are chord-related events and are all displayed)

[0098] [ins i+1 , time i+1 , pit i+1 , dur i+1 , (These 4 events in this line are all masked) ...

[0100] It should be noted that if the first word in this line is inschord, then these 4 events in this line are chord-related events and are all displayed; otherwise, these 4 events in this line are all masked; the masked events are set to the PAD identifier. According to the method proposed in this embodiment, as long as a specific masking strategy is implemented, any desired instrument and event e can be generated under the control of the given information. ij ​。The greatest value of this representation and masking method lies in supporting human-computer interaction for AI-generated content (AIGC). Specific interaction methods are as shown in the appendix Figure 2 For example, users can propose to upload a segment of missing musical notes and input their requirements. The method in this embodiment then generates appropriate music according to the requirements to fill in the gaps; users can also propose modification requirements for originally complete musical notes, or even just input a piece of text, and automatically match a corresponding segment of musical notes to generate music, etc.

[0101] Step S3 is a process of repeatedly training the original masked data pair in a preset music discrete diffusion framework to form a music discrete diffusion model; among them, in the process of repeatedly training the original masked data pair in a preset music discrete diffusion framework to form a music discrete diffusion model,

[0102] S31: Train the music discrete framework with the original masked data through normalized embedding learning; among them, the music discrete framework includes a diffusion model framework, a decoding model framework, and a word embedding module;

[0103] The norm loss of the diffusion model framework is calculated by a preset BERT model, and the attention mechanism of the diffusion model framework is data-related to the word embedding vectors involved in the word embedding module;

[0104] The objective function during the training of the music discrete diffusion framework is the sum of the loss function of the diffusion model framework, the loss function of the decoding model framework, and the loss function of the word embedding module.

[0105] In a specific embodiment, the variational lower bound L is derived according to the standard diffusion process VLB 。The first difference between this method and the prior art is the normalized embedding learning, that is, using the anchor loss~\cite{gao2022difformer} to replace the round loss of Diffuseq in the prior art to further alleviate the collapse phenomenon of the denoising objective. The anchor loss uses the prediction of the model after training as the input, so that the word embeddings are more informative and more distinguishable from each other. Secondly, different from Difformer, the optimization term of the decoding process, named L nll 。

[0106] To cooperate with the denoising process of masking, the final objective function is as shown in formula (3):

[0107] —Formula (3);

[0108] Among them, represents minimizing the variational lower bound, represents giving the independent variable to minimize, is the abbreviation of variational lower bound;

[0109] Z0 represents the initial state;

[0110] represents the state at step t;

[0111] represents the non-classifier discrete diffusion model;

[0112] represents the hidden layer state generated (output) by the non-classifier discrete diffusion model;

[0113] represents the three-dimensional word embedding;

[0114] represents the conditional probability of the parameter ;

[0115] represents for the parameter given the input the conditional probability of the output ;

[0116] represents the feature matrix.

[0117] Among them, the first term guides the diffusion process of the diffusion model framework, the second term is responsible for the decoder process of the decoding model framework, and the third term trains the word embedding module using the anchor loss. The first term (L2 norm loss) is calculated by BERT, and its attention mechanism also considers the explicit word embedding vector. Therefore, the gradient of backpropagation will affect the learning of the correlation with the conditional and generated vectors, and jointly optimize the word embedding module.

[0118] More specifically, during the training process, the work performed by the model is the same as the workflow of the following trained music discrete diffusion model when generating music for the user's music content representation and conditional matrix.

[0119] Step S4 is a process of preprocessing the music frequency band and music requirements input by the user to form a music content representation and a conditional matrix, inputting the music content representation and the conditional matrix into the music discrete diffusion model, and enabling the music discrete diffusion model to perform music generation to output a target music score; that is, the steps performed during training can refer to the steps when the trained music discrete diffusion model is applied, as shown in the following steps S421 - S428 and their embodiments.

[0120] Among them, the preprocessing of the music frequency band and music requirements input by the user to form a music content representation and a conditional matrix includes:

[0121] S411: Segment the music frequency band to form music segments;

[0122] S412: Encode each music segment to form a music content representation; obtain the element information of the music segment, combine the element information with the music requirements to obtain music combination information, and encode the music combination information to form a conditional matrix.

[0123] The music discrete diffusion model performs music generation to output a target music score, including:

[0124] S421: Concatenate the music content representation and the conditional matrix of each music segment to form a combined representation;

[0125] S422: Perform word embedding in the continuous domain on the combined representation using a preset three-dimensional word embedding space to form a combined word embedding vector;

[0126] S423: Use a preset layer normalization module to perform Markov transfer on the combined word embedding vector to obtain the noise initial state of the combined word embedding vector;

[0127] S424: Dynamically add noise to the masked part in the noise initial state through a preset diffusion module to generate a hidden state vector;

[0128] S425: Apply position word embedding and time step word embedding to the hidden state vector to generate a diffusion hidden vector, and then perform an inverse operation on the diffusion hidden vector through a preset bidirectional encoder network to obtain a target hidden vector;

[0129] S426: Perform category mapping on the target hidden vector through a preset multi-head output module to obtain a target music content representation matrix;

[0130] S427: Decode the target music content representation matrix to generate a target music score frequency band;

[0131] S428: Concatenate all the target music score frequency bands in chronological order to form a target music score.

[0132] Figure 3 A music discrete diffusion model is shown, which includes a discrete diffusion module for symbolic music generation and a deep learning framework, Figure 3The left part is the diffusion process, where the events to be learned and generated are masked and participate in diffusion and reconstruction, while the conditional events (such as chord events and rhythm events) remain unchanged; Figure 3 The right part is the encoder framework, which is a parameterized model of the discrete diffusion framework.

[0133] In a specific implementation, as shown on the left side of Figure 3, the Markov process composed of circular nodes is the standard Denoising Diffusion Probabilistic Models (DDPM), and other modules participate in the conversion process between the discrete and continuous spaces of the diffusion process.

[0134] To achieve parallel generation of long segments, the music score is segmented into segments of 120 notes each to obtain the music content representation matrix W of each segment m n and record the segment identifier fragment_n in the corresponding Wc, where n is the serial number of the identified segment. First, it is necessary to splice Wc and W m n to obtain the combined representation W of each segment c+m n;

[0135] After that, a learnable three-dimensional word embedding module as shown in the appendix Figure 4 is used to map the combined representation W in the discrete domain c+m n to the word embedding space EMB (W c+m n ) in the continuous domain to form a combined word embedding vector. As Figure 4 shown, in this embodiment, the size of the word embedding space is designed as [I, J, D], where I represents the note plus the description phrase, J represents the attribute, and D represents the word embedding vector, and I and J can be permuted.

[0136] In the music score, the frequencies of common events and rare events are extremely unbalanced, which has an adverse impact on the diversity of the content generated by the model. To solve this problem, in this implementation, a learnable layer normalization module (named LN) is adopted on the word embedding space to input the discrete event information into the initial state Z0 of the diffusion process by extending the standard forward chain to a new Markov transition process.

[0137] That is, mask positions are arranged for W m according to the specific task, where only noise is dynamically applied to the mask vector of Zt, and the visible part of Zt remains unchanged throughout the diffusion process.

[0138] Then perform the appendixFigure 3 The process of the discrete part, where the masked part in the initial state of the noise is dynamically noise-added through a preset diffusion module to generate a hidden state vector, that is, the process of converting Z0 to Zt; a non-classifier discrete diffusion model can be pre-trained for this process to complete the inverse process, and the applied formula (4) is as follows:

[0139] Formula (4)

[0140] where and are the parameterizations of the predicted mean and standard deviation in the forward process ; is the predicted mean, is the predicted standard deviation, is the state at step t. is the conditional probability of the forward process.

[0141] Then, a positional word embedding and a time step word embedding are applied to the hidden state vector to generate a diffusion hidden vector, and then an inverse operation is performed on the diffusion hidden vector through a preset bidirectional encoder network to obtain a target hidden vector. A bidirectional encoder network (BERT) can be used to simulate , by gradually transforming to , and then to the final target to complete the inverse operation process, thereby obtaining the target hidden vector, that is, the target .

[0142] Finally, a preset multi-head output module performs class mapping on the target hidden vector to obtain a target music content representation matrix. The multi-head output module is as shown in the appendix Figure 5 shown, which is connected to the end of the framework in the above appendix Figure 3 . The multi-head output module is used to accurately map the encoder hidden state (target hidden vector) to a certain attribute category. For example, the hidden vector belonging to the pitch dimension (the third column in W m n ) in the target hidden vector must be mapped to one of the 72 pitch categories, and cannot be classified as the start time or other categories. The multi-head output module is as Figure 5As shown in the figure, it consists of multiple linear layers and a classification mask. The linear layer (the cube part in Figure 5) uses a fully connected network to convert the hidden state vector output by BERT into a probability sequence with the same length as the vocabulary (vocabulary size), and this probability sequence reflects the probability scores of each word (Token). In this specific embodiment, the hidden state tensor is divided into 4 blocks by column (attribute), and 4 linear layers are arranged to predict their respective probability scores. The conditional information W c is connected to the first column (instrument attribute) and input into the first linear layer. Then, the obtained probability sequence is multiplied bit by bit by the 0-1 mask sequence Mask j . The token positions where the probability score sequence is set to 0 will not be selected. Thus, sorting by probability level, the predicted contents with high probabilities are concatenated according to the grammatical structure of the music representation to obtain the target music content representation matrix generated by the model .

[0143] Then, the target music content representation matrix is decoded to generate target music score frequency bands, and all the target music score frequency bands are concatenated according to time to form a target music score

[0144] Since the generated target music score is MIDI, which is a music score without sound, in this specific embodiment, the Fluidsynth library is used to load a music sampling file in the SoundFont2 format for the target music score, so as to export music audio and save and download it as a WAV file

[0145] Thus, the trained music discrete diffusion model is used for music generation to output a target music score

[0146] The training process of this music discrete diffusion model is similar to the usage process of the above music discrete diffusion model. Just replace the music frequency bands and music requirements input by the user above with the above original data set and its original masked data pairs, which will not be elaborated here

[0147] As described above, the controllable generation method of symbolic music provided by the present invention can significantly reduce the sequence length and retain the time and multi-track spatial information by using the chords of music content as the timeline; the original data set is masked according to the preset chord condition rules to form an original masked data pair, strengthening the restrictions of chords in composition and improving the effectiveness of feature learning and generation of pop music; when training and using the music discrete diffusion model, first splice the matrix, then perform discrete noise addition, and finally perform decoding. After decoding, probability mapping of the hidden vector is also required, which can not only achieve the richness of controllable music generation, but also improve the reliability and rationality

[0148] The following introduces a controllable symbol music generation system 100 provided by an embodiment of the present application. The controllable symbol music generation system described below can be correspondingly referred to the controllable symbol music generation method described above.

[0149] As Figure 6 shown, the present invention also provides a controllable symbol music generation system 100, which can implement the controllable symbol music generation method as described above. The system includes:

[0150] A music discrete diffusion model 101; wherein, the music discrete diffusion model 101 is learned and generated by repeatedly training a preset music discrete diffusion framework with the original masked data pair; wherein, obtaining the original masked data pair includes: preprocessing the pre-obtained music content and annotation information through a preset representation module to form an original data set; performing masking processing on the original data set according to a preset chord condition rule to form an original masked data pair;

[0151] The music discrete diffusion model 101 is used to receive the music content representation and the conditional matrix formed by preprocessing the user input music frequency band and music requirements, and perform music generation according to the music content representation and the conditional matrix to output a target music score.

[0152] The system further includes a preprocessing unit 102, which is used to preprocess the user input music frequency band and music requirements to form a music content representation and a conditional matrix, including:

[0153] Segmenting the music frequency band to form music segments;

[0154] Encoding each music segment to form a music content representation; obtaining the element information of the music segment, combining the element information with the music requirements to obtain music combination information, and encoding the music combination information to form a conditional matrix.

[0155] The controllable symbol music generation system 100 provided by the present invention performs calculations in the same way as the aforementioned controllable symbol music generation method. For a more specific implementation process, reference can be made to the specific embodiments of the above-mentioned controllable symbol music generation method.

[0156] As described above, the controllable symbol music generation system 100 provided by the present invention receives the music content representation and the conditional matrix formed by preprocessing the user input music frequency band and music requirements through the music discrete diffusion model, and performs music generation according to the music content representation and the conditional matrix to output a target music score, improving the reliability of symbol music generation.

[0157] See Figure 7, This figure is a schematic structural diagram of an electronic device provided by an embodiment of the present application, including: a memory 11 for storing computer programs; a processor 12 for implementing the steps of the controllable generation method of symbolic music described in any of the above method embodiments when executing the computer programs. In this embodiment, the device can be an in-vehicle computer, a PC (Personal Computer), or a terminal device such as a smart phone, a tablet computer, a palm computer, a portable computer, etc. The device may include a memory 11, a processor 12, and a bus 13. Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 11 may be an internal storage unit of the device in some embodiments, such as the hard disk of the device. The memory 11 may also be an external storage device of the device in other embodiments, such as a plug-in hard disk equipped on the device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 11 may also include both the internal storage unit and the external storage device of the device. The memory 11 can be used not only to store application software installed on the device and various types of data, such as program codes for executing the fault prediction method, etc., but also to temporarily store data that has been output or will be output. The processor 12 may be a central processing unit (CPU) in some embodiments. The processor 12 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments, for running program codes stored in the memory 11 or processing data, such as codes for executing the controllable generation method of symbolic music. The bus 13 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 7 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus. Further, the device may also include a network interface 14, and the network interface 14 may optionally include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the device and other electronic devices.

[0158] Optionally, the device may further include a user interface 15, which may include a display, an input unit such as a keyboard, and optionally, the user interface 15 may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the device and to display a visual user interface. Those skilled in the art can understand that Figure 7 The structures shown do not constitute a limitation on the device, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0159] Embodiments of the present invention also provide a computer-readable storage medium, which may be non-volatile or volatile. The storage medium stores a computer program, and when the computer program is executed by a processor, it realizes:

[0160] Preprocessing the pre-acquired music content and annotation information through a preset characterization module to form an original data set;

[0161] Performing masking processing on the original data set according to a preset chord condition rule to form an original masked data pair;

[0162] Feeding the original masked data pair into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model;

[0163] Performing preprocessing on the music frequency band and music requirements input by the user to form a music content representation and a condition matrix, and inputting the music content representation and the condition matrix into the music discrete diffusion model, so that the music discrete diffusion model performs music generation to output a target music score.

[0164] Specifically, the specific implementation method when the computer program is executed by the processor may refer to the description of the relevant steps in the controllable generation method of symbolic music in the embodiments, and will not be elaborated here.

[0165] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0166] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0167] In addition, in each embodiment of the present invention, each functional module may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0168] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0169] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be regarded as limiting the claimed invention.

[0170] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. The terms such as "second" are used to denote names and do not denote any particular order.

[0171] It should be noted that the embodiments in this specification are all described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for methods, devices, electronic devices, and computer-readable storage media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The methods, devices, electronic devices, and media described above are only illustrative. The units described as separation components may or may not be physically separated. The components indicated as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0172] The controllable generation method, system, electronic device and medium of symbolic music according to the present invention have been described above by way of example with reference to the accompanying drawings. However, those skilled in the art should understand that various improvements can be made to the above-mentioned controllable generation method, system, electronic device and medium of symbolic music proposed by the present invention without departing from the content of the present invention. Therefore, the protection scope of the present invention should be determined by the content of the appended claims.

Claims

1. A controllable method for generating symbolic music, characterized in that: include: Preprocessing the pre-acquired music content and annotation information through a preset representation module to form an original data set; The original data set is masked according to a preset chord condition rule to form an original masked data pair; the chord condition rule includes a relative time encoding and a masking method based on a chord condition; wherein the masking according to the preset chord condition rule to form an original masked data pair includes: displaying chord music events represented by music content in the original data set, and masking non-chord music events represented by the music content; wherein the masked music events are marked with PAD; Sending the original masked data pair into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model; The music frequency band and music demand input by the user are pre-processed to form a music content representation and a condition matrix, and the music content representation and the condition matrix are input into the music discrete diffusion model, so that the music discrete diffusion model performs music generation to output a target music score; wherein the music discrete diffusion model performs music generation to output a target music score, including: The music content representation and condition matrix of each music clip are concatenated to form a merged representation; Using a preset three-dimensional word embedding space to perform word embedding of the combined representation in a continuous domain to form a combined word embedding vector; Using a preset layer normalization module to perform Markov shift on the merged word embedding vector to obtain a noise initial state of the merged word embedding vector; Performing dynamic noise addition processing on the masked part in the initial noise state through a preset diffusion module to generate a hidden state vector; Applying position word embedding and time step word embedding to the hidden state vector to generate a diffuse hidden vector, and then performing an inverse operation on the diffuse hidden vector through a preset bidirectional encoder network to obtain a target hidden vector; Performing category mapping on the target hidden vector through a preset multi-head output module to obtain a target music content representation matrix; Decoding the target music content representation matrix to generate a target music score frequency band; All target music score frequency bands are concatenated in chronological order to form a target music score.

2. The controllable generation method of symbolic music according to claim 1, characterized in that: The pre-processing of the music frequency band and music requirements input by the user to form a music content representation and a condition matrix includes: Segmenting the music frequency band to form music clips; Each music clip is encoded to form a music content representation; element information of the music clip is obtained, the element information is combined with the music requirement to obtain music combination information, and the music combination information is encoded to form a condition matrix.

3. The controllable generation method of symbolic music as claimed in claim 2, characterized in that: Also includes: The target music is generated by loading a music sampling file into the target music score through a preset audio library.

4. The controllable generation method of symbolic music as claimed in claim 1, characterized in that: The pre-processing of the pre-acquired music content and annotation information by a preset characterization module to form an original data set includes: The chords of the music content are used as a timeline, the start time of the notes is automatically quantized in units of 32-notes, the duration of the notes is quantized in units of 16-notes, self-correction is performed, and empty tracks and empty notes are self-deleted to form standard music content, and the standard music content is self-annotated using the annotation information to form a music content set; wherein the music content set contains various music events of the standard music content, and the music events are annotated with music attributes; Combining all music events of a piece of music content into a matrix to form a music content representation, and combining the music content representations of all the music contents to form a music content representation data set; The music content representation dataset is used as the original dataset.

5. The controllable generation method of symbolic music as claimed in claim 4, characterized in that: In the process of sending the original masked data pair into a preset music discrete diffusion framework for repeated training to form a music discrete diffusion model, The music discrete diffusion framework is trained using original masked data through normalized embedding learning; wherein the music discrete diffusion framework includes a diffusion model framework, a decoding model framework and a word embedding module; The norm loss of the diffusion model framework is calculated by a preset BERT model, and the attention mechanism of the diffusion model framework is data-dependent with the word embedding vector involved in the word embedding module; The objective function when training the music discrete diffusion framework is the sum of the loss function of the diffusion model framework, the loss function of the decoding model framework and the loss function of the word embedding module.

6. A controllable generation system of symbolic music, characterized in that: A controllable method for generating symbolic music as claimed in any one of claims 1 to 5 is implemented, wherein the system includes a music discrete diffusion model; wherein: The music discrete diffusion model is generated by learning through repeated training of original masked data pairs using a preset music discrete diffusion framework; wherein obtaining the original masked data pairs includes: preprocessing the pre-acquired music content and annotation information through a preset representation module to form an original data set; performing masking processing on the original data set according to a preset chord condition rule to form an original masked data pair; The music discrete diffusion model is used to receive the music content representation and condition matrix formed by pre-processing the music frequency band and music requirements input by the user, and generate music according to the music content representation and condition matrix to output the target music score.

7. An electronic device, characterized in that: The device includes: a processor, a memory and a system bus; the processor and the memory are connected via the system bus; the memory is used to store one or more programs, and the one or more programs include instructions, which, when executed by the processor, enable the processor to execute the steps in the controllable generation method of symbolic music described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for controllable generation of symbolic music described in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Generation method of music works, training method of music generation model and equipment thereof

    CN116704980A

  • Automated Music Composition and Generation System and Method

    US20230326436A1