Music generation method, music generation device, electronic equipment and storage medium
By acquiring musical structure information to generate the first musical score sequence and converting it into target audio data, the problem of uncontrollable musical structure in deep learning music generation methods is solved, and high-quality music generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2024-10-21
- Publication Date
- 2026-04-24
AI Technical Summary
Existing deep learning-based music generation methods lack control over the music structure, resulting in music that lacks chord information and other professional music theory notations, leading to poor music quality and failure to meet user needs.
By acquiring music structure information, including the number of musical measures, the number of notes in each musical measure, and chord markings, a first musical score sequence is generated and converted into target audio data to output the target music, ensuring that the music structure matches the user's needs.
It achieves controllability in music generation effects, and the generated music is structurally consistent with user needs, containing rich chord information and other professional music theory markings, thus improving the quality of music generation.
Smart Images

Figure CN121922091A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of audio processing and artificial intelligence, and in particular to music generation methods, music generation devices, electronic devices, and storage media. Background Technology
[0002] Currently, music generation methods mainly include rule-based generation methods, probabilistic model-based generation methods, and deep learning-based generation methods. Given the significant limitations of rule-based and probabilistic model-based generation methods, and the rapid development of artificial intelligence technology, deep learning-based music generation is currently the main research direction in this field.
[0003] For music generation methods based on deep learning, how to make the music generated by the model more controllable and contain rich information is a problem faced by researchers in this field. Summary of the Invention
[0004] To overcome the problems existing in related technologies, this disclosure provides a music generation method, a music generation device, an electronic device, and a storage medium.
[0005] According to a first aspect of the present disclosure, a music generation method is provided, comprising: acquiring music structure information, the music structure information including the number of musical measures, the number of notes in each musical measure, and a chord notation corresponding to each musical measure, the chord notation representing the chord corresponding to the corresponding musical measure; obtaining a first musical score sequence based on the music structure information, the first musical score sequence including multiple musical measures, multiple notes in each musical measure, and a chord notation corresponding to each musical measure, the number of multiple musical measures being consistent with the number of musical measures, the number of multiple notes corresponding to each musical measure being consistent with the number of notes, and the multiple notes corresponding to each musical measure being consistent with the chord corresponding to the corresponding musical measure in musical expression; and converting the first musical score sequence into target audio data, the target audio data supporting being output as target music.
[0006] In one embodiment, obtaining the music structure information includes: obtaining input text for generating music, the input text including a unit musical beat, a musical measure value, a number of musical measures, and a chord corresponding to each musical measure to describe the music structure; performing text conversion on the input text according to a preset musical notation to obtain the music structure information, wherein the number of notes included in each musical measure in the music structure information corresponds to the unit musical beat and the musical measure value.
[0007] In one embodiment, the music structure information further includes music rhythm pattern information and unit rhythm information, wherein the rhythm pattern information and the unit rhythm information correspond to the unit musical beat and the musical measure time value; obtaining the first musical score sequence based on the music structure information includes: performing feature encoding on the number of notes included in each musical measure and the chord markings corresponding to the corresponding musical measure in the plurality of musical measures to obtain a first feature sequence, wherein the first feature sequence includes a plurality of first features arranged in chronological order, each first feature representing the chord markings and the number of notes corresponding to the corresponding musical measure; performing feature encoding on the rhythm pattern information and the unit rhythm information to obtain a second feature, wherein the second feature represents the music rhythm pattern and unit rhythm; obtaining the first musical score sequence based on the first feature sequence and the second feature, wherein the representation of the notes in the first musical score sequence corresponds to the rhythm pattern and the unit rhythm.
[0008] In one embodiment, obtaining the first musical score sequence based on the first feature sequence and the second feature includes: for the earliest chronological first feature in the first feature sequence, generating a third feature based on the first feature, the third feature representing multiple notes and chord markings corresponding to the multiple notes, the multiple notes being consistent in number with the number of notes represented by the corresponding first feature, and consistent in musical expression with the chord markings represented by the corresponding first feature; for the first features other than the earliest chronological first feature in the first feature sequence, generating a third feature based on the first feature and a first historical feature, the first historical feature being a previously generated third feature corresponding to the previous first feature chronologically adjacent to the corresponding first feature, the third feature representing multiple notes, the multiple notes being consistent in number with the number of notes represented by the corresponding first feature, consistent in musical expression with the chord markings represented by the corresponding first feature, and having a musical expression relationship with the first historical feature; determining the third features corresponding to the multiple first features as a third feature sequence, the multiple third features in the third feature sequence being arranged chronologically; and obtaining the first musical score sequence based on the third feature sequence and the second feature.
[0009] In one embodiment, obtaining the first musical score sequence based on the third feature sequence and the second feature includes: for the earliest chronological third feature in the third feature sequence, generating a fourth feature based on the third feature, the fourth feature representing multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes, wherein the additional musical markings corresponding to the multiple notes are musically consistent with the chord markings represented by the corresponding third feature, and the additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note; for the third features other than the earliest chronological third feature in the third feature sequence, generating a fourth feature based on the third feature and a second historical feature, the second historical feature... The fourth feature is the one that has been generated and is adjacent to the previous third feature in terms of temporal sequence. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. The fourth feature is also related to the second historical feature in terms of musical expression. The fourth features corresponding to the multiple third features are determined as a fourth feature sequence. The multiple fourth features in the fourth feature sequence are arranged in temporal sequence. The first musical score sequence is obtained based on the fourth feature sequence and the second feature.
[0010] According to a second aspect of the present disclosure, a music generation apparatus is provided, comprising: an acquisition unit, configured to acquire music structure information, the music structure information including the number of musical measures, the number of notes in each musical measure, and a chord notation corresponding to each musical measure, the chord notation representing a chord corresponding to a corresponding musical measure; a processing unit, configured to obtain a first musical score sequence based on the music structure information, the first musical score sequence including multiple musical measures, multiple notes in each musical measure, and a chord notation corresponding to each musical measure, the number of multiple musical measures being consistent with the number of musical measures, the number of multiple notes corresponding to each musical measure being consistent with the number of notes, and the multiple notes corresponding to each musical measure being consistent with the chord corresponding to the corresponding musical measure in musical expression; and a conversion unit, configured to convert the first musical score sequence into target audio data, the target audio data supporting output as target music.
[0011] In one embodiment, the acquisition unit acquires music structure information in the following manner: acquiring input text for generating music, the input text including unit musical beats, musical measure durations, number of musical measures, and chords corresponding to each musical measure for describing the music structure; and performing text conversion on the input text according to a preset musical notation to obtain the music structure information, wherein the number of notes included in each musical measure in the music structure information corresponds to the unit musical beats and the musical measure durations.
[0012] In one embodiment, the music structure information further includes music rhythm pattern information and unit rhythm information, wherein the rhythm pattern information and the unit rhythm information correspond to the unit musical beat and the musical measure time value; the processing unit obtains a first musical score sequence based on the music structure information in the following manner: feature encoding is performed on the number of notes included in each musical measure and the corresponding chord markings of the musical measures to obtain a first feature sequence, wherein the first feature sequence includes multiple first features arranged in chronological order, each first feature representing the chord markings and the number of notes corresponding to the corresponding musical measure; feature encoding is performed on the rhythm pattern information and the unit rhythm information to obtain a second feature, wherein the second feature represents the music rhythm pattern and unit rhythm; based on the first feature sequence and the second feature, the first musical score sequence is obtained, wherein the representation of notes in the first musical score sequence corresponds to the rhythm pattern and the unit rhythm.
[0013] In one embodiment, the processing unit obtains the first musical score sequence based on the first feature sequence and the second feature in the following manner: For the earliest chronological first feature in the first feature sequence, a third feature is generated based on the first feature. The third feature represents multiple notes and chord markings corresponding to the multiple notes. The number of multiple notes is consistent with the number of notes represented by the corresponding first feature, and the musical expression is consistent with the chord markings represented by the corresponding first feature. For the first features other than the earliest chronological first feature in the first feature sequence, a third feature is generated based on the first feature and a first historical feature. The first historical feature is the previously generated third feature corresponding to the previous first feature whose chronological order is adjacent to the corresponding first feature. The third feature represents multiple notes. The number of multiple notes is consistent with the number of notes represented by the corresponding first feature, and the musical expression is consistent with the chord markings represented by the corresponding first feature. The third feature is also associated with the first historical feature in terms of musical expression. The third features corresponding to the multiple first features are determined as a third feature sequence, and the multiple third features in the third feature sequence are arranged chronologically. The first musical score sequence is obtained based on the third feature sequence and the second feature.
[0014] In one embodiment, the processing unit obtains the first musical score sequence based on the third feature sequence and the second feature in the following manner: For the earliest chronological third feature in the third feature sequence, a fourth feature is generated based on the third feature. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are musically consistent with the chord markings represented by the corresponding third feature. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. For the third features other than the earliest chronological third feature in the third feature sequence, a fourth feature is generated based on the third feature and the second historical feature. The historical feature is the fourth feature that has been generated and is adjacent to the previous third feature in terms of temporal sequence. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. The fourth feature is also related to the second historical feature in terms of musical expression. The fourth features corresponding to the multiple third features are determined as a fourth feature sequence. The multiple fourth features in the fourth feature sequence are arranged in temporal sequence. The first musical score sequence is obtained based on the fourth feature sequence and the second feature.
[0015] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: execute the music generation method described in the first aspect or any embodiment of the first aspect.
[0016] According to a fourth aspect of the present disclosure, a storage medium is provided, the storage medium storing instructions that, when executed by a processor, enable the processor to perform the music generation method described in the first aspect or any embodiment of the first aspect.
[0017] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: Obtaining music structure information, which includes the number of musical measures, the number of notes in each musical measure, and the chord notation corresponding to each musical measure. Based on the music structure information, a first musical score sequence is obtained. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord notation corresponding to each musical measure. The number of multiple musical measures is consistent with the number of musical measures, the number of multiple notes corresponding to each musical measure is consistent with the number of notes, and the multiple notes corresponding to each musical measure are musically consistent with the chords corresponding to the corresponding musical measure. The first musical score sequence is converted into target audio data, which can be output as target music. Through this disclosure, a first musical score sequence corresponding to the music structure information is generated based on the music structure information. Target music data that can be output as target music is obtained based on the first musical score sequence, making the music generation effect more controllable, thereby obtaining target music that matches the user's needs in terms of music structure.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0020] Figure 1 This is a flowchart illustrating a music generation method according to an exemplary embodiment.
[0021] Figure 2 This is a flowchart illustrating a method for obtaining music structure information according to an exemplary embodiment.
[0022] Figure 3 This is a flowchart illustrating a method for obtaining a first musical score sequence based on musical structure information, according to an exemplary embodiment.
[0023] Figure 4 This is a flowchart illustrating a method for obtaining a first musical score sequence based on a first feature sequence and a second feature, according to an exemplary embodiment.
[0024] Figure 5 This is a flowchart illustrating a method for obtaining a first musical score sequence based on a third feature sequence and a second feature, according to an exemplary embodiment.
[0025] Figure 6 This is a block diagram illustrating a music generation apparatus according to an exemplary embodiment.
[0026] Figure 7 This is a block diagram illustrating an apparatus for music generation according to an exemplary embodiment.
[0027] Figure 8 This is a block diagram illustrating an apparatus for music generation according to an exemplary embodiment. Detailed Implementation
[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure.
[0029] The music generation method provided in this disclosure is applied to scenarios where corresponding target music is generated based on music structure information input by the user.
[0030] Currently, music generation methods mainly include rule-based generation methods, probabilistic model-based generation methods, and deep learning-based generation methods. Given the significant limitations of rule-based and probabilistic model-based generation methods, and the rapid development of artificial intelligence technology, deep learning-based music generation is the main research direction in this field.
[0031] In related technologies, when generating music from user-input text using deep learning methods, the resulting music often lacks structural control. For example, tempo, rhythm, number of measures, and the number of notes per measure frequently fail to meet user needs. Furthermore, deep learning-based music generation often simply produces a basic melody composed of notes, lacking chord markings and other musical notation (such as raising a note's degree by an octave or adding annotations). Moreover, these technologies directly extract features and generate music based on user-input prompts, which are often non-professional musical representations lacking in professional music theory information. This fails to accurately guide music generation, resulting in music of poor quality and significant discrepancies with user expectations.
[0032] In one example, a technique exists for music generation based on a chat-generated pre-trained converter (chatgpt). However, chatgpt can only generate musical note text sequentially from the user's input prompts, producing a short melody. Users cannot provide the chatgpt model with overall structural information about the music, such as the chord progressions, resulting in uncontrollable musical structure in the generated music. Furthermore, chatgpt can only generate simple melodies, lacking chord information, tempo, and other markings. If evaluated using professional music content / score standards, the music text generated by chatgpt contains too little information. In summary, the technique of music generation based on chatgpt suffers from both uncontrollable musical structure and insufficient information content in the generated music.
[0033] In another example, related technologies include methods for music generation based on songcomposers. Songcomposers lack music theory information (such as chords and other common compositional notations) and cannot input the overall structure of the model music. The overall generation process also uses absolute time values, making it difficult to unify the rhythm of the entire piece. This results in uncontrollable music structure during songcomposer generation and a lack of information contained in the generated music.
[0034] In view of this, this disclosure proposes a method for music generation. After obtaining music structure information, a first musical score sequence is obtained based on the number of musical measures, the number of notes in each musical measure, and the chord notation corresponding to each musical measure. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord notation corresponding to each musical measure. The number of multiple musical measures is consistent with the number of musical measures, the number of notes in each musical measure is consistent with the number of notes, and the multiple notes in each musical measure musically correspond to the chords corresponding to the corresponding musical measure. The first musical score sequence is converted into target audio data, which can be output as target music. Through this disclosure, a first musical score sequence corresponding to the music structure information is generated based on the music structure information. Target music data that can be output as target music is obtained based on the first musical score sequence, making the music generation effect more controllable, thereby obtaining target music that matches the user's needs in terms of music structure.
[0035] Figure 1 This is a flowchart illustrating a music generation method according to an exemplary embodiment. Figure 1 As shown, the method includes the following steps.
[0036] In step S101, music structure information is obtained, which includes the number of music measures, the number of notes in each music measure, and the chord mark corresponding to each music measure. The chord mark represents the chord corresponding to the music measure.
[0037] In step S102, a first musical score sequence is obtained based on the music structure information. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord markings corresponding to each musical measure. The number of multiple musical measures is consistent with the number of musical measures. The number of multiple notes corresponding to each musical measure is consistent with the number of notes. The multiple notes corresponding to each musical measure are consistent with the chords corresponding to the corresponding musical measure in terms of musical expression.
[0038] In step S103, the first musical score sequence is converted into target audio data, which can be output as target music.
[0039] In this embodiment, the music structure information is a structured text string, mainly including a music prefix and a music stem. The music prefix contains guiding information for the overall music (including each note of the music to be generated), guiding the generation of notes in each musical measure. The music stem contains chord information, bar lines, and note value information. Bar lines are used to divide musical measures, with each measure consisting of two bar lines. The music stem contains chord information (chord notation) and note value information (number of notes) to guide the corresponding musical measures; different musical measures can have different chord notations and note numbers.
[0040] In this embodiment, the music structure information specifies requirements for the number of musical measures and the number of notes in each musical measure of the music to be generated. This ensures that the number of musical measures in the final generated first musical score sequence matches the number of musical measures in the music structure information, and that the number of notes in each musical measure in the final generated first musical score sequence matches the number of musical measures corresponding to the musical measures in the music structure information. This ensures that the music structure of the final generated target music matches the music structure required by the user.
[0041] In this embodiment of the disclosure, the music structure information only requires the chord markings and the number of notes for each music measure, and does not include specific note information. This disclosure generates multiple notes for each music measure that are consistent in number with the number of notes required by the music structure information and in musical expression consistent with the chords pointed to by the corresponding chord markings.
[0042] In this embodiment of the disclosure, after obtaining the music structure information, a first musical score sequence is obtained based on the number of musical measures, the number of notes in each musical measure, and the chord notation corresponding to each musical measure. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord notation corresponding to each musical measure. The number of musical measures is consistent with the number of musical measures, the number of notes in each musical measure is consistent with the number of notes, and the multiple notes in each musical measure musically correspond to the chords corresponding to the corresponding musical measure. The first musical score sequence is then converted into target audio data, which can be output as target music.
[0043] This disclosure generates a first musical score sequence corresponding to the musical structure information based on the musical structure information, and obtains target music data that can be output as target music based on the first musical score sequence, making the music generation effect more controllable, thereby obtaining target music that matches the user's needs in terms of musical structure.
[0044] It is understandable that the prompt text directly entered by users during music generation is often a non-professional musical description. To ensure the quality of the generated target music, this disclosure, after obtaining the prompt text entered by the user, performs a conversion process on the prompt text, uses the converted professional musical description as music structure information, and generates music based on the music structure information. The following embodiments of this disclosure describe the method for obtaining music structure information.
[0045] Figure 2 This is a flowchart illustrating a method for obtaining music structure information according to an exemplary embodiment. Figure 2 As shown, the method includes the following steps.
[0046] In step S201, the input text for generating music is obtained. The input text includes the unit musical beat, musical measure time value, musical measure number, and chord corresponding to each musical measure, which are used to describe the musical structure.
[0047] In step S202, the input text is converted according to a preset musical notation to obtain musical structure information. The number of notes included in each musical measure in the musical structure information corresponds to the unit musical beat and the musical measure time value.
[0048] In this embodiment, after obtaining the user's input text, the user's input text is converted into a professional musical representation (i.e., musical structure information). The user's input text includes information corresponding to the overall music, such as basic unit beats and measure values, as well as information specific to individual measures, such as the number of notes in a single measure and the chord markings for that measure. It is understood that users have limited music theory knowledge and often do not understand the sound production of different chords. Furthermore, setting chord markings for different measures is too cumbersome. Therefore, this disclosure provides applications or web pages for dividing measures and setting chord markings for them. Users can divide measures, preview the sound effects of chords, and set chord markings for corresponding measures using these applications or web pages. This information, combined with the user's input text corresponding to the overall music, serves as the user's input text. By introducing applications or web pages that assist user input, the versatility of the music generation method in this disclosure is improved. This allows users without professional music theory knowledge to input relatively professional guidance information, set the musical structure, and obtain target music that meets their needs. In one example, the basic unit of time is an eighth note, and the measure value is six eighth notes, so the time signature of the entire piece is 6 / 8. It is understood that the above input for the basic unit of time and measure value is only an example, and this disclosure is not limited to this method of setting them.
[0049] In this embodiment of the disclosure, after obtaining the user's input text, the user's input text can be converted into a professional musical score representation using a preset notation method. For the user-input musical structure (basic unit beat, measure value, number of measures in the generated music, and chord markings for all measures in the entire piece), it can be converted into musical structure information (prefix: tonality, rhythmic pattern, unit rhythm; main musical part: chord information, bar lines, and time value information) using a preset musical notation method. The tonality, rhythmic pattern, and unit rhythm included in the prefix are obtained after conversion based on the user-input basic unit beat and measure value. The chord information, bar lines, and time value information of the main musical part are obtained based on the user-input measure value, number of measures, and chord markings for all measures in the entire piece. Bar lines are used to divide different measures. In one example, the preset musical notation method is ABC notation. The partial musical structure information obtained using ABC notation is denoted as follows:
[0050] L:1 / 4
[0051] M:4 / 4
[0052] K:C
[0053] |"C"z4|"Am7"z4|……
[0054] In this context, L represents the shortest note value, so L:1 / 4 indicates that the basic unit of time is a quarter note. M represents the time signature, so M:4 / 4 indicates that the entire piece is in 4 / 4 time. K represents the key, so K:C indicates that the entire piece is in C major. "|" represents a bar line, so |"C"z4| and |"Am7"z4| represent different musical measures. "·" represents a chord symbol, so the chord for |"C"z4| is C, and the chord for |"Am7"z4| is Am7. z represents a rest, so z4 indicates that there are four rests, meaning that the corresponding musical measure has four notes.
[0055] In this embodiment of the disclosure, the music structure information further includes rhythmic pattern information and unit rhythm information, where the rhythmic pattern information and unit rhythm information correspond to unit musical beats and musical measure values. The following embodiment of the disclosure describes a method for obtaining a first musical score sequence based on the music structure information.
[0056] Figure 3 This is a flowchart illustrating a method for obtaining a first musical score sequence based on musical structure information, according to an exemplary embodiment. Figure 3 As shown, the method includes the following steps.
[0057] In step S301, the number of notes in each musical measure and the corresponding chord mark of the musical measure are feature-encoded to obtain a first feature sequence. The first feature sequence includes multiple first features arranged in chronological order, and each first feature represents the chord mark and the number of notes corresponding to the musical measure.
[0058] In step S302, the rhythm pattern information and unit rhythm information are feature-encoded to obtain a second feature, which represents the rhythm pattern and unit rhythm of the music.
[0059] In step S303, a first musical score sequence is obtained based on the first feature sequence and the second feature. The representation of notes in the first musical score sequence corresponds to rhythmic patterns and unit rhythms.
[0060] In this embodiment, the music structure information includes not only the number of musical measures, the number of notes in each musical measure, and the chord notation corresponding to each musical measure, but also rhythmic information and unit rhythm information. The rhythmic information and unit rhythm information correspond to unit musical beats and musical measure durations. The rhythmic information and unit rhythm information are guiding information for the music as a whole, while the number of notes in each musical measure and the chord notation corresponding to each musical measure are guiding information for each specific musical measure. Therefore, when extracting features from the music structure information, feature encoding is performed on the rhythmic information and unit rhythm information to obtain a single second feature. Feature encoding is also performed on the number of notes in each musical measure and the chord notation corresponding to each musical measure to obtain a single first feature corresponding to each musical measure, resulting in a first feature sequence corresponding to all musical measures.
[0061] In an exemplary embodiment of this disclosure, a first feature sequence and a second feature can be obtained through a preset model in the following manner: For the prefix of the music structure information (rhythm pattern information and unit rhythm information), it is decomposed into a set of segments (patch) in the format of key vector (K): value vector (V). The music content part (the number of notes in each music measure and the corresponding chord markings in the corresponding music measures) is then divided into multiple patches according to the chord coverage, ultimately obtaining a patch sequence including all patches. An encoder is used to encode the patch sequence once to obtain the feature vector of each patch. The patch feature vector corresponding to the music structure information prefix is the second feature. A further model includes a first sub-model for further processing the feature vectors of all patches corresponding to the music content part. This first model mainly models the chord markings and the number of notes in the patch feature vectors corresponding to the music content part. A vocabulary is created for different markings and the number of notes under the corresponding chord markings, resulting in a vector corresponding to the chord markings and a vector corresponding to the number of notes. The two vectors are concatenated, and the concatenated vectors are passed through a fully connected layer (Linear layer) to obtain a feature vector 1. Simultaneously, a sequence encoding model is used to sequence encode the corresponding patch to obtain feature vector 2. Finally, these two features are added together to obtain the global feature (first feature sequence). When there are N patches, the global feature is a vector sequence of length N. The sequence encoding model can be a bidirectional encoder representation from transformers (BERT).
[0062] In this embodiment of the disclosure, after obtaining the first feature sequence, each first feature contained in the first feature sequence is processed in chronological order to obtain a third feature corresponding to each first feature. When obtaining the third feature, it is necessary to consider the historical third features corresponding to the processed first features to ensure the correlation between adjacent musical measures in the final generated first musical score sequence. The following embodiment of the disclosure describes the method for obtaining the first musical score sequence based on the first feature sequence and the second feature.
[0063] Figure 4 This is a flowchart illustrating a method for obtaining a first musical score sequence based on a first feature sequence and a second feature, according to an exemplary embodiment. Figure 4 As shown, the method includes the following steps.
[0064] In step S401, for the earliest time-series first feature in the first feature sequence, a third feature is generated based on the first feature. The third feature represents multiple notes and the chord markings corresponding to the multiple notes. The number of multiple notes is consistent with the number of notes represented by the corresponding first feature, and the musical expression is consistent with the chord markings represented by the corresponding first feature.
[0065] In step S402, for the first features other than the earliest first feature in the first feature sequence, a third feature is generated based on the first feature and the first historical feature. The first historical feature is the generated third feature corresponding to the previous first feature whose time sequence is adjacent to the corresponding first feature. The third feature represents multiple notes, and the number of notes is consistent with the number of notes represented by the corresponding first feature. In terms of musical expression, it is consistent with the chord notation represented by the corresponding first feature. Furthermore, the third feature has a correlation with the first historical feature in terms of musical expression.
[0066] In step S403, the third features corresponding to the multiple first features are determined as a third feature sequence, and the multiple third features in the third feature sequence are arranged in chronological order.
[0067] In step S404, the first musical score sequence is obtained based on the third feature sequence and the second feature.
[0068] In this embodiment of the disclosure, after obtaining the first feature sequence, each first feature contained in the first feature sequence is processed in chronological order to obtain a third feature corresponding to each first feature. Specifically, for the first first feature in the first feature sequence with the earliest chronological order, the corresponding third feature is obtained directly based on the number of notes and chord markings represented by the first feature. The third feature represents multiple notes and chord markings, and the number of notes represented by the third feature is consistent with the number of notes represented by the first feature. The multiple notes represented by the third feature are musically consistent with the corresponding chord markings. For the first feature in the first feature sequence that is chronologically after the first first feature, the corresponding third feature is obtained based on the number of notes and chord markings represented by the first feature, combined with the multiple notes represented by the previously generated third feature (first historical feature) corresponding to the previous first feature. The third feature represents multiple notes and chord markings, and the number of notes represented by the third feature is consistent with the number of notes represented by the first feature. The multiple notes represented by the third feature are musically consistent with the corresponding chord markings, and the multiple notes represented by the third feature are musically related to the multiple notes represented by the corresponding historical feature.
[0069] In this embodiment of the disclosure, after obtaining the third feature corresponding to each first feature, the third features corresponding to the multiple first features are determined as a third feature sequence, and then the first musical score sequence is obtained based on the third feature sequence and the second feature.
[0070] In an exemplary embodiment of this disclosure, accepting Figure 3 A corresponding exemplary embodiment can be obtained by using the following method: A second model is used to further generate feature information corresponding to each first feature. The second sub-model performs autoregressive processing on the first feature sequence. During the autoregressive process, the second sub-model uses the first feature sequence as a key vector (K) and a numerical vector (V) when calculating cross-attention, and uses this as a representation of the musical structure. The second sub-model can be an autoregressive model, and a TransformerDecoder can be used as its decoder. Then, the third sub-model processes the key vector (K) and numerical vector (V) obtained from the first feature sequence to obtain the third feature sequence. During the autoregressive process, the third sub-model receives the current output of the second sub-model (i.e., the key vector (K) and numerical vector (V) corresponding to a single first feature) at each step, and generates the content of a complete patch through autoregression, i.e., [a chord covering the notes and other information within the musical time]. This complete patch content is re-encoded into a patch feature vector (a single third feature) and used as the input of the second sub-model at the next time step (i.e., as a historical feature). Until the second model outputs the identifier representing the completion of the characterization process.<eos>This means predicting all the notes under all chords in the entire piece to obtain the complete third feature sequence. The third model is an autoregressive model, using a TransformerDecoder as the decoder for the third sub-model.
[0071] In an exemplary embodiment of this disclosure, the training process of the model is described as follows: 1. Constructing a prompt structure based on abcnotation to define the input and output format for the dataset. For example, " <input> {Chord + Duration (represented by rests)} <output>"{Chord + Melody Notes}", where the <> symbols contain special markers used for information, and the {} symbols contain strings representing the original input's overall chord distribution and the melody generated by the model for this overall structure, respectively. When constructing the data, data samples in abcnotation format including chord markers are filtered out. A pure chord version without melody notes is obtained by removing the melody and other musical markers and filling the original melody's time value with rests. The first sub-model, second sub-model, third sub-model, and encoder are constructed separately. The model first adds the abcnotation prefix (e.g., M:4 / 4, representing...) to the input string. The musical score (in 4 / 4 time) is decomposed into a set of patches in K:V format. Each chord notation and its corresponding duration (rests and their durations) is then further segmented into patches, resulting in a sequence of patches. Each patch represents either a prefix (a complete song attribute) or a chord and its corresponding musical content. An encoder (such as BERT) is used to encode each patch, yielding a feature vector. This process is then used to build and train the model. It's understandable that the loss function for training a Transformer-based model can be set accordingly, similar to that used for text generation tasks.
[0072] The following describes a method for obtaining a first musical score sequence based on a third feature sequence and a second feature, according to embodiments of this disclosure.
[0073] Figure 5 This is a flowchart illustrating a method for obtaining a first musical score sequence based on a third feature sequence and a second feature, according to an exemplary embodiment. Figure 5 As shown, the method includes the following steps.
[0074] In step S501, for the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note.
[0075] In step S502, for the third features other than the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature and the second historical feature. The second historical feature is the fourth feature that has been generated corresponding to the previous third feature whose time sequence is adjacent to the corresponding third feature. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. The fourth feature is related to the second historical feature in terms of musical expression.
[0076] In step S503, the fourth features corresponding to the multiple third features are determined as a fourth feature sequence, and the multiple fourth features in the fourth feature sequence are arranged in chronological order.
[0077] In step S504, the first musical score sequence is obtained based on the fourth feature sequence and the second feature.
[0078] In this embodiment of the disclosure, after obtaining the third feature sequence, each third feature contained in the third feature sequence is processed in chronological order to obtain a fourth feature corresponding to each third feature. Specifically, for the first third feature in the sequence with the earliest chronological order, the corresponding fourth feature is obtained directly based on the multiple notes and chord markings represented by the third feature. The fourth feature represents multiple notes, chord markings, and additional musical markings, and the additional musical markings represented by the fourth feature are musically consistent with the corresponding multiple notes and chord markings. For third features in the sequence that are chronologically following the first third feature, the corresponding fourth feature is obtained based on the multiple notes and chord markings represented by the third feature, combined with the multiple notes, chord markings, and additional musical markings represented by the previously generated fourth feature (second historical feature) corresponding to the previous third feature. The fourth feature represents multiple notes, chord markings, and additional musical markings, and the additional musical markings represented by the third feature are musically consistent with the corresponding chord markings and the corresponding multiple notes. The additional musical markings represented by the third feature are musically related to the multiple notes, chord markings, and additional musical markings represented by the corresponding historical feature. This disclosure enables the generation of harmonic and musical notation based on the entire melody score, thus providing the melody score with more professional information.
[0079] In an exemplary embodiment of this disclosure, accepting Figure 4 A corresponding exemplary embodiment can obtain the fourth feature sequence as follows: By including a fourth sub-model, one bit of the output of the decoder part in the third sub-model is used as the first input, and autoregression is performed to obtain a short sequence, i.e., the specific music text within a patch. Each time, the text of a patch output by the second model is re-encoded, first used as the next input to the TransformerDecoder in the third sub-model (i.e., as the second historical feature), to obtain the feature of a patch. This feature is then input to the second model, outputting the patch text at the corresponding position. An identifier representing the completion of processing is then provided. <eos>This involves predicting the additional musical markers corresponding to each feature sequence in the third feature sequence to obtain the complete fourth feature sequence. Based on the fourth feature sequence and the second feature, the first musical score sequence is obtained, which is a musical score with markers (chord symbols and musical symbols). The first musical score sequence can be converted into music data (target audio data) in the original Musical Instrument Digital Interface (MIDI) format.
[0080] In an exemplary embodiment of this disclosure, model training can be performed by constructing an input-output format for the dataset. For example, " <input> {Melody + Chord Markings} <output>The structure is defined as "{melody + chord notation + musical notation}". Here, the <> symbols contain special markers representing task names and other information, while the {} symbols contain strings representing the original input melody score and the enhanced melody score (including chord and musical notation). During data construction, data samples in abcnotation format including chord notation are selected, and a melody version with chord notation is obtained by removing the musical notation. A fourth sub-model and its corresponding encoder are then constructed. The model first decomposes the abcnotation prefix (e.g., M:4 / 4, indicating a 4 / 4 time signature) into a set of patches in K:V format. Then, the note parts are divided into patches by measure, resulting in a patch sequence. Each patch represents a prefix (a complete song attribute) or a measure of musical content. An encoder (e.g., BERT) encodes each patch to obtain its feature vector. The fourth sub-model is trained using the vectors obtained from the encoder as input to two Transformer models for autoregressive model training. It is understood that the loss function can be set appropriately for the text generation task when training the Transformer model.
[0081] In this embodiment, after obtaining the music structure information, a first musical score sequence is obtained based on the number of musical measures, the number of notes in each musical measure, and the chord markings corresponding to each musical measure. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord markings corresponding to each musical measure. The number of musical measures is consistent with the number of musical measures, the number of notes in each musical measure is consistent with the number of notes, and the multiple notes in each musical measure musically correspond to the chords corresponding to the corresponding musical measure. The first musical score sequence is converted into target audio data, which can be output as target music. Through this disclosure, a first musical score sequence corresponding to the music structure information is generated based on the music structure information. Target music data that can be output as target music is obtained based on the first musical score sequence, making the music generation effect more controllable, thereby obtaining target music that matches the user's needs in terms of music structure. In summary, this disclosure presents a model structure that solves the problem of generating chord and musical notation for melodic scores. It can generate harmony and musical notation based on the entire melody score, giving the score more professional information. Furthermore, this disclosure designs a multi-level model to perform hierarchical feature extraction of abbreviation score information, solving the problem that existing models do not generate melodies containing chord and musical notation information. This achieves the task of generating harmony and musical notation for existing complete melody scores, enhancing the score information. Moreover, given musical attributes and structure, this disclosure generates corresponding melodies with structures corresponding to the given structures, making the music generation effect more controllable. Compared to existing solutions, this disclosure is more suitable for generating musical melodies through explicit input information, conforming to the creative logic of composers or enthusiasts. This disclosure combines the composition process with text and a large model through explicit representation, enabling the generation of large model text input directly through the UI of applications or web pages, conforming to usage logic.
[0082] Based on the same concept, this disclosure also provides a music generation device 100.
[0083] It is understood that the music generation apparatus 100 provided in this disclosure includes hardware structures and / or software modules corresponding to each function in order to achieve the above-mentioned functions. In conjunction with the units and algorithm steps of the various examples disclosed in this disclosure, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the technical solutions of this disclosure.
[0084] Figure 6 This is a block diagram illustrating a music generation apparatus 100 according to an exemplary embodiment. (Refer to...) Figure 6 The device includes an acquisition unit 101, a processing unit 102, and a conversion unit 103.
[0085] The acquisition unit 101 is used to acquire music structure information, which includes the number of music measures, the number of notes in each music measure, and the chord mark corresponding to each music measure. The chord mark represents the chord corresponding to the music measure.
[0086] The processing unit 102 is used to obtain a first musical score sequence based on the musical structure information. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord markings corresponding to each musical measure. The number of multiple musical measures is consistent with the number of musical measures. The number of multiple notes corresponding to each musical measure is consistent with the number of notes. The multiple notes corresponding to each musical measure are consistent with the chords corresponding to the corresponding musical measure in musical expression.
[0087] The conversion unit 103 is used to convert the first musical score sequence into target audio data, which can be output as target music.
[0088] In one embodiment, the acquisition unit 101 acquires music structure information in the following manner: It acquires input text used to generate music, the input text including unit beats, measure values, number of measures, and chords corresponding to each measure, all describing the music structure. The input text is then converted according to a preset musical notation to obtain the music structure information, where the number of notes in each measure corresponds to the unit beat and measure value.
[0089] In one embodiment, the music structure information further includes rhythmic pattern information and unit rhythm information, where the rhythmic pattern information and unit rhythm information correspond to unit musical beats and musical measure values. The processing unit 102 obtains a first musical score sequence based on the music structure information as follows: It performs feature encoding on the number of notes in each musical measure and the corresponding chord markings within multiple musical measures to obtain a first feature sequence. The first feature sequence includes multiple first features arranged in chronological order, each first feature representing the chord markings and number of notes corresponding to the corresponding musical measure. It then performs feature encoding on the rhythmic pattern information and unit rhythm information to obtain second features, where the second features represent the music's rhythmic pattern and unit rhythm. Based on the first feature sequence and the second features, it obtains a first musical score sequence, where the representation of notes in the first musical score sequence corresponds to the rhythmic pattern and unit rhythm.
[0090] In one embodiment, the processing unit 102 obtains a first musical score sequence based on a first feature sequence and a second feature as follows: For the earliest first feature in the first feature sequence, a third feature is generated based on the first feature. The third feature represents multiple notes and their corresponding chord markings. The number of notes matches the number of notes represented by the corresponding first feature, and the musical expression matches the chord markings represented by the corresponding first feature. For first features other than the earliest first feature in the first feature sequence, a third feature is generated based on the first feature and a first historical feature. The first historical feature is the previously generated third feature corresponding to the previous first feature whose time is adjacent to the corresponding first feature. The third feature represents multiple notes. The number of notes matches the number of notes represented by the corresponding first feature, and the musical expression matches the chord markings represented by the corresponding first feature. The third feature is also associated with the first historical feature in terms of musical expression. The third features corresponding to the multiple first features are determined as a third feature sequence, and the multiple third features in the third feature sequence are arranged in chronological order. The first musical score sequence is obtained based on the third feature sequence and the second feature.
[0091] In one embodiment, the processing unit 102 obtains a first musical score sequence based on a third feature sequence and a second feature in the following manner: For the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. For third features other than the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature and a second historical feature. The second historical feature is the previously generated fourth feature corresponding to the previous third feature whose time is adjacent to the corresponding third feature. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. Furthermore, the fourth feature is associated with the second historical feature in terms of musical expression. The fourth features corresponding to the multiple third features are defined as a fourth feature sequence, and the multiple fourth features in the fourth feature sequence are arranged in chronological order. Based on the fourth feature sequence and the second feature, the first musical score sequence is obtained.
[0092] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0093] Figure 7 This is a block diagram illustrating an apparatus 200 for music generation according to an exemplary embodiment. The apparatus 200 can be provided as a terminal. For example, the apparatus 200 can be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.
[0094] Reference Figure 7 The device 200 may include one or more of the following components: processing component 202, memory 204, power component 206, multimedia component 208, audio component 210, input / output (I / O) interface 212, sensor component 214, and communication component 216.
[0095] Processing component 202 typically controls the overall operation of device 200, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 202 may include one or more processors 220 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 202 may include one or more modules to facilitate interaction between processing component 202 and other components. For example, processing component 202 may include a multimedia module to facilitate interaction between multimedia component 208 and processing component 202.
[0096] Memory 204 is configured to store various types of data to support the operation of device 200. Examples of such data include instructions for any application or method operating on device 200, contact data, phonebook data, messages, pictures, videos, etc. Memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0097] The power supply component 206 provides power to the various components of the device 200. The power supply component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 200.
[0098] Multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 208 includes a front-facing camera and / or a rear-facing camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0099] Audio component 210 is configured to output and / or input audio signals. For example, audio component 210 includes a microphone (MIC) configured to receive external audio signals when device 200 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 204 or transmitted via communication component 216. In some embodiments, audio component 210 also includes a speaker for outputting audio signals.
[0100] I / O interface 212 provides an interface between processing component 202 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0101] Sensor assembly 214 includes one or more sensors for providing status assessments of various aspects of device 200. For example, sensor assembly 214 may detect the on / off state of device 200, the relative positioning of components such as the display and keypad of device 200, changes in the position of device 200 or a component of device 200, the presence or absence of user contact with device 200, the orientation or acceleration / deceleration of device 200, and temperature changes of device 200. Sensor assembly 214 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 214 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 214 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0102] Communication component 216 is configured to facilitate wired or wireless communication between device 200 and other devices. Device 200 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 216 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 216 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0103] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0104] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, which can be executed by a processor 220 of the device 200 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0105] Figure 8 This is a block diagram illustrating an apparatus 300 for music generation according to an exemplary embodiment. For example, apparatus 300 may be provided as a server. (Refer to...) Figure 8 The device 300 includes a processing component 322, which further includes one or more processors, and memory resources represented by memory 332 for storing instructions, such as application programs, that can be executed by the processing component 322. The application programs stored in memory 332 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 322 is configured to execute instructions to perform the aforementioned music generation method.
[0106] Device 300 may also include a power supply component 326 configured to perform power management of device 300, a wired or wireless network interface 350 configured to connect device 300 to a network, and an input / output (I / O) interface 358. Device 300 may operate on an operating system stored in memory 332, such as Windows Server™, MacOSX™, Unix™, Linux™, FreeBSD™, or similar.
[0107] It is understood that in this disclosure, "multiple" refers to two or more, and other quantifiers are similar. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The singular forms "a," "the," and "the" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0108] It is further understood that the terms "first," "second," etc., are used to describe various types of information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another, and do not indicate a specific order or degree of importance. In fact, the expressions "first," "second," etc., are completely interchangeable. For example, without departing from the scope of this disclosure, first information can also be referred to as second information, and similarly, second information can also be referred to as first information.
[0109] It is further understood that the terms "center", "longitudinal", "lateral", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this embodiment and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation.
[0110] It can be further understood that, unless otherwise specified, "connection" includes both direct connections where no other components exist between the two parties and indirect connections where other elements exist between them.
[0111] It is further understood that although operations are described in a specific order in the accompanying drawings in this disclosure, this should not be construed as requiring these operations to be performed in the specific order or serial order shown, or requiring all of the shown operations to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.
[0112] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0113] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.< / output> < / eos> < / output> < / eos>
Claims
1. A method for generating music, characterized in that, include: Obtain music structure information, which includes the number of music measures, the number of notes in each music measure, and the chord markings corresponding to each music measure. The chord markings represent the chords corresponding to the music measures. Based on the musical structure information, a first musical score sequence is obtained. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord markings corresponding to each musical measure. The number of multiple musical measures is consistent with the number of musical measures. The number of multiple notes corresponding to each musical measure is consistent with the number of notes. The multiple notes corresponding to each musical measure are consistent with the chords corresponding to the corresponding musical measures in musical expression. The first musical score sequence is converted into target audio data, which can be output as target music.
2. The method according to claim 1, characterized in that, The acquisition of music structure information includes: Obtain the input text used to generate music, the input text including the unit musical beat, musical measure time value, musical measure number, and chord corresponding to each musical measure to describe the musical structure; The input text is converted according to a preset musical notation to obtain the musical structure information. The number of notes included in each musical measure in the musical structure information corresponds to the unit musical beat and the musical measure time value.
3. The method according to claim 2, characterized in that, The music structure information also includes music rhythm pattern information and unit rhythm information, wherein the rhythm pattern information and the unit rhythm information correspond to the unit music beat and the music measure value; The step of obtaining the first musical score sequence based on the musical structure information includes: The number of notes in each musical measure and the chord markings corresponding to the musical measure are feature-encoded to obtain a first feature sequence. The first feature sequence includes multiple first features arranged in chronological order, and each first feature represents the chord markings and the number of notes corresponding to the musical measure. The rhythm pattern information and the unit rhythm information are feature-encoded to obtain a second feature, which represents the rhythm pattern and unit rhythm of the music; Based on the first feature sequence and the second feature, the first musical score sequence is obtained, wherein the representation of the notes in the first musical score sequence corresponds to the rhythmic pattern and the unit rhythm.
4. The method according to claim 3, characterized in that, The step of obtaining the first musical score sequence based on the first feature sequence and the second feature includes: For the earliest time-series first feature in the first feature sequence, a third feature is generated based on the first feature. The third feature represents multiple notes and the chord markings corresponding to the multiple notes. The number of multiple notes is consistent with the number of notes represented by the corresponding first feature, and the musical expression is consistent with the chord markings represented by the corresponding first feature. For the first features other than the earliest first feature in the first feature sequence, a third feature is generated based on the first feature and the first historical feature. The first historical feature is the generated third feature corresponding to the previous first feature whose time sequence is adjacent to the corresponding first feature. The third feature represents multiple notes, and the number of notes is consistent with the number of notes represented by the corresponding first feature. In terms of musical expression, it is consistent with the chord notation represented by the corresponding first feature. Furthermore, the third feature has a correlation with the first historical feature in terms of musical expression. The third features corresponding to the multiple first features are determined as a third feature sequence, and the multiple third features in the third feature sequence are arranged in chronological order; The first musical score sequence is obtained based on the third feature sequence and the second feature.
5. The method according to claim 4, characterized in that, The step of obtaining the first musical score sequence based on the third feature sequence and the second feature includes: For the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. For the third feature other than the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature and the second historical feature. The second historical feature is the fourth feature that has been generated corresponding to the previous third feature whose time sequence is adjacent to the corresponding third feature. The fourth feature represents multiple notes, chord markings corresponding to multiple notes, and additional musical markings corresponding to multiple notes. The additional musical markings corresponding to multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. The fourth feature is related to the second historical feature in terms of musical expression. The fourth features corresponding to the multiple third features are determined as a fourth feature sequence, and the multiple fourth features in the fourth feature sequence are arranged in chronological order; The first musical score sequence is obtained based on the fourth feature sequence and the second feature.
6. A music generation device, characterized in that, include: The acquisition unit is used to acquire music structure information, which includes the number of music measures, the number of notes in each music measure, and the chord markings corresponding to each music measure. The chord markings represent the chords corresponding to the music measures. The processing unit is configured to obtain a first musical score sequence based on the musical structure information. The first musical score sequence includes multiple musical measures, multiple notes in each musical measure, and chord markings corresponding to each musical measure. The number of multiple musical measures is consistent with the number of musical measures. The number of multiple notes corresponding to each musical measure is consistent with the number of notes. The multiple notes corresponding to each musical measure are consistent with the chords corresponding to the corresponding musical measure in musical expression. A conversion unit is used to convert the first musical score sequence into target audio data, which can be output as target music.
7. The apparatus according to claim 6, characterized in that, The acquisition unit acquires music structure information in the following manner: Obtain the input text used to generate music, the input text including the unit musical beat, musical measure time value, musical measure number, and chord corresponding to each musical measure to describe the musical structure; The input text is converted according to a preset musical notation to obtain the musical structure information. The number of notes included in each musical measure in the musical structure information corresponds to the unit musical beat and the musical measure time value.
8. The apparatus according to claim 7, characterized in that, The music structure information also includes music rhythm pattern information and unit rhythm information, wherein the rhythm pattern information and the unit rhythm information correspond to the unit music beat and the music measure value; The processing unit obtains the first musical score sequence based on the music structure information in the following manner: The number of notes in each musical measure and the chord markings corresponding to the musical measure are feature-encoded to obtain a first feature sequence. The first feature sequence includes multiple first features arranged in chronological order, and each first feature represents the chord markings and the number of notes corresponding to the musical measure. The rhythm pattern information and the unit rhythm information are feature-encoded to obtain a second feature, which represents the rhythm pattern and unit rhythm of the music; Based on the first feature sequence and the second feature, the first musical score sequence is obtained, wherein the representation of the notes in the first musical score sequence corresponds to the rhythmic pattern and the unit rhythm.
9. The apparatus according to claim 8, characterized in that, The processing unit obtains the first musical score sequence based on the first feature sequence and the second feature in the following manner: For the earliest time-series first feature in the first feature sequence, a third feature is generated based on the first feature. The third feature represents multiple notes and the chord markings corresponding to the multiple notes. The number of multiple notes is consistent with the number of notes represented by the corresponding first feature, and the musical expression is consistent with the chord markings represented by the corresponding first feature. For the first features other than the earliest first feature in the first feature sequence, a third feature is generated based on the first feature and the first historical feature. The first historical feature is the generated third feature corresponding to the previous first feature whose time sequence is adjacent to the corresponding first feature. The third feature represents multiple notes, and the number of notes is consistent with the number of notes represented by the corresponding first feature. In terms of musical expression, it is consistent with the chord notation represented by the corresponding first feature. Furthermore, the third feature has a correlation with the first historical feature in terms of musical expression. The third features corresponding to the multiple first features are determined as a third feature sequence, and the multiple third features in the third feature sequence are arranged in chronological order; The first musical score sequence is obtained based on the third feature sequence and the second feature.
10. The apparatus according to claim 9, characterized in that, The processing unit obtains the first musical score sequence based on the third feature sequence and the second feature in the following manner: For the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature. The fourth feature represents multiple notes, chord markings corresponding to the multiple notes, and additional musical markings corresponding to the multiple notes. The additional musical markings corresponding to the multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to the multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. For the third feature other than the earliest third feature in the third feature sequence, a fourth feature is generated based on the third feature and the second historical feature. The second historical feature is the fourth feature that has been generated corresponding to the previous third feature whose time sequence is adjacent to the corresponding third feature. The fourth feature represents multiple notes, chord markings corresponding to multiple notes, and additional musical markings corresponding to multiple notes. The additional musical markings corresponding to multiple notes are consistent with the chord markings represented by the corresponding third feature in terms of musical expression. The additional musical markings corresponding to multiple notes are used to adjust the pitch of a single note and / or add annotations to a single note. The fourth feature is related to the second historical feature in terms of musical expression. The fourth features corresponding to the multiple third features are determined as a fourth feature sequence, and the multiple fourth features in the fourth feature sequence are arranged in chronological order; The first musical score sequence is obtained based on the fourth feature sequence and the second feature.
11. An electronic device, characterized in that, include: processor: Memory used to store processor-executable instructions; The processor is configured to execute the music generation method according to any one of claims 1 to 5.
12. A storage medium, characterized in that, The storage medium stores instructions that, when executed by a processor, enable the processor to perform the music generation method according to any one of claims 1 to 5.