Method and system for generating music score through musical instrument playing audio and readable storage medium

By converting instrument performance audio into MIDI data and using the beat tracking and score generation model of the GPT2 model, the problem of converting instrument performance audio into high-quality scores in the existing technology is solved, and efficient and accurate score generation is achieved, adapting to a variety of inputs and instrument types, improving user experience and creative efficiency.

CN120356446APending Publication Date: 2025-07-22BEIJING ZHIQU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510630086.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

It is difficult for the prior art to efficiently convert the audio of the instrument to high-quality standard music score, especially the lack of advanced music performance information such as musical methods, fingering methods, expression marks, and the effect is not good when processing non-standard input.

Method used

The instrument performance audio is converted into MIDI data, and the beat tracking and score generation model of the GPT2 model structure are processed separately to generate the simplest score and complete score with bars. Combined with the beam search strategy, it avoids repeated generation and supports multiple input forms and instrument types.

Benefits of technology

The generated scores are improved in accuracy and completeness, including advanced music performance information, adapting to different music styles and input forms, lowering the threshold for creation and learning, and improving processing efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356446A_ABST
    Figure CN120356446A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for generating a music score through musical instrument playing audio and a readable storage medium, and the method comprises the following steps: (1) converting the musical instrument playing audio into MIDI data which comprises pitch, time value and strength information; (2) converting the MIDI data into a first token sequence; (3) inputting the first token sequence into a beat tracking model to generate a second token sequence of the simplest music score with a nodal line; (4) inputting the second token sequence into a music score generation model to generate a third token sequence of a complete music score; and (5) converting the third token sequence into a music score file in a MusicXML format. According to the method, the complicated playing audio-to-music score task is decomposed into two sub-tasks of beat tracking and music score generation, and the special models are trained respectively, so that the model learning difficulty is reduced, and the generation quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and music information processing, and specifically refers to a method, system and readable storage medium for generating musical scores from instrument performance audio. Background Art

[0002] Generating structured musical scores from human performance recordings is a challenging task. Human performances are usually represented in the form of performance MIDI, which can be easily recorded by MIDI instruments or generated by audio transcription. However, high-quality standard score formats such as MusicXML usually need to be created by human experts. The process of converting performance MIDI to score (PM2S) is complex and includes many tasks such as rhythm quantization, beat tracking, voice allocation, and typesetting details. Therefore, PM2S and its subtasks have always been active research topics. Early methods relied on classical modeling and manual feature processing, and then gradually turned to statistical methods based on hidden Markov models (HMMs) combined with heuristic rules. Recent research has tried to use deep learning to solve the PM2S problem, such as using large models for end-to-end learning in the Seq2Seq manner, but the actual effects are poor and many score elements are missing, such as playing techniques, fingerings, expression marks, etc. Currently, there are also some tools that support converting MIDI to MusicXML, such as MuseScore, etc., but they are only limited to regularizing MIDI, and the effect of converting performance MIDI to MusicXML is also poor.

[0003] To solve the above technical problems, a method, system and readable storage medium for generating musical scores from instrument performance audio are urgently needed to be disclosed. Summary of the Invention

[0004] To solve the above technical problems, the technical solutions provided by the present invention are as follows: A method, system and readable storage medium for generating musical scores from instrument performance audio, including the following steps:

[0005] (1) Convert the instrument performance audio into MIDI data, where the MIDI data includes pitch, duration, and dynamics information;

[0006] (2) Convert the MIDI data into a first token sequence;

[0007] (3) Input the first token sequence into a beat tracking model to generate a second token sequence of the simplest score with bar lines;

[0008] (4) Input the second token sequence into a score generation model to generate a third token sequence of a complete score;

[0009] (5) Convert the third token sequence into a score file in MusicXML format.

[0010] Furthermore, both the beat tracking model and the music score generation model are implemented based on the GPT2 model structure and are trained through two stages: pre-training and fine-tuning.

[0011] Furthermore, in the pre-training stage of the beat tracking model, self-supervised learning is adopted, and the training data is the token sequence of the simplest music score with bar lines; in the fine-tuning stage, supervised learning is adopted, the input is the token sequence of the performance MIDI, and the output is the token sequence of the simplest music score with bar lines.

[0012] Furthermore, in the pre-training stage of the music score generation model, self-supervised learning is adopted, and the training data is the token sequence of the complete music score; in the fine-tuning stage, supervised learning is adopted, the input is the token sequence of the simplest music score with bar lines, and the output is the token sequence of the complete music score.

[0013] Furthermore, the beat tracking model and the music score generation model adopt the beam search strategy during inference and generation, and avoid the problem of repeated generation by optimizing the no_repeat_ngram_size parameter.

[0014] The advantages of the invention compared with the prior art are as follows:

[0015] 1. The innovative implementation method for generating music scores from performance MIDI splits complex tasks into two relatively simple tasks. Each simple task is implemented using a large GPT model, reducing the learning difficulty of the large model. The two simple tasks are then executed in series, ultimately achieving good results.

[0016] 2. Improve the accuracy and integrity of music score generation. By decomposing the complex task of converting performance audio to music score into two subtasks: beat tracking and music score generation, and training dedicated models respectively, the learning difficulty of the model is reduced and the generation quality is improved. The generated music score not only contains basic elements such as pitch, duration, and bar lines, but also can accurately restore advanced music performance information such as articulation (e.g., legato, staccato), fingering, expression marks (e.g., forte-piano, accelerando-rallentando), and ornaments (e.g., trill, glissando), making the generated music score more professional and practical.

[0017] 3. Support multiple input forms and have wide applicability. It not only supports regular MIDI input, but also can process non-standard inputs such as improvised performance recordings, video background music, and live performance segments, and adapts to different music styles (classical, pop, jazz, etc.). Users can directly record or upload audio through mobile devices such as mobile phones and tablets, and the system automatically processes and generates music scores, greatly simplifying the process of music creation and learning.

[0018] 4. Efficient processing and real-time feedback. The GPT2 large model + fine-tuning optimization method is used, combined with the Beam Search generation strategy to avoid repeated generation problems and improve reasoning efficiency. The system provides progress bar feedback during the processing process, so that users can understand the processing status in real time and improve the user experience.

[0019] 5. Lower the threshold for music creation and learning. Music creators can quickly convert inspirational performance recordings into standard music scores, avoiding the tedious manual notation and improving creative efficiency.

[0020] Music lovers: When you hear music you like but can't find the music score, you can generate it with one click, without the need for professional music transcription skills.

[0021] Music educators: Demonstration performances can be quickly converted into teaching scores, saving preparation time and improving teaching efficiency.

[0022] 6. The technology is advanced and scalable. It adopts the GPT2 large model + self-supervised pre-training + supervised fine-tuning method. The model has strong generalization ability and can adapt to the performance audio of different instruments (such as piano, guitar, violin, etc.). BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Schematic diagram of the user interaction process for generating music scores for musical instrument performance audio;

[0024] Figure 2 Another schematic diagram of a user interaction process for generating music scores for musical instrument performance audio;

[0025] Figure 3 The overall technical flow chart for generating music scores for instrument performance audio. DETAILED DESCRIPTION

[0026] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0027] The present invention is described in detail with reference to the accompanying drawings.

[0028] The present invention provides a method for generating music scores by playing audio of musical instruments in a specific implementation, comprising the following steps:

[0029] (1) converting the musical instrument performance audio into MIDI data, wherein the MIDI data includes pitch, duration, and force information;

[0030] (2) converting the MIDI data into a first token sequence;

[0031] (3) inputting the first token sequence into a beat tracking model to generate a second token sequence of a simplest musical score with bar lines;

[0032] (4) Input the second token sequence into the sheet music generation model to generate the third token sequence of the complete sheet music;

[0033] (5) Convert the third token sequence into a sheet music file in MusicXML format.

[0034] Example 1: Generating sheet music from a piano improvisation recording:

[0035] Application scenario:

[0036] A composer improvises a melody in front of a piano and hopes to quickly record the sheet music for subsequent arrangement.

[0037] Implementation process:

[0038] Audio input: The user records a 90 - second audio of piano improvisation (sampling rate 44.1kHz, mono) using a mobile phone. Audio processing:

[0039] The system calls the audio - to - MIDI model of ONNXruntime, and the processing time is 3.2 seconds;

[0040] Generate a MIDI file containing 328 note events, with an average transcription accuracy of 92%.

[0041] Beat tracking:

[0042] Convert MIDI to a token sequence (a total of 842 tokens);

[0043] The beat tracking model (6 - layer GPT2, hidden_size = 768) has a processing time of 1.8 seconds;

[0044] Output an intermediate representation with bar lines, automatically recognized as 4 / 4 time signature, tempo = 108 bpm.

[0045] Sheet music generation:

[0046] The sheet music generation model (8 - layer GPT2, hidden_size = 1024) has a processing time of 2.4 seconds;

[0047] Generate a complete sheet music containing the following elements:

[0048] Staff: Treble clef;

[0049] Key signature: C major;

[0050] Expression marks: p, mf, cresc., etc.;

[0051] Playing techniques: Slur, staccato marks;

[0052] Fingerings: Hints for some key fingerings.

[0053] Output result:

[0054] Generate a MusicXML file (size 28KB);

[0055] When opened in MuseScore, the display effect is good, and the matching degree with the original performance reaches 89%.

[0056] Example 2: Transcribing Guitar Singing Recordings into Sheet Music

[0057] Application scenario:

[0058] Music lovers hope to convert the accompaniment part in their favorite guitar singing videos into guitar sheet music.

[0059] Implementation process:

[0060] Audio input: Extract 2 minutes of audio (including vocals and guitar) from the video.

[0061] Sound source separation:

[0062] Use a pre-trained sound source separation model to extract the guitar track.

[0063] The signal-to-noise ratio is increased by 12dB, and the vocal suppression effect is good.

[0064] Feature extraction:

[0065] Identify standard tuning (EADGBE);

[0066] Detect that it is mainly in the C major chord progression.

[0067] Sheet music generation:

[0068] Generate a guitar sheet music with chord markings;

[0069] Automatically add right-hand picking fingering tips;

[0070] Identify special sliding and hammer-on techniques.

[0071] Output optimization:

[0072] Provide an editable output in Guitar Pro format;

[0073] Support export to PDF and image formats.

[0074] As a further elaboration of the present invention, both the beat tracking model and the sheet music generation model are implemented based on the GPT2 model structure and are trained through two stages of pre-training and fine-tuning.

[0075] As a further elaboration of the present invention, in the pre-training stage of the beat tracking model, self-supervised learning is adopted, and the training data is the token sequence of the simplest musical score with bar lines; in the fine-tuning stage, supervised learning is adopted, the input is the token sequence of the performed MIDI, and the output is the token sequence of the simplest musical score with bar lines.

[0076] As a further elaboration of the present invention, in the pre-training stage of the musical score generation model, self-supervised learning is adopted, and the training data is the token sequence of the complete musical score; in the fine-tuning stage, supervised learning is adopted, the input is the token sequence of the simplest musical score with bar lines, and the output is the token sequence of the complete musical score.

[0077] As a further elaboration of the present invention, when the beat tracking model and the musical score generation model perform inference and generation, the beam search strategy is adopted, and the problem of repeated generation is avoided by optimizing the no_repeat_ngram_size parameter.

[0078] As a further elaboration of the present invention, the token sequence of the complete musical score includes at least one element among duration, tempo, time signature, key signature, clef, slurs, repeated notes, and ornaments.

[0079] As a further elaboration of the present invention, the musical score file in MusicXML format includes at least one element among bar lines, pitch, duration, time signature, key signature, clef, accidentals, ornaments, articulation, fingering, and expression marks.

[0080] A system for generating a musical score from an instrument performance audio, comprising:

[0081] An audio transcription module for converting instrument performance audio into MIDI data;

[0082] A first conversion module for converting the MIDI data into a first token sequence;

[0083] A beat tracking module for inputting the first token sequence into a beat tracking model to generate a second token sequence of the simplest musical score with bar lines;

[0084] A musical score generation module for inputting the second token sequence into a musical score generation model to generate a third token sequence of the complete musical score;

[0085] A second conversion module for converting the third token sequence into a musical score file in MusicXML format.

[0086] A system for generating a musical score from an instrument performance audio, wherein both the beat tracking module and the musical score generation module are implemented by using a pre-trained and fine-tuned GPT2 model.

[0087] A computer-readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, it implements a method for generating a musical score for musical instrument performance audio.

[0088] The above describes the present invention and its implementation manners. Such a description is not restrictive, and what is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the spirit of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. A method for generating a musical score from musical instrument performance audio, characterized in that, It includes the following steps: (1) Convert the musical instrument performance audio into MIDI data, where the MIDI data includes pitch, duration, and dynamics information; (2) Convert the MIDI data into a first token sequence; (3) Input the first token sequence into a beat tracking model to generate a second token sequence of the simplest musical score with bar lines; (4) Input the second token sequence into a musical score generation model to generate a third token sequence of the complete musical score; (5) Convert the third token sequence into a musical score file in MusicXML format.

2. The method for generating a musical score from musical instrument performance audio according to claim 1, wherein, Both the beat tracking model and the musical score generation model are implemented based on the GPT2 model structure and are trained through two stages: pre-training and fine-tuning.

3. A method for generating a musical score from the audio of an instrument performance according to claim 2, characterized in that, In the pre-training stage of the beat tracking model, self-supervised learning is adopted, and the training data is the token sequence of the simplest musical score with bar lines; in the fine-tuning stage, supervised learning is adopted, the input is the token sequence of the performance MIDI, and the output is the token sequence of the simplest musical score with bar lines.

4. A method for generating a musical score from musical instrument performance audio according to claim 2, characterized in that, In the pre-training stage of the musical score generation model, self-supervised learning is adopted, and the training data is the token sequence of the complete musical score; in the fine-tuning stage, supervised learning is adopted, the input is the token sequence of the simplest musical score with bar lines, and the output is the token sequence of the complete musical score.

5. A method for generating a musical score from musical instrument performance audio according to claim 1, characterized in that, Both the beat tracking model and the musical score generation model adopt the beam search strategy during inference generation and avoid the problem of repeated generation by optimizing the no_repeat_ngram_size parameter.

6. A method for generating a musical score from musical instrument performance audio according to claim 1, characterized in that, The token sequence of the complete musical score includes at least one element among duration, tempo, time signature, key signature, clef, slurs, repeated notes, and ornaments.

7. A method for generating a musical score from musical instrument performance audio according to claim 1, characterized in that, The musical score file in MusicXML format contains at least one element among bar lines, pitch, duration, time signature, key signature, clef, accidentals, ornaments, articulation, fingering, and expression marks.

8. A system for generating musical scores from musical instrument performance audio, characterized in that, It includes: An audio transcription module for converting the musical instrument performance audio into MIDI data; A first conversion module for converting the MIDI data into a first token sequence; A beat tracking module for inputting the first token sequence into a beat tracking model to generate a second token sequence of the simplest musical score with bar lines; A musical score generation module for inputting the second token sequence into a musical score generation model to generate a third token sequence of the complete musical score; A second conversion module for converting the third token sequence into a musical score file in MusicXML format.

9. The system for generating a musical score from the audio of an instrument performance according to claim 8, wherein Both the beat tracking module and the musical score generation module are implemented using the pre-trained and fine-tuned GPT2 model.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.