Method and system for converting image into music score
By converting note images into MIDI files and performing multi-voice allocation, standardized sheet music data is generated, solving the problem of poor consistency between the sight-reading results and the original score structure in existing technologies. This achieves the generation of sheet music data that is closer to the standard of human input, supporting AI accompaniment and vocal synthesis functions in music teaching systems.
Patent Information
- Application Number
- CN202511140536.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-12-02
AI Technical Summary
Existing methods for collecting sheet music data, such as manual input, scanning of printed scores, or using music recognition software, result in poor consistency between the recognized scores and the original scores, failing to meet the needs of AI accompaniment and vocal synthesis functions in music teaching systems.
By converting note images into MIDI files, and then converting the MIDI files into sheet music data, a multi-part allocation method is used to ensure that notes are correctly assigned to parts and to generate a standardized sheet music data format.
It improves the structural consistency between the sight-reading results and the original score, making the generated score visually and logically closer to the standard of human input, and can be used for the regular functions of the music teaching system.
Smart Images

Figure CN121053932A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of music teaching systems, specifically relating to a method and system for converting images into sheet music. Background Technology
[0002] To enhance AI accompaniment and vocal synthesis capabilities, music teaching systems require structured sheet music data for large-scale AI accompaniment or vocal synthesis models. Existing methods for acquiring sheet music data include manual input, scanning printed scores, or using music recognition software to automatically identify melodies in audio files. When using scanned printed scores or music recognition software (such as OMR, Optical Music Recognition) for score reading, the structure of the readout results is poorly consistent with the original score, causing the generated score to differ significantly from the manual input standard in both visual and logical aspects. Furthermore, the readout results cannot be used for the regular functions of the music teaching system. Summary of the Invention
[0003] To address the above technical problems, one embodiment of the present invention proposes a method for converting images to musical scores. First, a note image is identified to obtain a MIDI file. Then, the MIDI file is converted to musical score data, whereby the musical score data includes several vocal parts. The method is characterized by the following multi-voice part assignment steps:
[0004] The MIDI file is parsed to obtain several notes to be assigned, the notes to be assigned including start time and duration;
[0005] When the note to be assigned is the first note, the note to be assigned is assigned to the first voice part, the end time of the first voice part, the start time of the first voice part, and the duration of the first voice part are recorded and marked as the first voice part;
[0006] When the note to be assigned is not the first note, if the non-first note has the same start time as the first note but a different duration, the note to be assigned is assigned to the second voice part, the end time of the second voice part, the start time of the second voice part and the duration of the second voice part are recorded, and it is marked as the second voice part;
[0007] When the note to be assigned is not the first note, if the non-first note has the same start time and the same duration as the first note, then the note to be assigned is assigned to the first voice and marked as a first voice chord.
[0008] When the note to be assigned begins immediately after the end of the first voice part, it continues to be assigned to the first voice part;
[0009] When the note to be assigned begins immediately after the end of the second voice part, it continues to be assigned to the second voice part;
[0010] If the voice part to which the note to be assigned belongs cannot be determined, it will be assigned to the first voice part by default.
[0011] The present invention also proposes a system for converting images to musical scores, the system comprising at least one processor; and a memory storing instructions that, when executed by the at least one processor, implement the above-described method for converting images to musical scores.
[0012] Compared with the prior art, the present invention has the following advantages: An embodiment of the present invention provides an image-to-music score conversion method that can directly convert the identified note data into music score data in a standardized data format for use by music teaching systems. This effectively improves the structural consistency between the music score recognition results and the original score, making the generated score closer to the standard of manual input in terms of both visual and logical aspects. Moreover, the obtained music score data can also be used for the regular functions of music teaching systems. Attached Figure Description
[0013] Figure 1 : A schematic diagram of the image-to-musical-score conversion process of this invention;
[0014] Figure 2 : A schematic diagram of the music score conversion process of this invention;
[0015] Figure 3 : A schematic diagram of the music recognition interface of this invention. Clicking the plus sign in the diagram allows you to select a music file.
[0016] Figure 4 : A schematic diagram of the file selection interface of this invention;
[0017] Figure 5 : A schematic diagram of the image-to-musical-score conversion interface of this invention;
[0018] Figure 6 : A schematic diagram of the interface for the spectrum identification process of this invention;
[0019] Figure 7 : A schematic diagram of an image-to-musical-sheet system. Detailed Implementation
[0020] The following embodiments further illustrate the content of the present invention, but should not be construed as limiting the present invention. Any modifications or substitutions made to the methods, steps, or conditions of the present invention without departing from the spirit and essence of the invention are within the scope of the present invention.
[0021] Music teaching systems (including any music teaching system disclosed in the series of patents such as CN109377818A and CN109345905A of Beijing Jin Sanhui Technology Co., Ltd.) require structured musical score data (also referred to as structured musical score data, score data, or score) to add AI accompaniment and vocal synthesis functions for large AI accompaniment or vocal synthesis models. Existing methods for acquiring musical score data include manual input, scanning printed scores, or using music recognition software to automatically identify melodies in audio files. When using scanned printed scores or music recognition software (such as OMR, Optical Music Recognition) for score reading, the structure of the score reading results is poorly consistent with the original score, causing the generated score to differ significantly from the manual input standard in both visual and logical aspects.
[0022] The method of obtaining sheet music data by scanning printed scores in this invention requires the use of external image recognition or optical music recognition (OMR) tools (using third-party technology). First, the sheet music image files (also referred to as images in this invention) in image format (such as ".jpg", ".jpeg", ".png") or PDF format (such as ".pdf") are converted into MIDI files. Then, based on the note events recorded in the MIDI files, the sheet music data is extracted and converted into standardized data formats suitable for use by music teaching systems (such as modules including AI accompaniment and vocal synthesis). Figure 1 As shown, the process involves "image → MIDI file data → internal data structure," effectively improving the structural consistency between the music score and the original score, making the generated score visually and logically closer to the standards of human input. Furthermore, the music teaching system also includes standard functions such as score playback and editing; therefore, the data obtained from converting images to music scores is also used in these standard functions.
[0023] like Figure 3 , Figure 4 , Figure 5 , Figure 6 As shown, this is a music teaching system configured with the "intelligent music reading" function, which is implemented by the image-to-music-score conversion method and system embodiment of this invention. The music reading images in this invention include any of the formats jpg, jpeg, png, and pdf. It should be noted that this invention does not limit the form of the intelligent music reading front-end interface.
[0024] "Musical score data" includes global musical score data and / or a portion thereof, and the data is represented in (but not limited to) staff notation and simplified notation. The "musical score data" includes objects in XML or JSON format. Preferably, the data is saved in JSON format, and the JSON object includes a global musical score class. Each global musical score class contains an array of data from several complex staff classes, each complex staff class contains an array of data from several single staff classes, and each single staff class contains an array of data from several measures, wherein: some measures include several voices, and some measures include corresponding secondary melody data.
[0025] The image-to-music-score conversion system of this invention belongs to a subsystem of the music teaching system, such as... Figure 7 As shown, it includes a bus, processor, memory, network interface, and I / O interface, all present in general-purpose computer devices. The memory stores instructions, which, when executed by at least one processor, implement the image-to-music-score method of the present invention, thereby solving several technical problems, including the aforementioned technical issues.
[0026] The method of converting images to musical scores is illustrated below through several embodiments.
[0027] In some embodiments, the method for converting an image to sheet music is as follows: Figure 1 As shown, it includes the following steps:
[0028] S1: First, the musical note image is processed by musical note recognition to obtain a MIDI file;
[0029] S2: The MIDI file is then converted into sheet music data, which includes several vocal parts.
[0030] In these embodiments, the note recognition of the present invention can employ any known technology as long as it can obtain MIDI file data. For example, by configuring the interface of a third-party image segmentation platform in the system backend of the present invention, the frontend only needs to request the interface for the third-party image segmentation platform to perform image segmentation and thus obtain MIDI file data. In a specific embodiment, the conversion of image to MIDI file data is divided into two steps: the first step is image to .mxl, based on the open-source project audiveris (https: / / github.com / Audiveris / audiveris?tab=readme-ov-file_2); the second step is .mxl to MIDI4, using the Python third-party library music21.
[0031] The present invention obtains musical score data for the generation of AI vocals in a music teaching system. When generating AI accompaniment, the AI algorithm involved is also implemented through known third-party technologies, such as the "chord theory".
[0032] In some embodiments, the music score conversion is as follows: Figure 2 As shown, the multi-voice distribution steps include the following:
[0033] S21: Parse the MIDI file to obtain several notes to be assigned. The notes to be assigned include the start time and duration. The ticks ("tick" is the most basic unit for calculating time in a MIDI file) of the MIDI file record the start time and duration of the notes. The notes to be assigned include the start time and duration, which are obtained based on the ticks of the notes recorded in the MIDI file.
[0034] S22: When the note to be assigned is the first note, the note to be assigned is assigned to the first voice part, the end time of the first voice part, the start time of the first voice part, and the duration of the first voice part are recorded and marked as the first voice part;
[0035] S23: When the note to be assigned is not the first note, if the non-first note has the same start time as the first note but a different duration, the note to be assigned is assigned to the second voice part, the end time of the second voice part, the start time of the second voice part, and the duration of the second voice part are recorded and marked as the second voice part; as can be seen from S22 and S23, the start time of the first voice part increases sequentially from 0, and if the start time of a note deviates from the end time of the previous note, this note is identified as the second voice part.
[0036] S24: When the note to be assigned is not the first note, if the non-first note has the same start time and the same duration as the first note, then the note to be assigned is assigned to the first voice and marked as a first voice chord.
[0037] S25: When the note to be assigned begins immediately after the end time of the first voice part, it continues to be assigned to the first voice part;
[0038] S26: When the note to be assigned begins immediately after the end time of the second voice part, it continues to be assigned to the second voice part;
[0039] S27: If it is impossible to determine the voice part to which the note to be assigned belongs, it will be assigned to the first voice part by default.
[0040] In these embodiments, on the one hand, by using the process of "image → MIDI file data → internal data structure," the identified note data can be directly converted into sheet music data in a standardized data format for use by the music teaching system. This effectively improves the structural consistency between the music recognition results and the original score, making the generated score visually and logically closer to the standard of manual input. Moreover, the obtained sheet music data can also be used for the regular functions of the music teaching system. On the other hand, a reliable method for multi-voice allocation is provided to address the structural inconsistencies that often occur during multi-voice conversion in sheet music conversion.
[0041] In some embodiments, the musical score data includes several complex staves, each complex stave including several simple staves, each simple stave including several measures, and each measure including several voice parts; the MIDI file includes header data and several tracks, each track including several notes to be assigned. This invention does not specifically limit the method and order of parsing the MIDI file. For example, the MIDI file can be parsed to obtain several tracks, the tracks can be parsed to obtain several notes to be assigned, the number of tracks parsed determines the number of voice parts, and based on the parsing results, the track, note, tempo, time signature, duration, total number of measures, etc., can be obtained.
[0042] In some embodiments, the music score conversion includes the following steps for creating a complex musical score:
[0043] Parse the header data of the MIDI file;
[0044] Calculate the number of sections to be created;
[0045] Create the basic data structure for the musical score data;
[0046] Create the complex spectrum table based on the number of subsections in each row;
[0047] Create a duration accumulation for the current track in the complex staff for each individual staff; create a note cycle index for each individual staff in the complex staff;
[0048] Create the last row of processing logic in the complex spectrum table;
[0049] Set the position for the complex spectrum.
[0050] In some embodiments, the music score conversion includes the following single staff creation steps:
[0051] Create a beat rate object;
[0052] Create a spectral object;
[0053] Create a sequence number object;
[0054] Process single-spectral data in a loop;
[0055] Create a single spectrum.
[0056] In some embodiments, the music score conversion includes the following steps for creating measures: 1) Creating a measure base object, specifically including:
[0057] Clone the note cycle index of the single staff;
[0058] Get the total duration of the current audio track;
[0059] Set the index position of the subsection in the complex spectrum table;
[0060] Create a section object and add it to the section array of the single spectrometer.
[0061] 2) Set the horizontal layout of the sections, specifically including:
[0062] Set the X coordinate of the musical line;
[0063] Set the X-coordinate of the simplified musical notation;
[0064] Calculate the starting position of the next section, which is the right boundary of the current section;
[0065] 3) Set the time attributes for the section, specifically including: setting the section start time;
[0066] 6) Calculate the duration of each segment, specifically including:
[0067] Calculate the duration of the current measure based on the current time signature;
[0068] Get the start time of the current section, add up the duration of the current section, and prepare for the calculation of the next section.
[0069] In some embodiments, the section creation step further includes:
[0070] 4) Assign time signatures, specifically including: if there is a time signature change at the current time, record the time signature used by the current audio track;
[0071] 5) Allocate beat rate, specifically including: if the single spectral is the first single spectral of the first complex spectral, allocate beat rate.
[0072] In some embodiments, the music score conversion includes the following note creation steps:
[0073] Create a complete musical note object (MusicScoreYinFu);
[0074] Create a MusicScoreYinFuContent object;
[0075] Set various attributes for the musical note, including duration, pitch, and display time;
[0076] Distribute the notes to the correct voices according to the polyphonic parting steps described above;
[0077] Set the clef information for the notes.
[0078] In some embodiments, the music score conversion includes the following note assignment steps:
[0079] Iterate through all the notes in the track and extract the notes belonging to the current measure;
[0080] The track data is split into notes for each measure.
[0081] The note array for the current measure;
[0082] Voice part time tracking object, used for multi-voice judgment;
[0083] Iterate through all notes in the current audio track;
[0084] Note start time (MIDI tick);
[0085] Note duration (MIDI tick);
[0086] Determine whether a note belongs to the time range of the current measure.
[0087] The order in which the processor of this invention executes programs from memory can be determined according to the development environment or business needs, and this invention does not impose specific limitations. For example, the following processing order is feasible when the development environment is Node.js 12, Vue 2, and Webpack.
[0088] / / Header data processing;
[0089] / / Calculate the number of sections created (run a separate loop to calculate the number of sections);
[0090] / / Create the basic data structure for the music score data (musicScoreData);
[0091] / / Create a complex spectrum based on the number of measures in each line;
[0092] / / The duration accumulation object for each individual spectrogram;
[0093] / / Note cycle index object for each single staff;
[0094] / / The last line needs to be processed;
[0095] / / Determine the number of sections in each line;
[0096] / / Set the position of the complex spectrum;
[0097] / / Create a beat rate object;
[0098] / / Create a spectral object;
[0099] / / Create a sequence number object;
[0100] / / Process single-spectrum data in a loop;
[0101] / / Create a single spectrum;
[0102] / / Create a section;
[0103] / / Create sections in a loop;
[0104] / / Record the beat number of the last position;
[0105] / / Create a section and calculate the total time of the current section based on the current section's time signature;
[0106] / / Accumulate the duration of the current section;
[0107] / / Create a musical note;
[0108] / / Splitting the track data to extract data from each subsection;
[0109] / / If there is no Prev, it starts with the first note; the default is the first part.
[0110] In some embodiments, the image-to-musical-score method of the present invention further includes logic for processing missing notes, specifically including the following steps:
[0111] Break down notes with incorrect duration;
[0112] Correcting the issues with the vocal parts.
[0113] In these embodiments, when a note duration is incorrect, the actual process involves splitting the MIDI notes within the current measure into voice parts to process the MIDI file data. The core objective is to label each note within the same measure based on its start time and duration, and determine whether it should be classified as a chord. A "chord" is multiple notes that start at the same time and have the same duration, and should be treated as the same notation unit (multiple pitches within a single note object). "Multiple voice parts" are notes that start at the same time but have different durations, or notes that progress along different timelines, and should be assigned to different voice parts, with each part independently arranged and timed.
[0114] In some embodiments, the following processing flow is included:
[0115] 1. Parse MIDI files → trackValue
[0116] Results: The output includes header(ppq(time resolution) / tempo(tempo) / timeSig(time signature)), tracks[](tracks), ticks(minimum time unit), durationTicks(note duration), MIDI, etc. for each track.notes[].
[0117] Purpose: To provide raw data for subsequent calculations of measure duration and note assignment.
[0118] 2. Create the basic data structure → musicScoreData
[0119] Result: Construct an empty MusicScoreData and write the spectral signature (puHao), beat number (paiHao), and beat speed (paiSu) mapping.
[0120] correspond:
[0121] Multi-score table: Pre-create FuPuBiao rows (insertMidiFuPuBiao) according to the total number of measures, with 4 measures per row.
[0122] 3. Process each audio track in a loop → parserTrackValue
[0123] Results: Traverse the tracks containing notes, initialize the time accumulator and index for each track; place a single staff on each complex staff for each track, and calculate its vertical position.
[0124] correspond:
[0125] Single Spectrum: Create a MusicScoreDanPuBiao for each valid track and add it to the current FuPuBiao.danPuBiaoArray.
[0126] The height of the complex spectral table is determined by the bottom of the last single spectral table.
[0127] 4. Create a single staff for each track → insertMidiDanPuBiao
[0128] Result: Returns MusicScoreDanPuBiao with index (MusicScoreIndex:fuPuBiaoIndex,danPuBiaoIndex).
[0129] correspond:
[0130] The single staff is ready (measures and notes will be added to it later).
[0131] 5. Create measures and assign notes → parserMusicScoreXiaoJie
[0132] Result: MusicScoreXiaoJie was created sequentially on each individual spectrogram:
[0133] Calculate the measure duration: totalDuration = ppq * (4 / time signature denominator) * time signature numerator.
[0134] Set the start time and horizontal position of the measure; write the time signature / tempo at the start ticks (if any changes are needed).
[0135] Scan track.notes and collect the notes that fall within the time window of that section into currentYinFuArray.
[0136] Multi-voice part allocation: Use shengBuTicks to determine
[0137] Simultaneous start + different durations → Voice 1
[0138] Simultaneous start + equal duration → chord (voice 0)
[0139] Immediately following the end of voice part 0 or 1 → continuing with the same voice part
[0140] Other → Default Voice Part 0
[0141] correspond:
[0142] Section: insertMidiXiaoJie creates MusicScoreXiaoJie and attaches it to DanPuBiao.xiaoJieArray.
[0143] Polyphonic assignment: Mark each note with note.shengBu as the basis for subsequent note construction.
[0144] 6. Create a specific note object → createMusicScoreYinFu
[0145] result:
[0146] Convert currentYinFuArray to MusicScoreYinFu by time and vocal part:
[0147] Calculate the spectral duration value yinChang (based on ppq and durationTicks), and split it using isSumOfTwoNPowers if necessary.
[0148] The 0th pitch of the same tick is combined into a chord (multiple MusicScoreYinFuContent).
[0149] Voice 0 is written to shengBuArray[0]; Voice 1 is written to shengBuArray[1] (if it does not exist, create MusicScoreShengBu).
[0150] Maintain two timelines, showTime and playTime (voice 0 and voice 1 advance independently).
[0151] Set the direction, UI, puHao, timing, and slur (across splits or across measures).
[0152] setFuGang: Groups sequences of notes smaller than a quarter note using bar groups.
[0153] correspond:
[0154] Note: MusicScoreYinFu + MusicScoreYinFuContent (including chords).
[0155] Multivoice: MusicScoreXiaoJie.shengBuArray[0 / 1].yinFuArray forms two independent voice chains.
[0156] 7. Set the layout and connection relationships → setupAllXiaoJieYinFuOrigin.
[0157] Results: Calculate the final coordinates and layout anchor points for all measures / notes; setupLianYinXian completes the slur endings.
[0158] correspond:
[0159] The geometric layout of the staff, clef, measure, and note is well-designed and can be rendered directly.
[0160] Final product:
[0161] A complete MusicScoreData:
[0162] fuPuBiaoArray[n]: Multiple rows (multiple staves), each row contains several danPuBiao (single staves / tracks).
[0163] Each DanPuBiao.xiaoJieArray[m]: sequential measure, containing the beat number / beat speed and time range.
[0164] Each XiaoJie.shengBuArray[0..1]: a multi-voice container.
[0165] Each voice part's yinFuArray[k]:MusicScoreYinFu (containing multiple YinFuContents for chords) includes slurs, bars, directions, and coordinates. This corresponds to the complete generation chain of compound staff → single staff → measure → note → multi-voice, and the data can be used for front-end score rendering and playback.
[0166] In some embodiments, the following processing flow is included:
[0167] S2-1 Initialization Phase:
[0168] Create an instance to save the current context;
[0169] Initialize the JMCData and EditorData objects, and store the editor data (EditorData) into the JMC data collection;
[0170] S2-2 MIDI file loading:
[0171] MIDI file data is loaded asynchronously using Midi.fromUrl(); upon successful loading, a Promise callback is executed to process the track data trackValue.
[0172] S2-3 Head Data Processing (Analyzing Shooting Speed and Shot Number):
[0173] Extract the MIDI header information `headerData`; process the time signature information `timeSignatures`, calculate the basic time unit `ppq` (the number of ticks per quarter note); parse the `trackValue.header.tempos` array to generate a tempo array of the corresponding length, where the tempo value is `trackValue.header.tempos[i].bpm` (where `i` is each item in the array). The rest uses default tempo data; `timeVakue=64` means the default unit is quarter notes, and `xuanPuOrigin` and `jianPuOrigin` are the tempo positions displayed on the score.
[0174] Call the createPaiSuWithMidiTempos method to process speed change information;
[0175] S2-4 Duration Calculation:
[0176] The formula for calculating the total time based on the race number is: (ppq × (4 / denominator)) × numerator
[0177] For example, the duration of a 4 / 4 beat is calculated as (ppq × (4 / 4)) × 4 = ppq × 4
[0178] S2-5 section processing (calculating the total number of sections created):
[0179] Call the getMusicScoreXiaoJieCount method to calculate the total number of sections;
[0180] Create the basic data structure musicScoreData.
[0181] S2-6 Spectrum Generation and Processing (Creating Complex Spectra):
[0182] The `insertMidiFuPuBiao` function creates a complex spectral structure; based on the previously calculated number of measures, a complex spectral is generated every four measures. The clef and time signature information are set; a default clef is generated here; the top position reference values for the staff notation and simplified notation are initialized.
[0183] Layout processing for section S2-7:
[0184] Each line displays 4 sections by default, with the last line having a special handling for the remaining sections;
[0185] The number of sections per line is controlled by xiaoJieCountFromFuPuBiao;
[0186] The `parserTrackValue` function parses the track data into a specific spectrogram; this method is for multi-voice parsing (see image below).
[0187] P1. Determine if it is a polyphonic voice based on the length of tack.notes, then insertMidiDanPuBiao;
[0188] Create the corresponding number of single spectra.
[0189] S2-8 Visual Layout Processing:
[0190] The Y-axis position of each complex spectrum table is dynamically calculated;
[0191] The spacing between the spectra is controlled by the constants XianPuFuPuBiaoSpace and JianPuFuPuBiaoSpace.
[0192] Finally, setupAllXiaoJieYinFuOrigin and setupLianYinXian are called to complete the note positions and slur layout;
[0193] S2-9 Shooting Speed Processing:
[0194] The createPaiSuWithMidiTempos method converts MIDI tempo information into a music tempo object;
[0195] Each beat rate object contains:
[0196] BPM value (type) rounded to the nearest whole number;
[0197] Fixed time value 64 (timeValue);
[0198] Coordinate offset of staff notation / simplified notation (xianPuOrigin / jianPuOrigin).
[0199] Section S2-10 Calculation Module:
[0200] The getMusicScoreXiaoJieCount method iterates through the MIDI track data, calculates the total number of ticks for each track, divides it by the number of ticks per measure, and rounds up to get the number of measures.
[0201] Math.max is used to preserve the maximum number of bars in all tracks, ensuring that the score can display all content completely.
[0202] S2-11 Spectrum Layout Module:
[0203] The insertMidiFuPuBiao method creates a complex spectrum object according to the rule of 4 measures per row;
[0204] The position of the spectrum is controlled by the fuPuBiaoIndex index, and the generated spectrum is stored in the editor data.
[0205] S2-12 Music Notation Generation Module:
[0206] `createPuHaoWithMidi` creates treble and bass clef objects, containing detailed indexing and positioning information.
[0207] The clef content is defined using the MusicScoreClefObject constant, supporting the display requirements of different spectral tables;
[0208] createPaiHaoWithTimeSignatures processes time signature data, converting MIDI time signature information into renderable objects;
[0209] S2-13 Track Analysis Module:
[0210] The parserTrackValue method processes single-track data and initializes note position indices and duration accumulation objects.
[0211] The vertical positions of the staff and simplified musical notation are dynamically calculated, and the spacing is controlled by xianPuSpace / jianPuSpace.
[0212] Finally, the spectral height is updated to fit the actual content.
[0213] Although the present invention has been described in detail above with general descriptions, specific embodiments, and experiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for converting an image to sheet music, comprising first obtaining a MIDI file from a note image through note recognition, and then converting the MIDI file to sheet music data, wherein the sheet music data includes several vocal parts, characterized in that, The music score conversion includes the following multi-voice part assignment steps: The MIDI file is parsed to obtain several notes to be assigned, the notes to be assigned including start time and duration; When the note to be assigned is the first note, the note to be assigned is assigned to the first voice part, the end time of the first voice part, the start time of the first voice part, and the duration of the first voice part are recorded and marked as the first voice part; When the note to be assigned is not the first note, if the non-first note has the same start time as the first note but a different duration, the note to be assigned is assigned to the second voice part, the end time of the second voice part, the start time of the second voice part and the duration of the second voice part are recorded, and it is marked as the second voice part; When the note to be assigned is not the first note, if the non-first note has the same start time and the same duration as the first note, then the note to be assigned is assigned to the first voice and marked as a first voice chord. When the note to be assigned begins immediately after the end of the first voice part, it continues to be assigned to the first voice part; When the note to be assigned begins immediately after the end of the second voice part, it continues to be assigned to the second voice part; If the voice part to which the note to be assigned belongs cannot be determined, it will be assigned to the first voice part by default.
2. The method as described in claim 1, characterized in that, The musical score data includes several complex staves, each complex stave includes several simple staves, each simple stave includes several measures, and each measure includes several voice parts; The MIDI file includes header data and several audio tracks, each containing several notes to be assigned.
3. The method as described in claim 2, characterized in that, The music score conversion includes the following steps for creating a complex musical score: Parse the header data of the MIDI file; Calculate the number of sections to be created; Create the basic data structure for the musical score data; Create the complex spectrum table based on the number of subsections in each row; The duration of the current track in each single clef is accumulated in the complex clef; Create a note cycle index for each single staff in the complex staff; Create the last row of processing logic in the complex spectrum table; Set the position for the complex spectrum.
4. The method as described in claim 3, characterized in that, The music score conversion includes the following single staff creation steps: Create a beat rate object; Create a spectral object; Create a sequence number object; Process single-spectral data in a loop; Create a single spectrum.
5. The method as described in claim 4, characterized in that, The music score conversion includes the following steps for creating sections: 1) Create the basic section object, specifically including: Clone the note cycle index of the single staff; Get the total duration of the current audio track; Set the index position of the subsection in the complex spectrum table; Create a section object and add it to the section array of the single spectrometer. 2) Set the horizontal layout of the sections, specifically including: Set the X coordinate of the musical line; Set the X-coordinate of the simplified musical notation; Calculate the starting position of the next section, which is the right boundary of the current section; 3) Set the time attributes for the section, specifically including: setting the section start time; 6) Calculate the duration of each segment, specifically including: Calculate the duration of the current measure based on the current time signature; Get the start time of the current section, add up the duration of the current section, and prepare for the calculation of the next section.
6. The method as described in claim 5, characterized in that, The section creation steps also include: 4) Assign time signatures, specifically including: if there is a time signature change at the current time, record the time signature used by the current audio track; 5) Allocate beat rate, specifically including: if the single spectral is the first single spectral of the first complex spectral, allocate beat rate.
7. The method as described in claim 6, characterized in that, The music score conversion includes the following note creation steps: Create a complete musical note object; Create a note content object; Set various attributes for the musical note, including duration, pitch, and display time; Distribute the notes to the correct voices according to the polyphonic parting steps described above; Set the clef information for the notes.
8. The method as described in claim 7, characterized in that, The music score conversion includes the following note assignment steps: Iterate through all the notes in the track and extract the notes belonging to the current measure; The track data is split into notes for each measure. The note array for the current measure; Voice part time tracking object, used for multi-voice judgment; Iterate through all notes in the current audio track; Note start time (MIDI tick); Note duration (MIDI tick); Determine whether a note belongs to the time range of the current measure.
9. The method as described in claim 1, characterized in that, It also includes the following steps: When the note duration is incorrect, split the MIDI notes in the current measure into voice parts.
10. A system for converting an image to sheet music, the system comprising at least one processor; and a memory storing instructions that, when executed by the at least one processor, implement the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Interactive digital music teaching system
CN109345905A
Music score play module assembly of digital music teaching system
CN109377818A