Music spectrum conversion method and device, storage medium and computer program product

Through the synergistic effect of the autoregressive lyrics and pitch prediction models, the problem of identifying lyrics corresponding to multiple pitches is solved, efficient and accurate music transcription is achieved, high-quality music scores are generated, and the operation process is simplified.

CN120673727APending Publication Date: 2025-09-19TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510895628.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing music transcription technology has difficulty accurately identifying complex situations where lyrics correspond to multiple pitches, resulting in inaccurate transcription results. It also requires manual labeling and multiple model comparisons, resulting in long links and low accuracy, and cannot meet the needs of efficient and accurate transcription.

Method used

By using the synergistic effect of the autoregressive lyrics and lyrics duration prediction model, the autoregressive lyrics and pitch number prediction model, and the non-autoregressive pitch prediction model, accurate music scores are generated by predicting the lyrics duration and pitch number of the target song audio.

Benefits of technology

It achieves efficient conversion from audio to musical scores, improves the accuracy of transcription, simplifies the operation process, and can automatically generate high-quality musical scores without the need for professional music theory knowledge and manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673727A_ABST
    Figure CN120673727A_ABST
Patent Text Reader

Abstract

The invention discloses a music score conversion method, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a target song audio; performing lyric time value prediction on the target song audio to obtain a lyric word sequence and a time value corresponding to each lyric word in the lyric word sequence; performing pitch quantity prediction on the target song audio to obtain a pitch quantity corresponding to each lyric word in the lyric word sequence; determining a pitch value corresponding to each lyric word in the lyric word sequence based on the hour value and the pitch quantity corresponding to each lyric word; and generating a music score corresponding to the target song audio based on the lyric word sequence and the hour value and the pitch value corresponding to each lyric word in the lyric word sequence. According to the invention, the operation process of music spectrum conversion is simplified, and the accuracy of music spectrum conversion is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of music technology, and more specifically, to a music transcription method, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In the field of music transcription, related technologies have many limitations. Traditional manual notation is costly, inefficient, relies on professionals, and is difficult to operate in batches. Automatic transcription technology also has obvious defects, especially when faced with complex situations where a lyric word corresponds to multiple pitches, which makes it difficult for existing technologies to accurately identify. For example, when there are inflections or pitch changes in the lyrics, existing algorithms are often unable to accurately capture the details of each pitch, resulting in inaccurate transcription results. In addition, existing technologies require manual annotation of lyrics or multiple model comparisons, which have long links and low accuracy, and cannot meet the needs of music creation and other fields for efficient and accurate transcription.

[0003] Therefore, how to improve the accuracy of music transcription is a technical problem that those skilled in the art need to solve. Summary of the Invention

[0004] The purpose of this application is to provide a music transcription method, electronic device, computer-readable storage medium and computer program product, which improve the accuracy of music transcription.

[0005] To achieve the above-mentioned purpose, the present application provides a first aspect of a music transcription method, comprising:

[0006] Get the target song audio;

[0007] Performing lyric duration prediction on the target song audio to obtain a lyric word sequence and a duration corresponding to each lyric word in the lyric word sequence;

[0008] Predicting the number of pitches of the target song audio to obtain the number of pitches corresponding to each word in the lyrics word sequence;

[0009] Determining the pitch value corresponding to each lyric word in the lyric word sequence based on the duration and pitch number corresponding to each lyric word;

[0010] A music score corresponding to the target song audio is generated based on the lyrics word sequence, the duration value and the pitch value corresponding to each lyrics word in the lyrics word sequence.

[0011] To achieve the above-mentioned object, the second aspect of the present application provides an electronic device, including:

[0012] memory for storing computer programs;

[0013] A processor is used to implement the steps of the above-mentioned music transcription method when executing the computer program.

[0014] To achieve the above-mentioned purpose, the third aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned music transcription method are implemented.

[0015] To achieve the above-mentioned purpose, the fourth aspect of the present application provides a computer program product, including a computer program, which implements the steps of the above-mentioned music transcription method when executed.

[0016] The music transcription method provided by the present application firstly predicts the duration of the lyrics of the target song audio, thereby accurately extracting the sequence of lyrics and determining the duration corresponding to each lyric word. Secondly, by predicting the number of pitches of the target song audio, the number of pitches corresponding to each lyric word can be effectively identified. In music, a lyric word may correspond to multiple pitches, especially when there are ornamental sounds such as glissando and portamento. Related technologies often have difficulty in accurately judging the number of pitches, resulting in inaccurate transcription results. However, the present application accurately identifies the number of pitches corresponding to the lyric words by analyzing the pattern of pitch changes in the audio signal, providing a basis for subsequent predictions. Furthermore, based on the duration and number of pitches of each lyric word, the pitch value corresponding to each lyric word is accurately determined. Through the synergistic effect of the above steps, efficient conversion from audio to music score is achieved, solving the technical problem of the related technology that it is unable to identify multiple pitches corresponding to a single lyric word, and improving the accuracy of transcription. Because the related technology cannot accurately identify the number of pitches of lyric words when processing complex audio environments, the transcription results are inaccurate. This application further improves the accuracy of music transcription by accurately identifying the number of pitches and combining it with precise pitch value prediction. Because the entire process requires no human intervention, users do not need professional music theory knowledge or complex operating skills; they only need to input the target audio to automatically generate the music score, simplifying the music transcription process. This application also discloses an electronic device, a computer-readable storage medium, and a computer program product that can also achieve the above-mentioned technical effects.

[0017] It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. The drawings are used to provide a further understanding of the present disclosure and constitute part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure, but do not constitute a limitation of the present disclosure. In the drawings:

[0019] Figure 1 A flowchart of a music transcription method provided in an embodiment of the present application;

[0020] Figure 2 A flowchart of a lyrics and lyrics duration prediction model provided in an embodiment of the present application;

[0021] Figure 3 A flowchart of a lyrics and pitch quantity prediction model provided in an embodiment of the present application;

[0022] Figure 4 A flowchart of a pitch prediction model provided in an embodiment of the present application;

[0023] Figure 5 A flowchart of another music transcription method provided in an embodiment of the present application;

[0024] Figure 6 An inference process for each target audio provided in an embodiment of the present application;

[0025] Figure 7 A flowchart of a timbre adaptation provided in an embodiment of the present application;

[0026] Figure 8 A flowchart of a timbre and lyrics adaptation provided in an embodiment of the present application;

[0027] Figure 9 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0029] In related technologies, music transcription is divided into manual and automatic transcription. Manual transcription primarily uses MuseScore or various vocal synthesis studio tools, such as ACE Studio and X Studio, to manually mark pitch and duration. However, this is time-consuming and requires the annotator to have a certain level of music theory knowledge, making it a relatively high barrier to entry. While automatic transcription can generate MIDI (Musical Instrument Digital Interface) notation, most of these only contain pitch and duration, without lyrics. This still requires manual annotation of lyrics note by note, and attention must be paid to issues such as pitch transitions during the annotation process. Alternatively, some solutions utilize lyric alignment technology, using lyrics and audio to obtain the timestamp of each word. This is then compared to the pitch and duration sequence of each MIDI note to automatically fill in the lyrics. However, this method is time-consuming, prone to lyric misalignment, and lacks high accuracy.

[0030] The music score conversion method provided by this application only requires the input of music files, without the need to prepare lyrics and melody in advance. The model directly generates the corresponding pitch, duration, and lyrics end-to-end, and the lyrics and pitch are automatically aligned. The complete music score is directly output through the coordinated cooperation of three models. These three models are: (1) an autoregressive lyrics and lyrics duration prediction model, which autoregressively outputs the words and word durations corresponding to the audio; (2) an autoregressive lyrics and pitch number prediction model, which autoregressively outputs the words and pitch numbers corresponding to the audio; and (3) a non-autoregressive pitch prediction model, which predicts the pitch value corresponding to each word based on the outputs of the first two models. This application achieves efficient conversion from audio to music score through the synergy of the three models, solves the technical problem in related technologies that cannot identify multiple pitches corresponding to a single lyric word, and improves the accuracy of the score conversion.

[0031] The embodiment of the present application discloses a method for transcribing music, which improves the accuracy of music transcribing.

[0032] See also Figure 1 , a flowchart of a music transcription method provided in an embodiment of the present application, such as Figure 1 As shown, including:

[0033] S101: Obtain target song audio;

[0034] The target song audio refers to the audio file that needs to be transcoded, which can be a song, a vocal segment, etc. In this step, the target song audio file can be uploaded by the user or obtained from historical data to provide the original material for the transcoding process.

[0035] S102: Predicting the duration of lyrics for the target song audio to obtain a lyrics word sequence and a duration corresponding to each word in the lyrics word sequence;

[0036] The lyric word sequence refers to the text content of the lyrics, a set of lyric words arranged in order. The duration corresponding to the lyric word refers to the duration of the lyric word, which can be expressed in time units such as milliseconds.

[0037] In this step, audio processing techniques, such as speech recognition, are used to extract the lyrics from the target song audio and convert them into a textual sequence of lyrics. Simultaneously, an audio analysis algorithm can be used to analyze the temporal position of each lyric word within the audio, calculating the start time and duration of each lyric word, i.e., the lyric duration.

[0038] As a feasible implementation method, lyrics duration prediction is performed on the target song audio to obtain a lyrics word sequence and the duration corresponding to each lyrics word in the lyrics word sequence, including: inputting the target song audio into an autoregressive lyrics and lyrics duration prediction model to predict the lyrics word sequence corresponding to the target song audio and the duration corresponding to each lyrics word in the lyrics word sequence.

[0039] Among them, the autoregressive model is a model that uses the output results of previous moments when predicting the output of the current moment. It has memory and can capture the time dependency in sequence data.

[0040] In a specific implementation, the target audio is input into an autoregressive lyrics and lyrics duration prediction model. The model can be a whisper structure based on the field of speech recognition. By analyzing the audio signal, the target audio corresponding lyrics text and the timestamp of each word in the lyrics are output, including the start time and the end time. Then, the duration of each word is calculated, that is, the end time minus the start time. For example, Figure 2 As shown, the input of the lyrics and lyrics duration prediction model is the time-frequency diagram on the left, such as a waveform diagram or a spectrum diagram, which reflects the characteristic distribution of audio in time and frequency. The output is the lyrics text and the duration corresponding to each lyric word. The lyrics text is: You are the origin of my SP, and the duration corresponding to each lyric word is: 5, 5, 7, 20, 5, 10, 12, 16.

[0041] It can be seen that this step can accurately convert the lyrics content in the audio into text form and determine the time position and duration of each word in the audio, providing accurate time alignment information for subsequent pitch prediction and music score generation, so that the generated music score can accurately correspond to the time and lyrics.

[0042] S103: Predicting the number of pitches of the target song audio to obtain the number of pitches corresponding to each word in the lyrics sequence;

[0043] Among them, the number of pitches refers to the number of notes corresponding to a word in the lyrics during the singing process. Usually one word corresponds to one pitch, but in some singing with techniques such as vibrato and tremolo, one word may correspond to multiple pitches.

[0044] In this step, the pitch change pattern of the target song audio is analyzed, and the number of different pitches corresponding to each lyric word is detected through an audio signal processing algorithm. For example, spectral analysis is performed on the target song audio to identify the characteristics of the pitch changes, and the number of pitches of each lyric word is determined based on these characteristics.

[0045] As a feasible implementation method, the pitch number of the target song audio is predicted to obtain the pitch number corresponding to each lyric word in the lyric word sequence, including: inputting the target song audio into an autoregressive lyrics and pitch number prediction model to predict the pitch number corresponding to each lyric word in the lyric word sequence.

[0046] In the specific implementation, the target audio is input into the autoregressive lyrics and pitch number prediction model, which can also be a whisper structure based on the field of speech recognition. By analyzing the audio, the lyrics text corresponding to the audio and the number of pitches contained in each word in the lyrics (note_number) are output. For example, if a word has no transposition, then the corresponding pitch number is 1; if a word has a transposition, then the corresponding pitch number is 2, and so on. For example, if Figure 3 As shown, the input of the lyrics and pitch number prediction model is the time-frequency diagram on the left, such as a waveform diagram or a spectrogram, which reflects the characteristic distribution of audio in time and frequency. The output is the lyrics text and the number of pitches corresponding to each word of the lyrics. The lyrics text is: You are the origin of my SP, and the number of pitches corresponding to each word of the lyrics are: 1, 1, 1, 2, 1, 1, 1, 1.

[0047] It can be seen that this step can determine in advance the number of pitches that each word of the lyrics may involve during the singing process, providing key information for the accurate prediction of subsequent pitch values, avoiding errors or inaccurate predictions caused by uncertainty about the number of pitches in the pitch prediction stage, and improving the accuracy and reliability of the entire transcription process.

[0048] S104: determining a pitch value corresponding to each lyric word in the lyric word sequence based on the duration and pitch number corresponding to each lyric word;

[0049] The pitch value refers to the specific numerical value of the pitch corresponding to each word in the lyrics, usually expressed in notes or frequencies.

[0050] In this step, the pitch value corresponding to each lyric word is further determined through an audio analysis algorithm in combination with the duration and number of pitches of the lyric words. For example, a spectral analysis is performed on the target song audio, and the corresponding pitch features are extracted based on the time position and number of pitches of the lyric words, and then converted into specific pitch values.

[0051] As a feasible implementation method, the pitch value corresponding to each lyric word in the lyric word sequence is determined based on the duration and pitch number corresponding to each lyric word, including: using a non-autoregressive pitch prediction model to predict the pitch value corresponding to each lyric word according to the duration and pitch number corresponding to each lyric word.

[0052] Among them, the non-autoregressive model does not rely on the output results of the previous moment during the prediction process, but predicts all the content that needs to be predicted in parallel, which is usually computationally efficient.

[0053] In practice, the duration of each lyric word output by the autoregressive lyrics and duration prediction model, along with the number of pitches for each lyric word output by the autoregressive lyrics and pitch count prediction model, are fed into a non-autoregressive pitch prediction model. This model primarily uses the Whisper encoder module, combined with two multi-classification models, to analyze the audio frame-level feature sequence (word dur) and predict the corresponding pitch value for each lyric word based on the duration and number of pitches. For example, for a lyric word with two pitches, the model predicts two specific pitch values ​​based on its duration and audio features.

[0054] As a feasible implementation method, the use of a non-autoregressive pitch prediction model to predict the pitch value corresponding to each lyric word according to the duration and pitch quantity corresponding to each lyric word includes: inputting the target audio into the non-autoregressive pitch prediction model; in the non-autoregressive pitch prediction model, using an encoding module to extract the feature sequence of the target audio frame level, and determining the feature sequence corresponding to each lyric word based on the duration corresponding to each lyric word; using a first multi-classification module to predict the duration of each pitch corresponding to each lyric word based on the feature sequence and pitch quantity corresponding to each lyric word; determining the feature sequence of each pitch corresponding to each lyric word based on the duration of each pitch corresponding to each lyric word; and using a second multi-classification module to predict the pitch value corresponding to each lyric word based on the feature sequence of each pitch corresponding to each lyric word.

[0055] In a specific implementation, the pitch prediction model includes an encoding module, a first multi-classification module, and a second multi-classification module. The pitch prediction model first receives the target audio and then uses the encoding module to extract a frame-level feature sequence from the audio. Frame-level features refer to the characteristic representation of the audio signal within each short period of time, reflecting information such as pitch and loudness at that point in time. The model then uses the duration of each lyric character to determine the feature sequence for each character. Duration refers to the duration of each lyric character in the audio. Based on this, the pitch prediction model uses the first multi-classification module. This step predicts the duration of each pitch based on the feature sequence and the number of pitches. The number of pitches refers to the number of pitches a lyric character may involve during a performance. For example, a lyric character may correspond to a single pitch or multiple pitches during a transition. The output of the first multi-classification module is the transition probability corresponding to each frame feature. For each lyric character's frame, the k pitch boundaries with the highest transition probability are extracted, where k = note_number - 1. This divides the duration corresponding to the lyric character into note_number parts, yielding the duration corresponding to each pitch. If note_number = 1, then k = 0. In this case, the duration corresponding to the lyric word is the pitch duration. If note_number = 2, then k = 1. In this case, the duration corresponding to the lyric is divided into two parts by the pitch boundary divided by topk = 1, and so on. After the first multi-classification module predicts each pitch duration, it further determines the feature sequence of each pitch. Finally, the pitch prediction model uses the second multi-classification module to predict the specific pitch value corresponding to each lyric word based on the feature sequence of each pitch. This prediction is non-autoregressive, which means that the model predicts all pitch values ​​in parallel instead of predicting them sequentially one by one, thereby improving computational efficiency.

[0056] For example, Figure 4As shown, the pitch prediction model is implemented using a whisper encoder and a multi-classification model. Its input is the audio's time-frequency graph (the waveform or spectrogram on the left), the duration of each lyric word (word dur), and the number of pitches (note num). The output is the predicted pitch value for each lyric word. The sequence "61|59|57|57 59|0|57|57" in the figure is an example of a pitch value sequence, representing the pitch value corresponding to each lyric word. Different lyric words are separated by "|". Within the pitch prediction model, the whisper encoder extracts a frame-level feature sequence from the audio. The pitch-related frame-level features extracted from the feature sequence output by the whisper encoder are also known as PPG frame-level features. Linear and Sigmoid layers are used to process these PPG frame-level features, generating the probability of pitch transitions for each frame and, in turn, obtaining a sequence of pitch durations. The example pitch duration sequence "5|5|7|911|10|12|16" in the figure represents the duration of each pitch corresponding to each lyric character, with different lyric characters separated by "|". The meaning pool performs pooling on the PPG frame-level features and pitch duration sequence, extracting the semantic features of each lyric character. The pooled features are then processed by linear and Sigmoid layers to predict the pitch value corresponding to each lyric character.

[0057] It can be seen that this step can accurately predict the pitch value corresponding to each lyric word based on the known duration and number of pitches of the lyric words, so that the generated music score is highly matched with the original audio in terms of pitch, further improving the generated content of the music score, so that the generated music score can truly reflect the melody information of the original audio.

[0058] S105: Generate a music score corresponding to the target song audio based on the lyrics word sequence, the duration value and the pitch value corresponding to each lyrics word in the lyrics word sequence.

[0059] Among them, musical notation refers to a written form that uses symbols to record the pitch, duration, rhythm and other information of music. Common ones include simplified notation, staff notation, MIDI notation, etc.

[0060] In this step, the predicted lyric sequence corresponding to the target song audio, along with the duration and pitch of each lyric, is integrated and, following the standard format and rules of musical notation, a corresponding score file is generated. For example, the pitch, duration, and lyrics of each lyric are arranged in chronological order to form a complete musical score. This can be in MIDI format or other common musical notation formats. For example, a MIDI score might look like this: each line contains the start time (ms), duration (ms), pitch, and lyric word: 40919 200 61 you; 41119 200 59 just; 41319 311 57 Yes; 41630 311 57 me; 42779 698 59 -; 43477 698 59; 44175 202 57 original; 44377 1003 57 points.

[0061] As a preferred implementation, this embodiment also includes: constructing a training data set; wherein, the training data set includes multiple groups of training data, and the training data includes training audio clips and training music scores corresponding to the training audio clips; based on the training data set, training an autoregressive lyrics and lyrics duration prediction model, an autoregressive lyrics and pitch quantity prediction model, and a non-autoregressive pitch prediction model.

[0062] In the specific implementation, a training dataset is constructed. This training dataset contains multiple sets of training data, each consisting of a training audio clip and its corresponding training musical score. An audio clip is part of a song, while the corresponding musical score contains information such as lyrics, pitch, and duration. To prepare this training data, a large number of MIDI files with lyrics and corresponding audio files must be collected. These files must be strictly aligned in terms of pitch, duration, and lyrics to ensure data consistency and accuracy. Alignment means that each audio segment can be matched to the corresponding note and lyrics in the musical score, which is crucial for subsequent model training. Since MIDI and audio files for entire songs are often long, directly inputting them into the model for training may result in insufficient graphics memory. Therefore, the entire song's MIDI and audio files need to be segmented into shorter segments to accommodate model training. The segmentation rule is typically based on rests in MIDI. A rest is a period of silence between two adjacent notes. When the duration of a rest exceeds a set threshold, the MIDI and audio files are segmented. This method allows a song to be broken down into multiple shorter MIDI and audio pairs, yielding a large number of training samples. This training dataset is used to train models, including an autoregressive lyrics and lyrics duration prediction model, an autoregressive lyrics and pitch quantity prediction model, and a non-autoregressive pitch prediction model. Through training on this data, the model learns how to extract lyrics and corresponding duration information from audio, and how to predict pitch values ​​based on this information, thereby achieving efficient and accurate transcription.

[0063] As a feasible implementation method, an autoregressive lyrics and lyrics duration prediction model is trained based on the training data set, including: determining the lyrics word sequence and lyrics word duration sequence corresponding to the training audio segment based on the training music score corresponding to the training audio segment; and training the autoregressive lyrics and lyrics duration prediction model based on multiple groups of the training audio segments, the corresponding lyrics word sequences and lyrics word duration sequences.

[0064] In practice, a triplet of <training audio clip, lyric word sequence, lyric word duration sequence> is extracted from the training dataset and fed into an autoregressive lyrics and lyric duration prediction model for training. This model takes audio as input and, through an autoregressive approach, progressively outputs the corresponding lyric word sequence and the duration sequence for each word, thereby learning to extract lyric text and timing information from the audio. This approach allows the model to learn the mapping between audio, lyrics, and timing, enabling it to accurately predict lyrics and duration from new audio in real-world applications.

[0065] As a feasible implementation method, an autoregressive lyrics and pitch quantity prediction model is trained based on the training data set, including: determining the lyrics word sequence and the number of lyrics word pitches corresponding to the training audio segment based on the training music score corresponding to the training audio segment; and training an autoregressive lyrics and pitch quantity prediction model based on multiple groups of the training audio segments, the corresponding lyrics word sequences and the number of lyrics word pitches.

[0066] In practice, a triplet of "training audio clip, lyric word sequence, and lyric word pitch count sequence" is extracted from the training dataset and fed into an autoregressive lyrics and pitch count prediction model for training. This model uses autoregression to learn the mapping between audio and lyric word pitch counts, enabling it to predict the corresponding pitch count for each lyric word from new audio.

[0067] As a feasible implementation method, a non-autoregressive pitch prediction model is trained based on the training data set, including: determining the lyrics word sequence, lyrics word duration sequence, number of lyrics word pitches, and lyrics word pitch value sequence corresponding to the training audio segment based on the training music score corresponding to the training audio segment; and training a non-autoregressive pitch prediction model based on multiple groups of the training audio segments, the corresponding lyrics word sequences, lyrics word duration sequences, number of lyrics word pitches, and lyrics word pitch value sequences.

[0068] In the specific implementation, a five-tuple consisting of <training audio clip, lyrics word sequence, lyrics word duration sequence, lyrics word pitch number sequence, lyrics word pitch value sequence> is extracted from the training dataset and fed into a non-autoregressive pitch prediction model for training. This model uses the word duration and pitch number information output by the first two models to learn how to predict the corresponding pitch value for each word from the audio. In this way, the model can learn the mapping relationship between audio features and pitch values, and thus accurately predict the corresponding pitch value based on the input audio and the predicted lyrics and timing information in practical applications.

[0069] The music transcription method provided in the embodiment of the present application firstly predicts the duration of lyrics for the target song audio, thereby accurately extracting the sequence of lyrics and determining the duration corresponding to each lyric word. Secondly, by predicting the number of pitches for the target song audio, the number of pitches corresponding to each lyric word can be effectively identified. In music, a lyric word may correspond to multiple pitches, especially when there are ornamental sounds such as glissando and portamento. Related technologies often have difficulty accurately determining the number of pitches, resulting in inaccurate transcription results. However, the present application accurately identifies the number of pitches corresponding to the lyric word by analyzing the pattern of pitch changes in the audio signal, providing a basis for subsequent predictions. Furthermore, based on the duration and number of pitches of each lyric word, the pitch value corresponding to each lyric word is accurately determined. Through the synergistic effect of the above steps, efficient conversion from audio to music score is achieved, solving the technical problem of the related art that it is unable to identify multiple pitches corresponding to a single lyric word, and improving the accuracy of transcription. Because the related art cannot accurately identify the number of pitches of lyric words when processing complex audio environments, the transcription results are inaccurate. This application further improves the accuracy of transcribing by accurately identifying the number of pitches and combining it with precise pitch value prediction. Since the entire process requires no human intervention, users do not need professional music theory knowledge or complex operating skills. Simply inputting the target audio automatically generates the sheet music, simplifying the process of transcribing music.

[0070] The present embodiment discloses a method for transcribing music. Compared with the previous embodiments, this embodiment further illustrates and optimizes the technical solution. Specifically:

[0071] See also Figure 5 , a flowchart of another music transcription method provided in an embodiment of the present application, such as Figure 5 As shown, including:

[0072] S201: Acquire an audio file of a target song and identify a silent segment in the audio file; wherein the silent segment is an audio segment with a silence duration greater than or equal to a preset value;

[0073] This step first requires obtaining the audio file to be processed. This can come from a variety of sources, such as user-uploaded audio or audio material stored in the system database. Voice activity detection (VAD) technology is then used to identify silent segments within the audio. These segments are defined as those whose duration meets or exceeds a preset threshold. This step allows for precise location of silent sections within the audio, providing a clear basis for subsequent audio segmentation and improving the accuracy and efficiency of the transcription process.

[0074] S202: dividing the audio file into a plurality of target song audios according to the audio segments;

[0075] In this step, the original audio file is segmented into multiple shorter target audio segments. Each segment is an independent portion of the audio, typically between two silent segments. This segmentation ensures that each target audio segment contains the complete non-silent content while avoiding the computational resource limitations and processing inefficiencies that can occur when processing very long audio.

[0076] S203: Inputting the target song audio into an autoregressive lyrics and lyrics duration prediction model to predict a lyrics word sequence corresponding to the target song audio and a duration corresponding to each lyric word in the lyrics word sequence;

[0077] S204: Inputting the target song audio into an autoregressive lyrics and pitch quantity prediction model to predict the number of pitches corresponding to each lyric word in the lyric word sequence;

[0078] S205: Inputting the target song audio into a non-autoregressive pitch prediction model, and using the non-autoregressive pitch prediction model to predict the pitch value corresponding to each word in the lyrics according to the duration and pitch number corresponding to each word in the lyrics;

[0079] S206: Generate a music score corresponding to the target song audio based on the lyrics word sequence, the duration and pitch value corresponding to each lyrics word in the lyrics word sequence;

[0080] S207: Splicing the music scores corresponding to the multiple target song audios into the music score corresponding to the target song.

[0081] The inference process for each target audio in this embodiment is as follows: Figure 6 As shown in the figure, after transcribing each target audio clip, a corresponding musical score is generated for each clip. In this step, the musical scores of these clips are spliced ​​together in the original audio order to form the final score corresponding to the complete audio file. This process requires ensuring the coherence and integrity of the score. During splicing, the start and end of each clip are precisely aligned to avoid misalignment of notes or lyrics. The spliced ​​musical score fully reflects the musical information of the entire audio.

[0082] As a preferred embodiment, after splicing the music scores corresponding to the multiple target song audios into the music score corresponding to the target song, it also includes: obtaining the dry audio of the target song, and generating a cover of the target song based on the dry audio, accompaniment audio and the lyrics sequence of the target song.

[0083] In specific implementation, the timbre adaptation focuses on changing the timbre itself while retaining the original lyrics and melody, using the user's timbre to cover classic songs, such as Figure 7 As shown. For example, platforms like Kugou and QQ Music allow users to create personalized covers of classic songs using this feature. Specifically, users upload dry vocal audio that reflects their voice. The system then uses vocal synthesis technology to combine the user's dry vocal audio with the selected song's melody (also known as the instrumental audio) and lyrics to create a unique cover. This timbre transformation enhances user interaction with classic songs, allowing users to deeply participate in the music creation process.

[0084] As another preferred embodiment, after splicing the music scores corresponding to the multiple target song audios into the music score corresponding to the target song, it also includes: obtaining the dry audio and adapted lyrics sequence of the target song, and generating a cover of the target song based on the dry audio, accompaniment audio and adapted lyrics sequence of the target song.

[0085] In practice, the adaptation of both timbre and lyrics not only changes the timbre, but also introduces a large language model (LLM) to generate new lyrics. In this process, the system will upload new lyrics based on the style and context of the original lyrics, or create new lyrics using a large language model. This is the adapted lyric word sequence. Subsequently, through vocal synthesis technology, the adapted lyric word sequence is combined with the user's dry voice audio and replaced with the original MIDI melody. This allows for an adapted cover that changes both the timbre and the lyrics, such as Figure 8 This dual adaptation method provides users with a broader creative space, allowing them to not only reinterpret songs with their own voices, but also give the songs new lyrics and create unique musical works.

[0086] This embodiment achieves efficient and accurate conversion from audio files to complete musical scores. During the inference phase, the audio segmentation step ensures that the model can process audio clips of appropriate length, improving processing efficiency and accuracy. The synergistic effect of the three prediction models enables the extraction of accurate lyrics text, duration, and pitch information from the audio, thereby generating high-quality musical scores. The final musical score splicing step ensures the integrity and coherence of the output results. As can be seen, this embodiment can not only automatically complete complex music transcription tasks, but also greatly reduce the cost and time consumption of manual music transcription.

[0087] The following describes a music transcribing device provided in an embodiment of the present application. The music transcribing device described below and the music transcribing method described above can be used in conjunction with each other. The music transcribing device provided in an embodiment of the present application includes:

[0088] Acquisition module, used to obtain the target song audio;

[0089] A first prediction module is used to predict the duration of lyrics for the target song audio, and obtain a lyrics word sequence and a duration corresponding to each word in the lyrics word sequence;

[0090] A second prediction module is used to predict the number of pitches of the target song audio to obtain the number of pitches corresponding to each word in the lyrics word sequence;

[0091] a third prediction module, configured to determine a pitch value corresponding to each lyric word in the lyric word sequence based on the duration and pitch number corresponding to each lyric word;

[0092] The first generating module is used to generate a music score corresponding to the target song audio based on the lyrics word sequence and the duration and pitch value corresponding to each lyrics word in the lyrics word sequence.

[0093] The music transcription device provided in the embodiment of the present application firstly predicts the duration of lyrics for the target song audio, accurately extracts the sequence of lyrics and determines the duration corresponding to each lyric word. Secondly, by predicting the number of pitches for the target song audio, it can effectively identify the number of pitches corresponding to each lyric word. In music, a lyric word may correspond to multiple pitches, especially when there are ornamental sounds such as glissando and portamento. Related technologies often have difficulty accurately determining the number of pitches, resulting in inaccurate transcription results. However, the present application accurately identifies the number of pitches corresponding to lyric words by analyzing the pattern of pitch changes in the audio signal, providing a basis for subsequent predictions. Furthermore, based on the duration and number of pitches of each lyric word, the pitch value corresponding to each lyric word is accurately determined. Through the synergistic effect of the above steps, efficient conversion from audio to music score is achieved, solving the technical problem of the related art that it is unable to identify multiple pitches corresponding to a single lyric word, and improving the accuracy of transcription. Because the related art cannot accurately identify the number of pitches of lyric words when processing complex audio environments, the transcription results are inaccurate. This application further improves the accuracy of transcribing by accurately identifying the number of pitches and combining it with precise pitch value prediction. Since the entire process requires no human intervention, users do not need professional music theory knowledge or complex operating skills. Simply inputting the target audio automatically generates the sheet music, simplifying the process of transcribing music.

[0094] Based on the above embodiment, as a preferred implementation, the first prediction module is specifically configured to: input the target song audio into an autoregressive lyrics and lyrics duration prediction model to predict the lyrics word sequence corresponding to the target song audio and the duration corresponding to each lyrics word in the lyrics word sequence;

[0095] The second prediction module is specifically configured to: input the target song audio into an autoregressive lyrics and pitch quantity prediction model to predict the number of pitches corresponding to each lyric word in the lyric word sequence;

[0096] The third prediction module is specifically used to: input the target song audio into a non-autoregressive pitch prediction model; in the non-autoregressive pitch prediction model, use the encoding module to extract the feature sequence of the target audio frame level, and determine the feature sequence corresponding to each lyric word based on the time value corresponding to each lyric word; use the first multi-classification module to predict the time value of each pitch corresponding to each lyric word based on the feature sequence and pitch number corresponding to each lyric word; determine the feature sequence of each pitch corresponding to each lyric word based on the time value of each pitch corresponding to each lyric word; use the second multi-classification module to predict the pitch value corresponding to each lyric word based on the feature sequence of each pitch corresponding to each lyric word.

[0097] Based on the above embodiment, as a preferred implementation, it further includes:

[0098] A construction module is used to construct a training data set; wherein the training data set includes multiple sets of training data, and the training data includes training audio segments and training music scores corresponding to the training audio segments;

[0099] A training module is used to train an autoregressive lyrics and lyrics duration prediction model, an autoregressive lyrics and pitch quantity prediction model, and a non-autoregressive pitch prediction model based on the training data set.

[0100] Based on the above embodiments, as a preferred implementation mode, the training module is specifically used to: determine the lyrics word sequence and lyrics word duration sequence corresponding to the training audio segment according to the training music score corresponding to the training audio segment; and train an autoregressive lyrics and lyrics duration prediction model based on multiple groups of the training audio segments, the corresponding lyrics word sequences and lyrics word duration sequences.

[0101] Based on the above embodiments, as a preferred implementation mode, the training module is specifically used to: determine the lyrics word sequence and the number of lyrics word pitches corresponding to the training audio segment according to the training music score corresponding to the training audio segment; and train an autoregressive lyrics and pitch number prediction model based on multiple groups of the training audio segments, the corresponding lyrics word sequences and the number of lyrics word pitches.

[0102] On the basis of the above embodiments, as a preferred implementation mode, the training module is specifically used to: determine the lyrics word sequence, lyrics word duration sequence, lyrics word pitch number, lyrics word pitch value sequence corresponding to the training audio segment according to the training music score corresponding to the training audio segment; train a non-autoregressive pitch prediction model based on multiple groups of the training audio segments, corresponding lyrics word sequences, lyrics word duration sequences, lyrics word pitch number, lyrics word pitch value sequences.

[0103] Based on the above embodiment, as a preferred implementation, the acquisition module is specifically configured to: acquire an audio file of a target song, identify silent segments in the audio file; wherein the silent segments are audio segments with a silence duration greater than or equal to a preset value; and segment the audio file into a plurality of target audios according to the audio segments;

[0104] Accordingly, the device further includes:

[0105] The splicing module is used to splice the music scores corresponding to the multiple target song audios into the music score corresponding to the target song.

[0106] Based on the above embodiment, as a preferred implementation, it further includes:

[0107] The second generation module is used to obtain the dry audio of the target song, and generate a cover of the target song based on the dry audio, accompaniment audio and the lyrics sequence of the target song.

[0108] Based on the above embodiment, as a preferred implementation, it further includes:

[0109] The third generation module is used to obtain the dry audio and adapted lyrics sequence of the target song, and generate a cover of the target song based on the dry audio, accompaniment audio and adapted lyrics sequence of the target song.

[0110] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0111] This application also provides an electronic device, see Figure 9 , a structural diagram of an electronic device 90 provided in an embodiment of the present application, such as Figure 9 As shown, a processor 91 and a memory 92 may be included.

[0112] Processor 91 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 91 may be implemented in at least one of the following hardware forms: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 91 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 91 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content displayed on the display screen. In some embodiments, processor 91 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0113] The memory 92 may include one or more computer-readable storage media, which may be non-transitory. The memory 92 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In this embodiment, the memory 92 is used to store at least the following computer program 921, wherein, after the computer program is loaded and executed by the processor 91, it can implement the relevant steps in the music transcription method performed by the electronic device side disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 92 may also include an operating system 922 and data 923, and the storage method may be temporary storage or permanent storage. Among them, the operating system 922 may include Windows, Unix, Linux, etc.

[0114] In some embodiments, the electronic device 90 may further include a display screen 93 , an input / output interface 94 , a communication interface 95 , a sensor 96 , a power supply 97 , and a communication bus 98 .

[0115] certainly, Figure 9 The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiment of the present application. In actual applications, the electronic device may include Figure 9 More or fewer components than shown, or combinations of certain components.

[0116] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the music transcription method performed by the electronic device in any of the above embodiments are implemented.

[0117] In another exemplary embodiment, a computer program product is further provided, comprising a computer program, which, when executed, implements the steps of the music transcription method performed by the electronic device in any of the above embodiments.

[0118] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method section. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

[0119] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

Claims

1. A music transcription method, characterized in that: include: Get the target song audio; Performing lyric duration prediction on the target song audio to obtain a lyric word sequence and a duration corresponding to each lyric word in the lyric word sequence; Predicting the number of pitches of the target song audio to obtain the number of pitches corresponding to each word in the lyrics word sequence; Determining the pitch value corresponding to each lyric word in the lyric word sequence based on the duration and pitch number corresponding to each lyric word; A music score corresponding to the target song audio is generated based on the lyrics word sequence, the duration value and the pitch value corresponding to each lyrics word in the lyrics word sequence.

2. The music transcription method according to claim 1, wherein: Performing a lyrics duration prediction on the target song audio to obtain a lyrics word sequence and a duration corresponding to each word in the lyrics word sequence includes: Inputting the target song audio into an autoregressive lyrics and lyrics duration prediction model to predict the lyrics word sequence corresponding to the target song audio and the duration corresponding to each lyrics word in the lyrics word sequence; Accordingly, the pitch number prediction is performed on the target song audio to obtain the pitch number corresponding to each lyric word in the lyric word sequence, including: Inputting the target song audio into an autoregressive lyrics and pitch quantity prediction model to predict the number of pitches corresponding to each lyric word in the lyric word sequence; Accordingly, determining the pitch value corresponding to each lyric word in the lyric word sequence based on the duration and pitch number corresponding to each lyric word includes: Inputting the target song audio into a non-autoregressive pitch prediction model; In the non-autoregressive pitch prediction model, a coding module is used to extract a feature sequence at the target audio frame level, and a feature sequence corresponding to each lyric word is determined based on the duration corresponding to each lyric word; Using the first multi-classification module to predict the duration of each pitch corresponding to each lyric word based on the feature sequence and the number of pitches corresponding to each lyric word; Determine a characteristic sequence of each pitch corresponding to each word of the lyrics based on the duration of each pitch corresponding to each word of the lyrics; The second multi-classification module is used to predict the pitch value corresponding to each lyric word based on the feature sequence of each pitch corresponding to each lyric word.

3. The music transcription method according to claim 2, wherein: Also includes: Constructing a training data set; wherein the training data set includes multiple sets of training data, and the training data includes training audio segments and training music scores corresponding to the training audio segments; An autoregressive lyrics and lyrics duration prediction model, an autoregressive lyrics and pitch quantity prediction model, and a non-autoregressive pitch prediction model are trained based on the training data set.

4. The music transcription method according to claim 3, wherein: Training an autoregressive lyrics and lyrics duration prediction model based on the training data set includes: Determining a lyrics word sequence and a lyrics word duration sequence corresponding to the training audio segment based on the training music score corresponding to the training audio segment; Training an autoregressive lyrics and lyrics duration prediction model based on multiple groups of training audio clips, corresponding lyrics word sequences and lyrics word duration sequences; Accordingly, the autoregressive lyrics and pitch quantity prediction model is trained based on the training data set, including: Determining the lyrics word sequence and the number of lyrics word pitches corresponding to the training audio segment based on the training music score corresponding to the training audio segment; Training an autoregressive lyrics and pitch quantity prediction model based on multiple sets of training audio clips, corresponding lyrics word sequences and lyrics word pitch quantities; Accordingly, training a non-autoregressive pitch prediction model based on the training dataset includes: Determining a lyrics word sequence, a lyrics word duration sequence, a lyrics word pitch number, and a lyrics word pitch value sequence corresponding to the training audio segment based on the training music score corresponding to the training audio segment; A non-autoregressive pitch prediction model is trained based on multiple groups of training audio clips, corresponding lyrics word sequences, lyrics word time value sequences, lyrics word pitch quantities, and lyrics word pitch value sequences.

5. The music transcription method according to claim 1, wherein: The step of obtaining the target song audio includes: Obtain an audio file of a target song and identify a silent segment in the audio file; wherein the silent segment is an audio segment whose silence duration is greater than or equal to a preset value; Dividing the audio file into a plurality of target song audios according to the audio segments; Accordingly, after generating the music score corresponding to the target song audio based on the lyrics word sequence and the duration and pitch value corresponding to each lyrics word in the lyrics word sequence, the method further includes: The music scores corresponding to the multiple target song audios are spliced ​​into the music score corresponding to the target song.

6. The music transcription method according to claim 5, wherein: After the music scores corresponding to the plurality of target song audios are spliced ​​into the music score corresponding to the target song, the method further includes: The dry audio of the target song is obtained, and a cover of the target song is generated based on the dry audio, accompaniment audio and the lyrics sequence of the target song.

7. The music transcription method according to claim 5, wherein: After the music scores corresponding to the plurality of target song audio files are spliced ​​into the music score corresponding to the audio file, the method further includes: The dry audio and adapted lyrics sequence of the target song are obtained, and a cover of the target song is generated based on the dry audio, accompaniment audio and adapted lyrics sequence of the target song.

8. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the music transcription method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed, implements the steps of the music transcription method according to any one of claims 1 to 7.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the music transcription method according to any one of claims 1 to 7 when the computer program is executed.