Song annotation methods, computer equipment, and readable storage media

By identifying boundary audio frames and syllable boundary frames in the singing audio signal, and combining spectral and fundamental frequency features, pitch information is calculated using a note prediction model, which solves the problem of insufficient accuracy in singing annotation in existing technologies and achieves higher accuracy singing annotation.

CN119785827BActive Publication Date: 2025-10-28TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411905809.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-28
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

In existing methods for annotating vocal performance, the accuracy of pitch information calculation is low, resulting in insufficient precision in vocal performance annotation.

Method used

By acquiring boundary audio frames and syllable boundary frames from the singing audio signal, combining spectral features and fundamental frequency features, a pre-trained note prediction model is used to identify note boundaries, and pitch information is calculated based on the signal features within the boundary frames of syllables and notes, thus achieving more accurate singing annotation.

Benefits of technology

It improves the accuracy of vocal transcription, enabling more accurate identification and transcription of pitch information for syllables and notes in vocal performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785827B_ABST
    Figure CN119785827B_ABST
Patent Text Reader

Abstract

This application relates to a method for annotating vocal recordings, a computer device, and a computer-readable storage medium. The method includes: acquiring the vocal audio signal of a target song to be annotated; acquiring the boundary audio frames corresponding to each note in the vocal audio signal, and the boundary audio frames corresponding to each syllable in the vocal audio signal; obtaining the pitch information of the notes contained within the two boundary audio frames of each syllable based on the boundary audio frames of each syllable and the boundary audio frames of each note; and using the two boundary audio frames of each syllable, the boundary audio frames of the notes contained within the two boundary audio frames of each syllable, and the pitch information of the notes contained within the two boundary audio frames of each syllable as vocal annotation information for the vocal audio signal. This method can combine note boundaries to obtain the pitch information of each note contained in a syllable, and use the pitch information of the notes for vocal annotation, thereby improving the accuracy of vocal annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a method for annotating singing voices, a computer device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the development of audio processing technology, a technique for synthesizing singing voices has emerged. This technique can identify the pitch information of each syllable by annotating the singing voice, such as annotating the pitch of the syllables in the singing voice.

[0003] The aforementioned song annotation methods typically involve directly identifying the time range of each syllable and then averaging the fundamental frequency signal within that time range to obtain the pitch information for each syllable. However, this annotation method has low accuracy in calculating pitch information, resulting in low precision in current song annotation methods. Summary of the Invention

[0004] Therefore, it is necessary to provide a singing voice annotation method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the accuracy of singing voice annotation in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a method for annotating singing voices, including:

[0006] Obtain the audio signal of the target song to be labeled;

[0007] Obtain the two boundary audio frames corresponding to each note of the target song in the singing audio signal, and obtain the two boundary audio frames corresponding to each syllable of the target song in the singing audio signal.

[0008] Based on the boundary audio frames of each syllable and the boundary audio frames of each note, determine the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable, and obtain the pitch information of the notes contained within the range of the two boundary audio frames of each syllable based on the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable.

[0009] The pitch information of the two boundary audio frames of each syllable, the note boundary audio frames contained within the two boundary audio frames of each syllable, and the notes contained within the two boundary audio frames of each syllable are used as the singing annotation information of the singing audio signal.

[0010] In one embodiment, obtaining the two boundary audio frames corresponding to each note of the target song in the singing audio signal includes: obtaining the spectral features and fundamental frequency features of the singing audio signal; inputting the spectral features and the fundamental frequency features into a pre-trained note prediction model, and obtaining the two boundary audio frames corresponding to each note of the target song in the singing audio signal through the note prediction model.

[0011] In one embodiment, before inputting the spectral features and the fundamental frequency features into the pre-trained note prediction model, the method further includes: acquiring sample vocal audio signals of a sample song, and sample spectral features and sample fundamental frequency features corresponding to the sample vocal audio signals; inputting the sample spectral features and the sample fundamental frequency features into the note prediction model to be trained to obtain the predicted boundary audio frames corresponding to each note of the sample song in the sample vocal audio signal; acquiring the actual boundary audio frames corresponding to each note of the sample song in the sample vocal audio signal; and training the note prediction model based on the difference between the predicted boundary audio frames and the actual boundary audio frames to obtain the pre-trained note prediction model.

[0012] In one embodiment, obtaining the pitch information of the notes contained within the two boundary audio frames of each syllable based on the boundary audio frames of the notes contained within the two boundary audio frames of each syllable includes: determining the current syllable of the target song, wherein the current syllable is any one of the syllables of the target song; when there are boundary audio frames of notes between the two boundary audio frames corresponding to the current syllable, obtaining the note duration of the notes contained in the current syllable within the current syllable based on the two boundary audio frames corresponding to the current syllable and the existing boundary audio frames of the notes; and obtaining the pitch information of the notes contained in the current syllable based on the note duration of the notes contained in the current syllable within the current syllable.

[0013] In one embodiment, obtaining the pitch information of the notes contained in the current syllable based on the note duration of the notes contained in the current syllable within the current syllable includes: obtaining a first sub-fundamental frequency signal that matches the note duration of the notes contained in the current syllable from the fundamental frequency signal corresponding to the singing audio signal, based on the note duration of the notes contained in the current syllable within the current syllable; and obtaining the pitch information of the notes contained in the current syllable based on the signal values ​​of each of the first sub-fundamental frequency signals.

[0014] In one embodiment, after determining the current syllable of the target song, the method further includes: if there are no boundary audio frames for notes between the two boundary audio frames corresponding to the current syllable, obtaining the syllable duration of the current syllable based on the two boundary audio frames corresponding to the current syllable; obtaining a second sub-fundamental frequency signal matching the syllable duration from the fundamental frequency signal corresponding to the singing audio signal based on the syllable duration; obtaining the pitch information of the current syllable based on the signal value of the second sub-fundamental frequency signal, and using the pitch information of the current syllable as the pitch information of each note contained in the current syllable.

[0015] In one embodiment, obtaining the two boundary audio frames corresponding to each syllable of the target song in the singing audio signal includes: obtaining the lyrics text information of the target song, the lyrics text information including each syllable of the target song; and performing forced alignment processing on the singing audio signal and the lyrics text information to obtain the two boundary audio frames corresponding to each syllable of the target song in the singing audio signal.

[0016] Secondly, this application also provides a singing voice annotation device, comprising:

[0017] The audio signal acquisition module is used to acquire the audio signal of the target song that needs to be labeled.

[0018] The note boundary acquisition module is used to acquire two boundary audio frames corresponding to each note of the target song in the singing audio signal, and to acquire two boundary audio frames corresponding to each syllable of the target song in the singing audio signal.

[0019] The pitch information acquisition module is used to determine the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable based on the boundary audio frames of each syllable and the boundary audio frames of each note, and to obtain the pitch information of the notes contained within the range of the two boundary audio frames of each syllable based on the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable.

[0020] The vocal annotation acquisition module is used to take the two boundary audio frames of each syllable, the note boundary audio frames contained within the two boundary audio frames of each syllable, and the pitch information of the notes contained within the two boundary audio frames of each syllable as vocal annotation information of the vocal audio signal.

[0021] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in any embodiment of the first aspect.

[0022] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any embodiment of the first aspect.

[0023] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any embodiment of the first aspect.

[0024] The aforementioned singing voice annotation method, apparatus, computer equipment, computer-readable storage medium, and computer program product, when annotating singing voice audio signals, can first obtain the boundary audio frames contained in the singing voice audio signal. Then, by combining the boundary audio frames of each syllable and the boundary audio frames of each note contained in the singing voice audio signal, the pitch information of each note contained in each syllable can be obtained. Thus, the above pitch information, the boundary audio frames of each note contained in each syllable, and the boundary audio frames of each syllable are used as singing voice annotation information. In this way, the pitch information of each note contained in a syllable can be obtained by combining the boundary audio frames of the notes, and the singing voice annotation can be performed using the pitch information of the notes, thereby improving the accuracy of singing voice annotation. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating a singing voice annotation method in one embodiment;

[0027] Figure 2 This is a schematic diagram of the process of training a note prediction model in one embodiment;

[0028] Figure 3 This is a schematic diagram of the process for obtaining the pitch information of each note contained in each syllable in one embodiment;

[0029] Figure 4 This is a schematic diagram of training data annotation in one embodiment;

[0030] Figure 5 This is a schematic diagram of the training process of a note prediction model in one embodiment;

[0031] Figure 6 This is a schematic diagram of the model structure of a note prediction model in one embodiment;

[0032] Figure 7 This is a schematic diagram of the overall process for song annotation in one embodiment;

[0033] Figure 8 This is a structural block diagram of a singing voice annotation device in one embodiment;

[0034] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0035] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0036] In one embodiment, such as Figure 1 As shown, a method for annotating singing voices is provided. This embodiment illustrates the application of this method to a server. It is understood that this method can also be applied to a terminal, or to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0037] Step S101: Obtain the audio signal of the target song to be labeled;

[0038] Step S102: Obtain the two boundary audio frames corresponding to each note of the target song in the singing audio signal, and obtain the two boundary audio frames corresponding to each syllable of the target song in the singing audio signal.

[0039] The audio signal to be labeled refers to the audio signal of the target song for which the labeling needs to be completed. The boundary audio frame corresponding to the note in the audio signal can also be called the syllable boundary. It refers to the audio frame corresponding to the boundary of the note in the target song in the audio signal. For example, if the audio signal of a 2-second target song consists of two notes, where the duration of note 1 is 0s-1s and the duration of note 2 is 1s-2s, then the two boundary audio frames of note 1 can be the audio signal frame corresponding to 0s and the audio signal frame corresponding to 1s, respectively. The two boundary audio frames of note 2 can be the audio signal frame corresponding to 1s and the audio signal frame corresponding to 2s, respectively. Other audio frames of the audio signal are not considered as boundary audio frames of the notes.

[0040] The boundary audio frame corresponding to a syllable in a singing audio signal can also be called a syllable boundary. It refers to the audio frame corresponding to the boundary of the syllable in the target song in the singing audio signal. It can be used to distinguish different syllables in the singing audio signal. This syllable boundary can be obtained from the singing audio signal and the lyrics text corresponding to the singing audio signal.

[0041] Specifically, after obtaining the audio signal of the target song, the server can further extract the two boundary audio frames of each note and the two boundary audio frames of each syllable contained in the audio signal.

[0042] Step S103: Based on the boundary audio frames of each syllable and the boundary audio frames of each note, determine the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable, and obtain the pitch information of the notes contained within the range of the two boundary audio frames of each syllable based on the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable.

[0043] Since a syllable can contain multiple notes, such as a note boundary between two syllable boundaries, it means that the syllable can be composed of two notes. Therefore, after the server obtains the boundary audio frames of each syllable and the boundary audio frames of each note, it can determine the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable, thereby determining the notes contained within the range of the two boundary audio frames of each syllable, and further obtaining the pitch information of the notes contained within the range of the two boundary audio frames of each syllable.

[0044] Step S104: Use the two boundary audio frames of each syllable, the note boundary audio frames contained within the two boundary audio frames of each syllable, and the pitch information of the notes contained within the two boundary audio frames of each syllable as the singing annotation information of the singing audio signal.

[0045] Finally, the server can also use the pitch information of each note contained in each syllable, the boundary audio frames of each note contained in each syllable, and the boundary audio frames corresponding to each syllable as singing annotation information of the singing audio signal. In this way, pitch information annotation at the note level can be achieved, thereby improving the accuracy of singing annotation.

[0046] For example, if the boundary audio frames of a syllable are frames 10 and 30, and a note boundary audio frame appears in frame 20, it indicates that the syllable contains two notes. This syllable corresponds to two pitch information points: pitch information 1 corresponds to the pitch of note 1 in frames 10 to 20, and pitch information 2 corresponds to the pitch of note 2 in frames 20 to 30. However, if existing techniques are used for vocal annotation, since only syllables are identified, the syllable will only correspond to one overall pitch information point—that is, the pitch information of the syllables in frames 10 to 30 will be directly annotated. Therefore, the vocal annotation method provided in this embodiment has higher accuracy than existing vocal annotation methods.

[0047] In the above-mentioned singing voice annotation method, when annotating the singing voice audio signal, the boundary audio frames contained in the singing voice audio signal can be obtained first. Then, the boundary audio frames of each syllable and each note in the singing voice audio signal can be combined to obtain the pitch information of each note in each syllable. Thus, the above pitch information, the boundary audio frames of each note in each syllable, and the boundary audio frames of each syllable are used as singing voice annotation information. In this way, the pitch information of each note in each syllable can be obtained by combining the boundary audio frames of the notes, and the pitch information of the notes can be used for singing voice annotation, thereby improving the accuracy of singing voice annotation.

[0048] In one embodiment, step S102 may further include: acquiring the spectral features and fundamental frequency features of the singing audio signal; inputting the spectral features and fundamental frequency features into a pre-trained note prediction model, and obtaining the two boundary audio frames corresponding to each note of the target song in the singing audio signal through the note prediction model.

[0049] Spectral features refer to the feature information extracted from the spectral signal corresponding to the singing audio signal. This spectral signal can refer to the Mel spectrum signal corresponding to the singing audio signal, while fundamental frequency features refer to the feature information extracted from the fundamental frequency signal corresponding to the singing audio signal.

[0050] Specifically, when the server receives the audio signal of the song that needs to be labeled, it can also perform Mel spectrum extraction and fundamental frequency extraction on the audio signal to obtain the corresponding spectrum signal and fundamental frequency signal. Then, the server can perform feature extraction on the spectrum signal and fundamental frequency signal to obtain the spectral features corresponding to the spectrum signal and the fundamental frequency features corresponding to the fundamental frequency signal.

[0051] The pre-trained note prediction model refers to the trained note prediction model. This model can be used to identify the frame boundaries corresponding to each note in the singing audio signal, that is, the boundary audio frames of the notes. After the server obtains the spectral features and fundamental frequency features of the singing audio signal, it can input the above spectral features and fundamental frequency features into the pre-trained note prediction model, and the note prediction model can identify the two audio signal frames corresponding to the note boundaries in the singing audio signal.

[0052] In this embodiment, the note boundaries in the singing audio signal can be obtained by inputting the spectral features and fundamental frequency features of the singing audio signal into a pre-trained note prediction model, and then outputting the note prediction model. This method can improve the accuracy of note boundary acquisition.

[0053] In addition, such as Figure 2 As shown, before step S101, the following may also be included:

[0054] Step S201: Obtain the sample singing audio signal of the sample song, as well as the sample spectrum features and sample fundamental frequency features corresponding to the sample singing audio signal.

[0055] Among them, the sample singing audio signal refers to the singing audio signal of the sample song used to train the note prediction model, the sample spectral feature refers to the feature information extracted from the spectral signal corresponding to the sample singing audio signal, which can refer to the Mel spectrum signal corresponding to the sample singing audio signal, and the sample fundamental frequency feature refers to the feature information extracted from the fundamental frequency signal corresponding to the sample singing audio signal.

[0056] Specifically, when training the note prediction model, the server first collects sample singing audio signals and performs Mel spectrum extraction and fundamental frequency extraction on the sample singing audio signals to obtain the corresponding spectrum signal and fundamental frequency signal. Then, the server can perform feature extraction on the aforementioned spectrum signal and fundamental frequency signal to obtain the sample spectrum features corresponding to the spectrum signal and the sample fundamental frequency features corresponding to the fundamental frequency signal.

[0057] Step S202: Input the sample spectral features and sample fundamental frequency features into the note prediction model to be trained to obtain the prediction boundary audio frames corresponding to each note of the sample song in the sample song audio signal.

[0058] The note prediction model to be trained refers to the note prediction model that needs to be trained by a neural network model, which can be a convolutional neural network model. The predicted boundary audio frames of each note are the audio signal frames corresponding to the note boundaries in the sample singing audio signal output by the note prediction model. Specifically, after the server obtains the sample spectral features and sample fundamental frequency features, it can input the above sample spectral features and sample fundamental frequency features into the note prediction model to be trained, and the note prediction model outputs the predicted note frame boundaries of each note contained in the sample singing audio signal.

[0059] Step S203: Obtain the actual boundary audio frames corresponding to each note of the sample song in the sample song audio signal.

[0060] The actual boundary audio frame of each note refers to the actual boundary of the note contained in the sample singing audio signal. This note boundary can be obtained by the user annotating the sample singing audio signal. For example, it can be obtained by annotating the singing audio signal frame corresponding to the note boundary in the sample singing audio signal as the actual boundary audio frame of each note contained in the sample singing audio signal.

[0061] Step S204: Based on the difference between the predicted boundary audio frame and the actual boundary audio frame, train the note prediction model to obtain the pre-trained note prediction model.

[0062] After obtaining the predicted boundary audio frames and the actual boundary audio frames, the terminal can use the difference between the predicted boundary audio frames and the actual boundary audio frames to train the note prediction model, thereby obtaining a pre-trained note prediction model.

[0063] In this embodiment, the server can also use the sample singing audio signal and the actual note boundaries of the sample singing audio signal to train a note prediction model. This model can accurately predict note boundaries, thereby improving the accuracy of note boundary acquisition.

[0064] In one embodiment, such as Figure 3 As shown, step S103 may further include:

[0065] Step S301: Determine the current syllable of the target song. The current syllable is any one of the syllables in the target song.

[0066] The current syllable refers to any one of the multiple syllables contained in the singing audio signal. The two boundary audio frames corresponding to the current syllable are used to distinguish the current syllable from the previous syllable and to distinguish the current syllable from the next syllable, respectively.

[0067] In this embodiment, a syllable can be uniquely determined between two adjacent syllable boundaries. After obtaining the current syllable, the server can also obtain a boundary audio frame used to distinguish the current syllable from the previous syllable and a boundary audio frame used to distinguish the current syllable from the next syllable, which are respectively used as the two boundary audio frames corresponding to the current syllable.

[0068] Step S302: If there is a boundary audio frame for a note between the two boundary audio frames corresponding to the current syllable, the note duration of the note contained in the current syllable is obtained based on the two boundary audio frames corresponding to the current syllable and the boundary audio frames of the existing note.

[0069] If there is a note boundary audio frame between the two boundary audio frames corresponding to the current syllable, it indicates that the current syllable contains at least two notes. Then the server can obtain the note duration of the notes contained in the current syllable based on the two boundary audio frames and the existing note boundary audio frames.

[0070] For example, if the two boundary audio frames of the current syllable are the 10th and 30th frames of the vocal audio signal, and a note boundary appears in the 20th frame of the vocal audio signal, this indicates that the current syllable contains two notes. Note 1's duration within the current syllable is from frame 10 to frame 20, while note 2's duration is from frame 20 to frame 30. This method allows us to obtain the duration of each note contained within the current syllable within that syllable.

[0071] Step S303: Based on the duration of the notes contained in the current syllable within the current syllable, obtain the pitch information of the notes contained in the current syllable.

[0072] After determining the duration of each note in the current syllable, the pitch information of each note in the current syllable can be obtained based on the duration of the notes. For example, the pitch information of each note in the current syllable can be obtained by using the part of the singing audio signal that matches the duration of each note in the current syllable. That is, the pitch information of note 1 can be obtained by using the part of the singing audio signal from frame 10 to frame 20, and the pitch information of note 2 can be obtained by using the part of the singing audio signal from frame 20 to frame 30.

[0073] In this embodiment, if there is a boundary audio frame for a note between two boundary audio frames corresponding to a certain syllable, it indicates that the syllable contains at least two notes. Thus, the duration of each note in the current syllable can be determined by using the two boundary audio frames corresponding to the current syllable and the boundary audio frames for the existing notes, thereby determining the pitch information of the notes in the current syllable. This method can improve the accuracy of note marking.

[0074] Furthermore, step S303 may further include: obtaining a first sub-fundamental frequency signal that matches the note duration of the notes contained in the current syllable from the fundamental frequency signal corresponding to the singing audio signal, based on the note duration of the notes contained in the current syllable; and obtaining the pitch information of the notes contained in the current syllable based on the signal values ​​of each first sub-fundamental frequency signal.

[0075] In this embodiment, pitch information can be calculated from the fundamental frequency signal. The first sub-fundamental frequency signal refers to the portion of the fundamental frequency signal corresponding to the singing audio signal that matches the note duration of each note in the current syllable. Specifically, the server can extract the fundamental frequency from the singing audio signal to obtain the corresponding fundamental frequency signal, and extract the first sub-fundamental frequency signal matching the note duration from the fundamental frequency signal based on the note duration of each note in the current syllable. The pitch information of each note in the current syllable can be calculated using the signal value of the first sub-fundamental frequency signal, which can be the average or median of the signal value as the pitch information of the corresponding note.

[0076] For example, the server could take the portion of the baseband signal from frame 10 to frame 20 as the first sub-baseband signal corresponding to note 1, and take the average value of the signal corresponding to the first sub-baseband signal as the pitch information of note 1.

[0077] In this embodiment, the server can obtain a first sub-fundamental frequency signal that matches the note duration of each note from the fundamental frequency signal, and then use the signal value of the first sub-fundamental frequency signal to obtain the pitch information of each note. This method can improve the accuracy of obtaining the pitch information of the notes.

[0078] Additionally, after step S301, the method may further include: if there are no boundary audio frames for notes between the two boundary audio frames corresponding to the current syllable, obtaining the syllable duration of the current syllable based on the two boundary audio frames corresponding to the current syllable; obtaining a second sub-fundamental frequency signal that matches the syllable duration from the fundamental frequency signal corresponding to the singing audio signal based on the syllable duration; obtaining the pitch information of the current syllable based on the signal value of the second sub-fundamental frequency signal, and using the pitch information of the current syllable as the pitch information of each note contained in the current syllable.

[0079] If there are no boundary audio frames for notes between the two boundary audio frames corresponding to the current syllable, it indicates that the current syllable contains only one note and not multiple notes. In this case, only the pitch information of the current syllable needs to be obtained, and it can be used as the pitch information of each note contained in the current syllable.

[0080] Specifically, if the server determines that there is no note boundary audio frame between the start boundary audio frame and the end boundary audio frame corresponding to the current syllable, the server can directly calculate the syllable duration of the current syllable based on the singing audio signal frames corresponding to the start boundary audio frame and the end boundary audio frame.

[0081] The second sub-fundamental frequency signal refers to the portion of the fundamental frequency signal corresponding to the singing audio signal that matches the syllable duration of the current syllable. Specifically, the server can extract the fundamental frequency from the singing audio signal to obtain the corresponding fundamental frequency signal, and extract the second sub-fundamental frequency signal that matches the syllable duration from the fundamental frequency signal based on the syllable duration of the current syllable. The pitch information of the current syllable can be calculated using the signal value of the second sub-fundamental frequency signal, which can be the average or median of the signal value. This pitch information is then used as the pitch information of each note contained in the current syllable.

[0082] For example, if the starting boundary audio frame of the current syllable is the 30th frame of the singing audio signal, and the ending boundary audio frame is the 50th frame of the singing audio signal, and there are no boundary audio frames of notes between the 30th and 50th frames of the singing audio signal, then it indicates that the syllable is composed of the same note, and the syllable duration is from the 30th to the 50th frame. The server can then use the portion of the fundamental frequency signal from the 30th to the 50th frame as the second sub-fundamental frequency signal corresponding to the current syllable, and use the average value of the signal corresponding to the second sub-fundamental frequency signal as the pitch information of the current syllable.

[0083] In this embodiment, if there is no boundary audio frame of a note between the boundary of the first syllable and the boundary of the second syllable, the server can directly use the pitch information of the current syllable as the pitch information of each note contained in the current syllable. This method can improve the efficiency of obtaining the pitch information of each note contained in the current syllable.

[0084] In one embodiment, step S102 may further include: obtaining the lyrics text information of the target song, the lyrics text information including each syllable of the target song; performing forced alignment processing on the singing audio signal and the lyrics text information to obtain two boundary audio frames corresponding to each syllable of the target song in the singing audio signal.

[0085] Lyrics text information refers to the lyrics text corresponding to the singing audio signal. In this embodiment, syllable boundaries can be obtained by forcibly aligning the singing audio signal and the corresponding lyrics text. Specifically, after obtaining the singing audio signal, the server can also simultaneously obtain the lyrics text corresponding to the singing audio signal as lyrics text information. Then, the singing audio signal and lyrics text information can be forcibly aligned, thereby identifying each syllable contained in the singing audio signal and the boundaries between the two syllables corresponding to each syllable.

[0086] In this embodiment, the syllable boundaries can be obtained by the server through forced alignment of the singing audio signal and the corresponding lyrics text information. This method can ensure the accuracy of syllable boundary acquisition.

[0087] In one embodiment, an automated method for annotating singing voices is also provided. This method uses a regression model to predict note boundaries and combines this with syllable boundaries calculated through forced alignment, thereby achieving automated generation of highly accurate singing voice annotations. Compared to existing technologies that average the effective fundamental frequency over a syllable's time range as its pitch, the singing voice annotation method provided in this embodiment offers higher accuracy. This can be achieved through the following steps:

[0088] 1. Training data preparation:

[0089] (1) Audio and note annotation paired data: The smallest granularity is the note. For example, if there are multiple notes in a syllable, all notes will be annotated. Figure 4 The syllable "shen" corresponds to multiple pitches, namely 61 and 59, while the syllable "yi" corresponds to only one pitch, namely 61.

[0090] (2) The operation of fundamental frequency extraction and Mel spectrum extraction of audio signals can be performed using existing tools.

[0091] 2. Model Training:

[0092] Throughout the process, forced alignment can be achieved using existing tools, while note prediction requires training a note prediction model.

[0093] The training steps for the note prediction model can be as follows: Figure 5 As shown, the structure of the note prediction model used can be as follows: Figure 6As shown, the model's input consists of Mel features and fundamental frequency features. The fundamental frequency is normalized before being fed into the model. The model predicts the probability of whether a segment is a note boundary frame by frame. This model is a regression model. The model's annotations are derived from the annotations of whether a segment is a note boundary. Consonant and vowel boundaries are not labeled as note boundaries, while syllable boundaries and the intersection of note boundaries are used as model annotations.

[0094] 3. Process Reasoning:

[0095] The overall process of song annotation is as follows: Figure 7 As shown, syllable boundaries are provided by the results of forced alignment; if a transition occurs within a syllable, the note boundary is output by the note prediction model. Then, based on the syllable and note boundaries, the corresponding fundamental frequency sequence is taken. By averaging or taking the median of the fundamental frequency sequence, the note value can be obtained. For example, for the syllable "hao", the duration is from frame 10 to frame 30, and the note boundary appears in frame 20. Therefore, "hao" has two pitches: one is mean(pitch[10:20]), and the other is mean(pitch[20:30]). Through the above method, the singing annotation can be completed, including: syllables and syllable boundaries, notes and note boundaries.

[0096] In this embodiment, a regression model is used to predict the boundaries of notes, and the syllable boundaries calculated by forced alignment are combined to improve the accuracy of vocal annotation.

[0097] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0098] Based on the same inventive concept, this application also provides a singing voice annotation device for implementing the singing voice annotation method described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more singing voice annotation device embodiments provided below can be found in the limitations of the singing voice annotation method above, and will not be repeated here.

[0099] In one embodiment, such as Figure 8As shown, a singing voice annotation device is provided, including: an audio signal acquisition module 801, a note boundary acquisition module 802, a pitch information acquisition module 803, and a singing voice annotation acquisition module 804, wherein:

[0100] The audio signal acquisition module 801 is used to acquire the audio signal of the target song to be labeled.

[0101] The note boundary acquisition module 802 is used to acquire the two boundary audio frames corresponding to each note of the target song in the singing audio signal, and to acquire the two boundary audio frames corresponding to each syllable of the target song in the singing audio signal.

[0102] The pitch information acquisition module 803 is used to determine the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable based on the boundary audio frames of each syllable and the boundary audio frames of each note, and to obtain the pitch information of the notes contained within the range of the two boundary audio frames of each syllable based on the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable.

[0103] The vocal annotation acquisition module 804 is used to take the two boundary audio frames of each syllable, the note boundary audio frames contained within the two boundary audio frames of each syllable, and the pitch information of the notes contained within the two boundary audio frames of each syllable as vocal annotation information of the vocal audio signal.

[0104] Each module in the aforementioned song annotation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0105] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores vocal audio signal data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a vocal annotation method.

[0106] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0107] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0108] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0109] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0110] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0111] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0112] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0113] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for annotating singing voices, characterized in that, The method includes: Obtain the audio signal of the target song to be labeled; Obtain the two boundary audio frames corresponding to each note of the target song in the singing audio signal, and obtain the two boundary audio frames corresponding to each syllable of the target song in the singing audio signal. Based on the boundary audio frames of each syllable and the boundary audio frames of each note, determine the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable, and obtain the pitch information of the notes contained within the range of the two boundary audio frames of each syllable based on the boundary audio frames of the notes contained within the range of the two boundary audio frames of each syllable. The pitch information of the two boundary audio frames of each syllable, the note boundary audio frames contained within the two boundary audio frames of each syllable, and the notes contained within the two boundary audio frames of each syllable are used as the singing annotation information of the singing audio signal.

2. The method according to claim 1, characterized in that, The step of obtaining the two boundary audio frames corresponding to each note of the target song in the singing audio signal includes: Obtain the spectral characteristics and fundamental frequency characteristics of the singing audio signal; The spectral features and the fundamental frequency features are input into a pre-trained note prediction model, and the note prediction model is used to obtain the two boundary audio frames corresponding to each note of the target song in the singing audio signal.

3. The method according to claim 2, characterized in that, Before inputting the spectral features and the fundamental frequency features into the pre-trained note prediction model, the method further includes: Obtain the sample vocal audio signal of the sample song, as well as the sample spectral features and sample fundamental frequency features corresponding to the sample vocal audio signal; The sample spectral features and the sample fundamental frequency features are input into the note prediction model to be trained to obtain the prediction boundary audio frames corresponding to each note of the sample song in the sample singing audio signal. Obtain the actual boundary audio frames corresponding to each note of the sample song in the sample singing audio signal; The note prediction model is trained based on the difference between the predicted boundary audio frame and the actual boundary audio frame to obtain the pre-trained note prediction model.

4. The method according to claim 1, characterized in that, The process of obtaining pitch information of the notes contained within the two boundary audio frames of each syllable, based on the boundary audio frames of the notes contained within the two boundary audio frames of each syllable, includes: Determine the current syllable of the target song, wherein the current syllable is any one of the syllables in the target song; When there is a boundary audio frame for a note between the two boundary audio frames corresponding to the current syllable, the note duration of the note contained in the current syllable within the current syllable is obtained based on the two boundary audio frames corresponding to the current syllable and the boundary audio frame for the existing note. The pitch information of the notes contained in the current syllable is obtained based on the duration of the notes in the current syllable.

5. The method according to claim 4, characterized in that, The step of obtaining the pitch information of the notes contained in the current syllable based on the duration of the notes contained in the current syllable within the current syllable includes: Based on the note duration of the notes contained in the current syllable within the current syllable, a first sub-fundamental frequency signal matching the note duration of the notes contained in the current syllable is obtained from the fundamental frequency signal corresponding to the singing audio signal; Based on the signal values ​​of each of the first sub-fundamental frequency signals, the pitch information of the notes contained in the current syllable is obtained.

6. The method according to claim 4, characterized in that, After determining the current syllable of the target song, the method further includes: If there are no boundary audio frames of a note between the two boundary audio frames corresponding to the current syllable, the syllable duration of the current syllable is obtained based on the two boundary audio frames corresponding to the current syllable. Based on the syllable duration, a second sub-fundamental frequency signal matching the syllable duration is obtained from the fundamental frequency signal corresponding to the singing audio signal; Based on the signal value of the second sub-fundamental frequency signal, the pitch information of the current syllable is obtained, and the pitch information of the current syllable is used as the pitch information of each note contained in the current syllable.

7. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining the two boundary audio frames corresponding to each syllable of the target song in the singing audio signal includes: Obtain the lyrics text information of the target song, wherein the lyrics text information includes each syllable of the target song; The singing audio signal and the lyrics text information are forcibly aligned to obtain two boundary audio frames corresponding to each syllable of the target song in the singing audio signal.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Viewing and singing pitch detection method, system and device based on target detection and medium

    CN115206339A

  • Speech synthesis method and system based on adaptive attention mechanism

    CN116030786A