Fundamental frequency sequence recognition model training and fundamental frequency sequence recognition method, device and product

By acquiring the dry audio and spectrum of a user's song, verifying the fundamental frequency sequence based on standard pitch, determining training labels, and training a fundamental frequency sequence recognition model, the problem of low fundamental frequency recognition accuracy in existing technologies is solved, achieving higher recognition accuracy and data fit.

CN115510911BActive Publication Date: 2026-04-14TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
Filing Date
2022-09-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing fundamental frequency identification models are difficult to accurately identify fundamental frequencies in audio in practical applications, resulting in low accuracy.

Method used

By acquiring the dry audio and spectrum of a user's song, fundamental frequency sequence recognition is performed. The fundamental frequency sequence is verified based on the standard pitch of the song, and training labels are determined. The fundamental frequency sequence recognition model is trained using the spectrum of the dry audio and the training labels, thereby improving the model's training data to better match actual application data and calibration accuracy.

Benefits of technology

This improves the accuracy of the fundamental frequency sequence recognition model in practical applications, making the model more closely match the actual input data and enhancing the accuracy and reliability of the training data, thereby improving the accuracy of fundamental frequency sequence recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115510911B_ABST
    Figure CN115510911B_ABST
Patent Text Reader

Abstract

The application relates to a base frequency sequence identification model training method, a base frequency sequence identification method, a computer device and a computer program product, which can improve base frequency identification accuracy. The method comprises the following steps: acquiring dry audio of a song sung by a user and a spectrum of the dry audio, performing base frequency sequence identification on the dry audio to obtain a base frequency sequence of the dry audio; determining a standard pitch of a song content of the song, and verifying the base frequency sequence of the dry audio based on the standard pitch; in the case of passing the verification, determining the base frequency sequence as a training label of the spectrum of the dry audio; training a base frequency sequence identification model to be trained according to the spectrum of the dry audio and the training label of the spectrum, to obtain a trained base frequency sequence identification model; and the base frequency sequence identification model is used for predicting a base frequency sequence of an input spectrum.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a training method for a fundamental frequency sequence recognition model, a fundamental frequency sequence recognition method, a computer device, and a computer program product. Background Technology

[0002] With the development of audio processing technology, fundamental frequency extraction technology has been increasingly widely used. For example, after identifying the fundamental frequency in the audio, various processing methods such as pitch shifting and mixing are performed on the audio based on the fundamental frequency.

[0003] In related technologies, a fundamental frequency recognition model can be trained using deep learning methods. During the training process, a large amount of artificially synthesized audio data can be acquired to train the fundamental frequency recognition model, so that the final trained model can recognize the fundamental frequency in the audio.

[0004] However, the recognition model obtained through the above method is difficult to accurately identify the fundamental frequency in practical applications, resulting in poor recognition accuracy. Summary of the Invention

[0005] Therefore, it is necessary to provide a training method for a fundamental frequency sequence recognition model, a fundamental frequency sequence recognition method, a computer device, and a computer program product to address the above-mentioned technical problems.

[0006] Firstly, this application provides a method for training a fundamental frequency sequence recognition model. The method includes:

[0007] Obtain the dry audio of the user singing a song and the spectrum of the dry audio, and perform fundamental frequency sequence identification on the dry audio to obtain the fundamental frequency sequence of the dry audio;

[0008] Determine the standard pitch of the song content, and verify the dry audio fundamental frequency sequence based on the standard pitch;

[0009] If the verification passes, the fundamental frequency sequence is determined as the training label for the spectrum of the dry audio.

[0010] Based on the spectrum of the dry audio and the training labels of the spectrum, the fundamental frequency sequence recognition model to be trained is trained to obtain the trained fundamental frequency sequence recognition model; the fundamental frequency sequence recognition model is used to predict the fundamental frequency sequence of the input spectrum.

[0011] Secondly, this application also provides a fundamental frequency sequence identification method. The method includes:

[0012] Obtain the spectrum of the target dry audio;

[0013] The spectrum of the target dry audio is input into a pre-trained fundamental frequency sequence recognition model, and the fundamental frequency sequence output by the fundamental frequency sequence recognition model is determined as the fundamental frequency sequence of the target dry audio.

[0014] The fundamental frequency sequence identification model is trained according to the training method of the fundamental frequency sequence identification model described in any of the above claims.

[0015] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0016] Obtain the dry audio of the user singing a song and the spectrum of the dry audio, and perform fundamental frequency sequence identification on the dry audio to obtain the fundamental frequency sequence of the dry audio;

[0017] Determine the standard pitch of the song content, and verify the dry audio fundamental frequency sequence based on the standard pitch;

[0018] If the verification passes, the fundamental frequency sequence is determined as the training label for the spectrum of the dry audio.

[0019] Based on the spectrum of the dry audio and the training labels of the spectrum, the fundamental frequency sequence recognition model to be trained is trained to obtain the trained fundamental frequency sequence recognition model; the fundamental frequency sequence recognition model is used to predict the fundamental frequency sequence of the input spectrum.

[0020] Fourthly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0021] Obtain the spectrum of the target dry audio;

[0022] The spectrum of the target dry audio is input into a pre-trained fundamental frequency sequence recognition model, and the fundamental frequency sequence output by the fundamental frequency sequence recognition model is determined as the fundamental frequency sequence of the target dry audio.

[0023] The fundamental frequency sequence identification model is trained according to the training method of the fundamental frequency sequence identification model described in any of the above claims.

[0024] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0025] Obtain the dry audio of the user singing a song and the spectrum of the dry audio, and perform fundamental frequency sequence identification on the dry audio to obtain the fundamental frequency sequence of the dry audio;

[0026] Determine the standard pitch of the song content, and verify the dry audio fundamental frequency sequence based on the standard pitch;

[0027] If the verification passes, the fundamental frequency sequence is determined as the training label for the spectrum of the dry audio.

[0028] Based on the spectrum of the dry audio and the training labels of the spectrum, the fundamental frequency sequence recognition model to be trained is trained to obtain the trained fundamental frequency sequence recognition model; the fundamental frequency sequence recognition model is used to predict the fundamental frequency sequence of the input spectrum.

[0029] Sixthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0030] Obtain the spectrum of the target dry audio;

[0031] The spectrum of the target dry audio is input into a pre-trained fundamental frequency sequence recognition model, and the fundamental frequency sequence output by the fundamental frequency sequence recognition model is determined as the fundamental frequency sequence of the target dry audio.

[0032] The fundamental frequency sequence identification model is trained according to the training method of the fundamental frequency sequence identification model described in any of the above claims.

[0033] The aforementioned training method, fundamental frequency sequence recognition method, computer equipment, and computer program product for the fundamental frequency sequence recognition model can acquire the dry audio and spectrum of a user singing a song, perform fundamental frequency sequence recognition on the dry audio to obtain the fundamental frequency sequence of the dry audio, and then determine the standard pitch of the song content. Based on the standard pitch, the fundamental frequency sequence of the dry audio is verified. If the verification passes, the fundamental frequency sequence is used to locate the training label of the spectrum of the dry audio. Based on the spectrum of the dry audio and the training label of the spectrum, the fundamental frequency sequence recognition model to be trained is trained to obtain a trained fundamental frequency sequence recognition model. This fundamental frequency sequence recognition model can be used to predict the fundamental frequency sequence corresponding to the input spectrum. The solution of this application, on the one hand, by using the dry audio collected when the user sings a song to train the fundamental frequency sequence recognition model, makes the training data of the model more closely match the input data in the actual application process. On the other hand, by calibrating the extracted fundamental frequency using the standard pitch of the song and obtaining the training label, the accuracy and reliability of the training data can be improved, thereby improving the accuracy of the fundamental frequency sequence recognition model in recognizing the fundamental frequency sequence. Attached Figure Description

[0034] Figure 1 This diagram illustrates the training method for the fundamental frequency sequence recognition model and its application environment in one embodiment.

[0035] Figure 2 This is a flowchart illustrating a fundamental frequency sequence identification model method in one embodiment;

[0036] Figure 3a This is a time-domain waveform of an audio frame in one embodiment;

[0037] Figure 3b This is a spectrogram of an audio frame in one embodiment;

[0038] Figure 4 This is a flowchart illustrating a fundamental frequency sequence verification process in one embodiment.

[0039] Figure 5a This is a schematic diagram comparing a fundamental frequency sequence predicted by a fundamental frequency sequence recognition model with a training label in one embodiment;

[0040] Figure 5b This is a schematic diagram comparing another fundamental frequency sequence predicted by a fundamental frequency sequence recognition model with the training label in one embodiment;

[0041] Figure 6 This is a flowchart illustrating a fundamental frequency sequence identification method in one embodiment;

[0042] Figure 7 This is an internal structural diagram of a computer device in one embodiment;

[0043] Figure 8 This is an internal structural diagram of another computer device in one embodiment. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0045] Figure 1 This diagram illustrates the training method for a fundamental frequency sequence recognition model provided in one embodiment, as well as the application environment of the fundamental frequency sequence recognition method. For example... Figure 1As shown, the application environment can include a terminal and a server. The server can train the fundamental frequency sequence recognition model to obtain a trained model. After obtaining the trained model, the server can deploy it in an audio processing application, which can be installed on the terminal. When the terminal obtains the audio with the fundamental frequency sequence to be recognized, the user can issue a fundamental frequency sequence recognition command through corresponding operations. The terminal can receive the command, obtain the audio spectrum, and perform fundamental frequency recognition based on the spectrum using the pre-deployed fundamental frequency sequence recognition model to obtain the fundamental frequency sequence of the audio.

[0046] It is understood that the above application scenario is merely an example and does not constitute a limitation on the training method and fundamental frequency sequence recognition method of the fundamental frequency sequence recognition model provided in the embodiments of this application. For example, the fundamental frequency sequence recognition model can be trained through a terminal, or the fundamental frequency sequence recognition model can be deployed on a server. The server can receive the target dry audio of the fundamental frequency sequence to be recognized sent by the terminal, perform fundamental frequency sequence recognition on the target dry audio to obtain the fundamental frequency sequence, and return it to the terminal.

[0047] The server can be a standalone physical server or a server cluster consisting of multiple physical servers. It can also be a cloud server providing basic cloud computing services such as cloud servers, cloud databases, or cloud storage. The terminal can be a smartphone, tablet, laptop, desktop computer, or smartwatch, but is not limited to these. The terminal and computer equipment can connect via a network, which is not limited herein.

[0048] In one embodiment, such as Figure 2 As shown, a training method for a fundamental frequency sequence recognition model is provided, which is then applied to... Figure 1 Taking the server in the example, the following steps are included:

[0049] S201, acquire the dry audio of the user's song and the spectrum of the dry audio, and perform fundamental frequency sequence recognition on the dry audio to obtain the fundamental frequency sequence of the dry audio.

[0050] The dry audio can be non-synthesized audio obtained by collecting the human voice signal emitted by the user during the singing of a song. The dry audio can contain only human voice signal, and the human voice signal is the user's own human voice signal, that is, it does not contain the human voice signal in the original song.

[0051] In practice, multiple dry audio files can be pre-acquired and stored. For example, users can record and upload songs through audio / video applications. The uploaded audio content can include dry audio recordings of a specific song, which can then be used as the dry audio file in this step. Alternatively, high-quality dry audio files with high sampling rates, such as those with a sampling rate of 44.1kHz or higher, can be obtained from open-source datasets of dry audio files. In this step, multiple dry audio files can be obtained from a database, for example, multiple dry audio files can be randomly selected from the database; or, multiple dry audio files that meet preset filtering conditions can be obtained, such as filtering based on audio quality, singer's gender, or audio duration, or one or more other conditions.

[0052] In addition, the spectrum of the dry audio signal can be obtained. The spectrum is a representation of a time-domain signal in the frequency domain; it records the relationship between the time and frequency of a sound signal. The spectrum can be stored in the form of a spectrogram. Figure 3a and Figure 3b The time-domain waveform and spectrogram of an audio signal at time 3:46.386 are shown respectively. Of course, the spectrum of the dry audio signal can also be stored in other data formats, such as tabular data. In some optional embodiments, to obtain a suitable amount of sound characteristics, the Mel spectrum of the dry audio signal can be obtained. For example, the original spectrum of the dry audio signal can be processed through Mel-scale filter banks to convert it into a Mel spectrum, which better matches the human auditory characteristics of being more sensitive to low-frequency sounds.

[0053] After acquiring multiple dry audio samples, fundamental frequency sequence identification can be performed on the dry audio samples to obtain their fundamental frequency sequence. Specifically, existing fundamental frequency identification methods can be used to identify the fundamental frequency of the dry audio samples. This embodiment does not limit the method of identifying the fundamental frequency sequence of the dry audio samples in this step; those skilled in the art can choose the fundamental frequency identification method according to the actual situation.

[0054] S202, determine the standard pitch of the song content, and verify the fundamental frequency sequence of the dry audio based on the standard pitch.

[0055] Among them, the song content can be the content sung by the user according to the melody of the song, such as lyrics sung according to the melody of the song; the standard pitch can be the pitch that the corresponding melody should have when the song content is sung.

[0056] In practical applications, after identifying the fundamental frequency sequence corresponding to the dry audio of a user singing a song, the standard pitch of the song content can be further determined, and the fundamental frequency sequence of the identified dry audio can be verified based on the standard pitch to determine whether the currently identified fundamental frequency sequence is accurate and reliable.

[0057] Specifically, the pitch of a sound can be determined based on its fundamental frequency. Sound, as a mechanical wave, has a pitch determined by its frequency. Pitch is positively correlated with frequency; the higher the frequency, the higher the pitch, and vice versa. After identifying the fundamental frequency sequence of the dry audio, the pitch of the user singing a song can be determined based on this sequence. The pitch of the sung song often has a preset standard (e.g., the lyric 'A' should be sung at the preset pitch 'a'). Therefore, in this step, the standard pitch of the song content can be compared with the identified fundamental frequency sequence to determine if they match, thus verifying the accuracy of the identified fundamental frequency sequence.

[0058] S203, if the verification passes, the fundamental frequency sequence is determined as the training label for the spectrum of the dry audio.

[0059] After verification, if the fundamental frequency sequences of several audio frequencies pass the verification, the training labels of the spectrum of the dry audio frequencies can be obtained based on the fundamental frequency sequences. In other words, the identified fundamental frequency sequences can be used as reference standards for the fundamental frequency sequences output by subsequent models.

[0060] S204. Based on the spectrum of the dry audio and the training labels of the spectrum, train the fundamental frequency sequence recognition model to be trained to obtain the trained fundamental frequency sequence recognition model; the fundamental frequency sequence recognition model is used to predict the fundamental frequency sequence of the input spectrum.

[0061] Specifically, the fundamental frequency sequence recognition model can predict the fundamental frequency sequence of the audio corresponding to the input spectrum. That is, if only the audio spectrum is input into the fundamental frequency sequence recognition model, the model can output the fundamental frequency sequence.

[0062] After obtaining the spectrum of the dry audio and the fundamental frequency sequence used as training labels, the spectrum of the dry audio can be used as input to train the fundamental frequency sequence recognition model. Specifically, the spectrum of the dry audio can be input into the fundamental frequency sequence recognition model, which will predict the fundamental frequency sequence of the spectrum. Then, the model parameters of the fundamental frequency sequence recognition model can be adjusted according to the difference between the fundamental frequency sequence predicted by the model and the fundamental frequency sequence used as training labels, until the training termination condition is met, and a trained fundamental frequency sequence recognition model can be obtained.

[0063] The training method for the aforementioned fundamental frequency sequence recognition model can acquire the dry audio and spectrum of a user singing a song, perform fundamental frequency sequence recognition on the dry audio to obtain the fundamental frequency sequence of the dry audio, and then determine the standard pitch of the song content. Based on the standard pitch, the fundamental frequency sequence of the dry audio is verified. If the verification passes, the fundamental frequency sequence is determined as the training label for the spectrum of the dry audio. Based on the spectrum of the dry audio and the training label, the fundamental frequency sequence recognition model to be trained is trained to obtain a trained fundamental frequency sequence recognition model. This model can be used to predict the fundamental frequency sequence corresponding to the input spectrum. The solution in this application, on the one hand, by using the dry audio collected when the user sings a song to train the fundamental frequency sequence recognition model, makes the training data of the model more closely match the input data in the actual application process. On the other hand, by calibrating the extracted fundamental frequency using the standard pitch of the song and obtaining the training label, the accuracy and reliability of the training data can be improved, thereby improving the accuracy of the fundamental frequency sequence recognition model in recognizing the fundamental frequency sequence.

[0064] In one embodiment, S202, determining the standard pitch of the song's content and verifying the fundamental frequency sequence of the dry audio based on the standard pitch, may include the following steps:

[0065] S301, Obtain the sheet music of the song, and determine the standard pitch of multiple phonemes in the song content based on the sheet music.

[0066] Among them, a phoneme is the smallest unit of speech that is divided according to the natural attributes of speech. It is determined by the articulation action in a syllable, and one action constitutes one phoneme.

[0067] In practice, the sheet music corresponding to the song sung by the user can be obtained. For example, if there are different versions of the same song, the sheet music corresponding to the version chosen by the user can be obtained. The song content sung by the user (such as lyrics) can be broken down into multiple phonemes. After obtaining the sheet music, the standard pitch of each phoneme contained in the song content can be determined based on the sheet music.

[0068] S302, Based on the fundamental frequency sequence of the dry audio, determine the fundamental frequency of each audio frame in the dry audio.

[0069] Specifically, the fundamental frequency sequence of the dry audio can be obtained by identifying the fundamental frequency of each audio frame. In other words, after obtaining the dry audio, it can be divided into multiple audio frames, and the fundamental frequency of each audio frame can be obtained. This results in a fundamental frequency sequence composed of the fundamental frequencies of multiple audio frames, which serves as the fundamental frequency sequence of the dry audio. In this step, after obtaining the fundamental frequency sequence of the dry audio, the fundamental frequency of each audio frame in the dry audio can be determined.

[0070] S303 verifies the fundamental frequency of each audio frame in dry audio based on the standard pitch of multiple phonemes.

[0071] After obtaining the standard pitch of multiple phonemes in the song content, the fundamental frequency of multiple audio frames in the dry audio can be verified based on the standard pitch of multiple phonemes, thereby obtaining the verification result of the entire fundamental frequency sequence of the dry audio.

[0072] In this embodiment, the fundamental frequency verification at the audio frame level can be performed by using the standard pitch of each phoneme in the song content, which effectively improves the reliability of the fundamental frequency sequence verification results and provides a foundation for obtaining accurate and reliable fundamental frequency sequences as training labels in the future.

[0073] In one embodiment, S303, based on the standard pitch of multiple phonemes, verifies the fundamental frequency of each audio frame in the dry audio and may include the following steps:

[0074] Determine the phoneme associated with each audio frame in the dry audio; for each audio frame, obtain the pitch difference between the standard pitch of the phoneme associated with the audio frame and the pitch corresponding to the fundamental frequency of the audio frame; if the pitch difference of each audio frame in the dry audio is less than the difference threshold, determine that the fundamental frequency sequence of the dry audio passes the verification.

[0075] In practical applications, after the dry audio is segmented, the phonemes contained in each audio frame can be identified, and the phonemes contained in the audio frame can be determined as the phonemes associated with that audio frame. Alternatively, the phonemes in the song content corresponding to the time position can be determined based on the time position of the audio frame in the entire dry audio, and these phonemes can be used as the phonemes associated with that audio frame.

[0076] After obtaining the phonemes associated with each audio frame in the dry audio, for each audio frame, the standard pitch of the phoneme associated with that audio frame and the pitch corresponding to the fundamental frequency of that audio frame can be determined, and the pitch difference between the standard pitch of the phoneme and the pitch corresponding to the fundamental frequency can be obtained. Then, it can be determined whether the corresponding pitch difference of that audio frame is less than the difference threshold, such as whether it is less than 12 semitones (i.e., one octave).

[0077] like Figure 4 As shown, if it is determined that the pitch difference of each audio frame in the dry audio is less than the difference threshold, it can be determined that the current identified fundamental frequency sequence has passed the verification. If it is determined that the pitch difference of one or more audio frames is greater than the difference threshold, it can be determined that the current identified fundamental frequency sequence may be abnormal. The fundamental frequency sequence of the current identified dry audio can be sent to relevant staff for manual verification.

[0078] In this embodiment, by comparing the standard pitch of the phoneme associated with the audio frame with the pitch corresponding to the fundamental frequency of the audio frame, the corresponding pitch difference is determined, which can accurately and quickly identify whether there is a fundamental frequency recognition abnormality.

[0079] In one embodiment, determining the training labels for the spectrum of the fundamental frequency sequence as the dry audio in S203 may include the following steps:

[0080] The dry audio is segmented based on predetermined audio segmentation points to obtain multiple audio segments. For each audio segment, the fundamental frequency sequence associated with the audio segment in the fundamental frequency sequence of the dry audio is used as the training label of the audio segment's spectrum.

[0081] The audio segmentation point is determined based on the location of silence and / or the location of breathing sounds in the dry audio.

[0082] In practical applications, complete dry audio recordings may be quite long, making model training and fundamental frequency prediction difficult. Furthermore, the audio content in dry audio may be discontinuous, potentially leading to poor fundamental frequency continuity in the predicted fundamental frequency sequence, thus affecting model training results. Therefore, breath detection can be performed on the dry audio to identify locations where breathing sounds occur and / or where there is silence. These identified locations are then designated as audio segmentation points, and the dry audio is segmented based on these points to obtain multiple audio segments.

[0083] After acquiring multiple audio segments, for each audio segment, the fundamental frequency sequence associated with that audio segment can be determined from the fundamental frequency sequence of the dry audio. For example, based on the time segment corresponding to the segmented audio segment, the fundamental frequency sequence corresponding to that time segment can be obtained from the entire fundamental frequency sequence of the dry audio to obtain the fundamental frequency sequence associated with that audio segment, and this can be used as the training label for the spectrum of the audio segment.

[0084] In this embodiment, by determining the audio segmentation point based on the location of silence and / or the location of breathing sounds in the dry audio, and segmenting the dry audio, multiple audio segments with continuous content can be obtained, ensuring that the fundamental frequency sequence predicted by the subsequently trained model can have good fundamental frequency continuity.

[0085] In one embodiment, step S204, training the fundamental frequency sequence recognition model to be trained based on the spectrum of the dry audio and the training labels of the spectrum, to obtain a trained fundamental frequency sequence recognition model, may include the following steps:

[0086] The spectrum of the audio segment is input into the fundamental frequency sequence recognition model to be trained, and the predicted fundamental frequency sequence is obtained based on the fundamental frequency sequence output by the fundamental frequency sequence recognition model. The model parameters of the fundamental frequency sequence recognition model are adjusted according to the difference between the predicted fundamental frequency sequence and the fundamental frequency sequence associated with the audio segment until the training termination condition is met, and the trained fundamental frequency sequence recognition model is obtained.

[0087] In practical applications, after segmenting the dry audio and obtaining multiple audio segments, the spectrum of the audio segments can be obtained. The spectrum of the audio segments is then input into the fundamental frequency sequence recognition model to be trained. Based on the fundamental frequency sequence output by the model, the predicted fundamental frequency sequence is obtained.

[0088] In an optional embodiment, the fundamental frequency sequence recognition model to be trained can be a convolutional neural network, which can consist of three convolutional layers and one linear layer. Each convolutional layer can include a convolutional kernel (such as a one-dimensional convolutional kernel conv1D) and a normalization layer (such as a layer norm layer). After the input dry audio signal is processed through multiple convolutional layers for feature extraction, the extracted features are passed through the linear layer, outputting a fundamental frequency mapping value. For example, an activation function (such as the sigmoid activation function or other equivalent activation functions) can be used to process the extracted features to obtain a value between 0 and 1, representing the position of the fundamental frequency within its frequency range. For example, if the sigmoid activation function is used, the function S(x) can be as follows:

[0089]

[0090] Here, x represents the extracted features, and the output value of the function S(x) is in the range of (0,1).

[0091] After obtaining the predicted fundamental frequency sequence, the difference between the predicted fundamental frequency sequence and the fundamental frequency sequence associated with the audio segment can be determined. The model parameters of the fundamental frequency sequence recognition model are then adjusted based on this difference, and the model parameters are iteratively updated until the training termination condition is met, resulting in a trained fundamental frequency sequence recognition model. When adjusting the model parameters, the model loss value can be determined based on the difference between the two sequences, and the model parameters can be adjusted based on this loss value. The model loss value can be determined based on a loss function; for example, the loss function can be loss1 or loss2 as shown below:

[0092]

[0093]

[0094] Among them, y true For the fundamental frequency sequence associated with the audio segment, y predThe predicted fundamental frequency sequence output by the fundamental frequency sequence recognition model is denoted by n, where n is the number of audio segment samples. The loss function loss2 can amplify the difference between the predicted spectrum sequence and the fundamental frequency sequence used as training labels. During training, the appropriate loss function can be selected according to the training requirements.

[0095] After multiple iterations, the difference between the predicted fundamental frequency sequence and the fundamental frequency sequence used as the training label in the fundamental frequency sequence recognition model will gradually decrease. Figure 5a and Figure 5b As shown, Figure 5a The predicted fundamental frequency sequence is obtained when the model is first trained. It deviates from the fundamental frequency sequence used as the training label. When the model training ends, the predicted fundamental frequency sequence obtained by the model based on the fundamental frequency sequence is basically the same as the fundamental frequency sequence of the training label.

[0096] In this embodiment, by training the fundamental frequency sequence recognition model with the spectrum corresponding to continuous audio segments of audio content and their corresponding training labels, the continuity of the fundamental frequency sequence predicted by the model can be improved.

[0097] In one embodiment, obtaining the predicted fundamental frequency sequence based on the fundamental frequency sequence output by the fundamental frequency sequence identification model may include the following steps:

[0098] Obtain the fundamental frequency sequence output by the fundamental frequency sequence recognition model and determine the sound frequency range of the sound source of the dry audio; adjust the fundamental frequencies in the fundamental frequency sequence output by the fundamental frequency sequence recognition model that are outside the sound frequency range to preset values, and use the adjusted fundamental frequency sequence as the predicted fundamental frequency sequence.

[0099] In specific implementation, after the fundamental frequency sequence recognition model outputs the corresponding fundamental frequency sequence based on the spectrum of the input audio segment, the sound frequency range of the sound source of the dry audio can be determined. For example, since dry audio is obtained by collecting human voice signals, the sound frequency range of the sound source can be determined based on the sound frequency range of human voice. Alternatively, the gender of the user who speaks in the dry audio can be determined, and the corresponding sound frequency range can be determined based on the gender of the user who speaks.

[0100] After obtaining the sound frequency range, each fundamental frequency in the output fundamental frequency sequence can be compared with the sound frequency range. If the frequency value of the fundamental frequency exceeds the sound frequency range, it can be determined that the currently identified fundamental frequency is abnormal. Therefore, the fundamental frequencies in the fundamental frequency sequence that exceed the sound frequency range can be adjusted to a preset value, such as setting them to 0, and the adjusted fundamental frequency sequence can be used as the predicted fundamental frequency sequence.

[0101] In this embodiment, by adjusting the fundamental frequencies that are outside the sound frequency range in the fundamental frequency sequence output by the fundamental frequency sequence recognition model to a preset value, and using the adjusted fundamental frequency sequence as the predicted fundamental frequency sequence, it can be ensured that the frequency value of the fundamental frequency in the final predicted fundamental frequency sequence matches the frequency range of the user's voice, thereby enhancing the stability of the model output.

[0102] In one embodiment, segmenting the dry audio based on predetermined audio segmentation points to obtain multiple audio segments may include the following steps:

[0103] The dry audio is segmented based on predetermined audio segmentation points to obtain multiple initial audio segments; noise is added to at least one of the multiple initial audio segments to obtain an audio segment containing noise; multiple audio segments are obtained based on the initial audio segment without added noise and the audio segment containing noise.

[0104] As an example, noise can be an audio signal that affects the accuracy of fundamental frequency sequence recognition results, such as ambient noise, song accompaniment, etc.

[0105] In practice, after segmenting the dry audio according to the audio segmentation points, the resulting audio segments can be used as initial audio segments. For the multiple initial audio segments obtained, one or more can be randomly selected from them, and noise can be added to the selected initial audio segments to obtain one or more audio segments containing noise. Then, based on the initial audio segments without added noise and the audio segments containing noise, multiple audio segments can be obtained for training the fundamental frequency sequence recognition model.

[0106] In this embodiment, by adding noise to a portion of the initial audio segments to obtain noisy audio segments, and by training the model based on the initial audio segments without added noise and the noisy audio segments, the robustness of the fundamental frequency sequence recognition model can be increased.

[0107] In one embodiment, such as Figure 6 As shown, a fundamental frequency sequence identification method is provided, which can be applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps:

[0108] S601, Obtain the spectrum of the target dry audio.

[0109] Specifically, the dry audio of the fundamental frequency sequence to be identified can be obtained and used as the target dry audio. For example, the dry audio sent by the user for mixing or pitch shifting can be used as the target dry audio, or the dry audio for melody identification can be used as the target dry audio. After obtaining the target dry audio, time-frequency conversion can be performed on the target dry audio to obtain its spectrum.

[0110] S602, input the spectrum of the target dry audio into the pre-trained fundamental frequency sequence recognition model, and determine the fundamental frequency sequence output by the fundamental frequency sequence recognition model as the fundamental frequency sequence of the target dry audio.

[0111] The fundamental frequency sequence recognition model is trained according to the training method of the fundamental frequency sequence recognition model in any of the above embodiments.

[0112] After obtaining the spectrum of the target dry audio, the spectrum can be input into a pre-trained fundamental frequency sequence recognition model. Based on the fundamental frequency sequence output by the model, the fundamental frequency sequence of the target dry audio can be determined. For example, the fundamental frequency sequence output by the model can be used as the fundamental frequency sequence of the target dry audio, or the model's output can be further processed to obtain the final fundamental frequency sequence.

[0113] In this embodiment, the spectrum of the target dry audio can be obtained, and the spectrum of the target dry audio can be input into a pre-trained fundamental frequency sequence recognition model. The fundamental frequency sequence output by the fundamental frequency sequence recognition model is determined as the fundamental frequency sequence of the target dry audio. This fundamental frequency sequence recognition model is trained according to the training method of the fundamental frequency sequence recognition model in any of the above embodiments. In the solution of this application, by inputting the spectrum of the dry audio into the trained fundamental frequency sequence recognition model, the accurate fundamental frequency sequence of the dry audio can be obtained quickly, effectively improving the recognition efficiency of the fundamental frequency sequence.

[0114] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0115] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores baseband data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a training method for a baseband sequence recognition model or a baseband sequence recognition method.

[0116] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a training method for a fundamental frequency sequence recognition model or a fundamental frequency sequence recognition method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0117] Those skilled in the art will understand that Figure 7 and Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0118] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the training method of the fundamental frequency sequence recognition model or the fundamental frequency sequence recognition method as described in any of the above embodiments.

[0119] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the training method or the fundamental frequency sequence identification method of the fundamental frequency sequence identification model as described in any of the above embodiments.

[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A training method for a fundamental frequency sequence recognition model, characterized in that, The method includes: Obtain the dry audio of the user singing a song and the spectrum of the dry audio, and perform fundamental frequency sequence identification on the dry audio to obtain the fundamental frequency sequence of the dry audio; The standard pitch of multiple phonemes in the song content is determined, the fundamental frequency of each audio frame in the dry audio is determined based on the fundamental frequency sequence of the dry audio, and the fundamental frequency sequence of the dry audio is verified based on the pitch difference between the standard pitch of the phoneme associated with each audio frame and the pitch corresponding to the fundamental frequency of each audio frame. If the verification passes, the fundamental frequency sequence is determined as the training label for the spectrum of the dry audio. Based on the spectrum of the dry audio and the training labels of the spectrum, the fundamental frequency sequence recognition model to be trained is trained to obtain the trained fundamental frequency sequence recognition model; the fundamental frequency sequence recognition model is used to predict the fundamental frequency sequence of the input spectrum.

2. The method according to claim 1, characterized in that, Determining the standard pitch of multiple phonemes in the song content includes: Obtain the sheet music of the song, and determine the standard pitch of multiple phonemes in the song content based on the sheet music.

3. The method according to claim 2, characterized in that, The step of verifying the fundamental frequency sequence of the dry audio based on the pitch difference between the standard pitch of the phoneme associated with each audio frame and the pitch corresponding to the fundamental frequency of each audio frame includes: If the pitch difference of each audio frame in the dry audio is less than the difference threshold, the fundamental frequency sequence of the dry audio is determined to pass the verification.

4. The method according to claim 1, characterized in that, The step of determining the fundamental frequency sequence as the training label for the spectrum of the dry audio includes: The dry audio is segmented based on predetermined audio segmentation points to obtain multiple audio segments; the audio segmentation points are determined based on the locations of silence and / or breathing sounds in the dry audio. For each audio segment, the fundamental frequency sequence associated with the audio segment in the fundamental frequency sequence of the dry audio is used as the training label of the spectrum of the audio segment.

5. The method according to claim 4, characterized in that, The step of training the fundamental frequency sequence recognition model to be trained based on the spectrum of the dry audio and the training labels of the spectrum to obtain the trained fundamental frequency sequence recognition model includes: The spectrum of the audio segment is input into the fundamental frequency sequence recognition model to be trained, and the predicted fundamental frequency sequence is obtained based on the fundamental frequency sequence output by the fundamental frequency sequence recognition model. Based on the difference between the predicted fundamental frequency sequence and the fundamental frequency sequence associated with the audio segment, the model parameters of the fundamental frequency sequence recognition model are adjusted until the training termination condition is met, thus obtaining a trained fundamental frequency sequence recognition model.

6. The method according to claim 5, characterized in that, The process of obtaining a predicted fundamental frequency sequence based on the fundamental frequency sequence output by the fundamental frequency sequence recognition model includes: Obtain the fundamental frequency sequence output by the fundamental frequency sequence recognition model, and determine the sound frequency range of the sound source of the dry audio. The fundamental frequencies in the fundamental frequency sequence output by the fundamental frequency sequence recognition model that are outside the sound frequency range are adjusted to preset values, and the adjusted fundamental frequency sequence is used as the predicted fundamental frequency sequence.

7. The method according to claim 4, characterized in that, The dry audio is segmented based on predetermined audio segmentation points to obtain multiple audio segments, including: The dry audio is segmented based on predetermined audio segmentation points to obtain multiple initial audio segments; Noise is added to at least one of the plurality of initial audio segments to obtain an audio segment containing noise; Multiple audio segments are obtained based on an initial audio segment without added noise and an audio segment containing noise.

8. A fundamental frequency sequence identification method, characterized in that, The method includes: Obtain the spectrum of the target dry audio; The spectrum of the target dry audio is input into a pre-trained fundamental frequency sequence recognition model, and the fundamental frequency sequence output by the fundamental frequency sequence recognition model is determined as the fundamental frequency sequence of the target dry audio. The fundamental frequency sequence identification model is trained according to the method described in any one of claims 1-7.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and apparatus for determining pitch deviation of audio content

    CN108206026A

  • Audio processing method and device, computing equipment and medium

    CN112992110A