A fundamental frequency prediction method and device

By obtaining the fundamental frequency sequence, spectrum envelope and target note sequence of the audio data to be processed, and combining the pre-trained fundamental frequency prediction model, the predicted fundamental frequency sequence of the audio data to be processed is solved, and the problem of complex and low accuracy of fundamental frequency adjustment in the prior art is achieved, achieving more natural sound editing effects and better user characteristics retention.

CN113990346BActive Publication Date: 2025-06-20TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111249248.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-06-20
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

In the prior art, works with complex basic frequency adjustment and sound edited sound sound mechanical and low accuracy.

Method used

By obtaining the fundamental frequency sequence, spectrum envelope and target note sequence of the audio data to be processed, combined with the pre-trained fundamental frequency prediction model, the predicted fundamental frequency sequence of the audio data to be processed is determined.

Benefits of technology

It reduces the complexity of the basic frequency adjustment, improves the accuracy of the obtained basic frequency, makes the modified works sound more natural, and retains the user's personal singing characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113990346B_ABST
    Figure CN113990346B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a fundamental frequency prediction method and apparatus. The method includes: obtaining a fundamental frequency sequence and a spectral envelope of audio data to be processed, and obtaining a target note sequence corresponding to the audio data to be processed, where the target note sequence is a pre-stored note sequence associated with the audio data to be processed; determining a vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence; and inputting the spectral envelope, the vuv sequence, and the target note sequence into a pre-trained fundamental frequency prediction model to determine a predicted fundamental frequency sequence corresponding to the audio data to be processed. By adopting the embodiment of the present application, the complexity of fundamental frequency adjustment can be reduced, and the accuracy of the obtained fundamental frequency can be improved, so that the tuned work sounds more natural and has more personal singing characteristics of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing, and in particular to a fundamental frequency prediction method and device. Background Art

[0002] Singing is a way of conveying emotions and entertainment that is loved by people. However, not everyone can grasp the pitch of songs as accurately as professional singers or professional performers. Therefore, non-professionals often hope to improve the pitch of their own singing works through tuning technology, so that their singing works are not out of tune and more beautiful.

[0003] In recent years, the application of online karaoke platforms in intelligent technology has become increasingly prominent. K-song platforms can support the "one-click tuning" function for most songs, automatically optimize the user's singing works, and reduce the difficulty and threshold of online karaoke. Generally speaking, the "one-click tuning" technology mainly determines whether the rhythm of the song sung by the user is accurate through the rhythm judgment algorithm, and then processes the two situations of accurate rhythm and inaccurate rhythm based on different rules. For example, when the rhythm is accurate, the basic trend of the fundamental frequency of the song sung by the user is retained, and then the fundamental frequency of the same song sung by a professional singer is used as the standard fundamental frequency, and the fundamental frequency of the user is shifted up and down for tuning; when the rhythm is inaccurate, the standard fundamental frequency is directly set to the note sequence, but this will result in the loss of the direction of the user's corresponding fundamental frequency, making the tuning result more mechanical. In general, the tuning scheme used in the relevant technology is relatively complex, and the tuned works sound very mechanical and the accuracy is not high. Summary of the invention

[0004] The embodiments of the present application provide a method and device for predicting a fundamental frequency, which can reduce the complexity of fundamental frequency adjustment and improve the accuracy of the acquired fundamental frequency.

[0005] In a first aspect, an embodiment of the present application provides a fundamental frequency prediction method, the method comprising:

[0006] Acquire a fundamental frequency sequence and a spectrum envelope of the audio data to be processed, and acquire a target note sequence corresponding to the audio data to be processed, wherein the target note sequence is a pre-stored note sequence associated with the audio data to be processed;

[0007] Determine the unvoiced and voiced sound vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence;

[0008] Input the spectral envelope, the vuv sequence, and the target note sequence into a pre-trained fundamental frequency prediction model to determine the predicted fundamental frequency sequence corresponding to the audio data to be processed, where the fundamental frequency prediction model is trained by a sample fundamental frequency sequence, a sample spectral envelope, and a sample vuv sequence and a sample note sequence obtained by converting the sample fundamental frequency sequence for the sample audio data.

[0009] Combined with the first aspect, in a possible implementation, the method further includes:

[0010] Obtain a training sample set and the sample fundamental frequency sequence and the sample spectral envelope corresponding to each sample audio data in the training sample set, and determine the sample vuv sequence and the sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence;

[0011] Input the sample spectral envelope, the sample vuv sequence, and the sample note sequence into an initial fundamental frequency prediction model to obtain the sample predicted fundamental frequency sequence output by the initial fundamental frequency prediction model;

[0012] Train the initial fundamental frequency prediction model according to the sample fundamental frequency sequence and the sample predicted fundamental frequency sequence corresponding to each sample audio data in the training sample set to obtain the fundamental frequency prediction model.

[0013] Combined with the first aspect, in a possible implementation, the obtaining the training sample set includes:

[0014] Obtain a set of candidate sample audio data, where the candidate sample audio data includes m candidate sample audio data, and m is an integer greater than 0;

[0015] Obtain the pitch distribution vector corresponding to each candidate sample audio data in the m candidate sample audio data;

[0016] Determine the sample sampling probability corresponding to each candidate sample audio data in the m candidate sample audio data according to the m pitch distribution vectors and the m;

[0017] Determine the sample audio data included in the training sample set from the m candidate sample audio data according to the sample sampling probability corresponding to each candidate sample audio data.

[0018] By specifying the screening rules for the training sample set, the samples put into model training cover various pitch distribution ranges, so that the trained model is more reliable.

[0019] Combined with the first aspect, in a possible implementation, the determining the sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence includes:

[0020] Calculate the absolute value of the difference between each adjacent sample fundamental frequency value in the sample fundamental frequency sequence;

[0021] If there are multiple consecutive absolute values of the differences greater than or equal to a preset threshold, determine at least one fundamental frequency segmentation boundary of the sample fundamental frequency sequence;

[0022] Segment the sample fundamental frequency sequence according to the at least one fundamental frequency segmentation boundary to obtain multiple fundamental frequency segmented sequences;

[0023] Perform smoothing processing on the multiple fundamental frequency segmented sequences, and splice the sequences obtained after the smoothing processing as the sample note sequence corresponding to the sample audio data.

[0024] Combined with the first aspect, in a possible implementation manner, the obtaining the target note sequence corresponding to the audio data to be processed includes:

[0025] Obtain a set of candidate note sequences corresponding to the audio data to be processed, where the set of candidate note sequences includes at least two candidate note sequences;

[0026] Determine a first parameter value based on each fundamental frequency value included in the fundamental frequency sequence;

[0027] Determine the target note sequence from the at least two candidate note sequences according to the first parameter value.

[0028] Combined with the first aspect, in a possible implementation manner, the at least two candidate note sequences include at least one candidate note sequence corresponding to a female voice and at least one candidate note sequence corresponding to a male voice; the determining the target note sequence from the at least two candidate note sequences according to the first parameter value includes:

[0029] Determine the note average value corresponding to each candidate note sequence in the at least two candidate note sequences, and determine a first parameter judgment value according to the note average values corresponding to the at least two candidate note sequences;

[0030] Determine the voice category of the audio data to be processed according to the first parameter judgment value and the first parameter value, where the voice category includes a female voice or a male voice;

[0031] Determine at least one candidate note sequence corresponding to the voice category from the at least two candidate note sequences according to the voice category of the audio data to be processed;

[0032] Determine one candidate note sequence from the at least one candidate note sequence corresponding to the voice category as the target note sequence.

[0033] By classifying the audio data to be processed into male and female voices, the target note sequence is selected from the candidate note sequences, so that the predicted fundamental frequency is more accurate and the user's personal singing characteristics can be better retained.

[0034] Combined with the first aspect, in a possible implementation manner, the determining the vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence includes:

[0035] Traverse each fundamental frequency value included in the fundamental frequency sequence;

[0036] Set the vuv value corresponding to the fundamental frequency value equal to 0 to a first preset value, and set the vuv value corresponding to the fundamental frequency value not equal to 0 to a second preset value, to obtain the vuv sequence corresponding to the audio data to be processed.

[0037] Combined with the first aspect, in a possible implementation manner, the obtaining the fundamental frequency sequence of the audio data to be processed includes:

[0038] Perform frame addition and windowing processing on the audio data to be processed to obtain N first frame signals, where N is an integer greater than 0;

[0039] Filter the first frame signals respectively through M low-pass filters with different cut-off frequencies to obtain M filtered signals corresponding to the first frame signals, where M is an integer greater than 1;

[0040] Determine a cut-off frequency from the M cut-off frequencies as the fundamental frequency value corresponding to the first frame signal according to the period information of the M filtered signals;

[0041] Generate the fundamental frequency sequence corresponding to the audio data to be processed according to the fundamental frequency values corresponding to the N first frame signals.

[0042] Combined with the first aspect, in a possible implementation manner, the obtaining the spectral envelope of the audio data to be processed includes:

[0043] Obtain the power spectrum of the audio data to be processed;

[0044] Perform inverse Fourier transform on the power spectrum of the audio data to be processed to obtain the cepstrum corresponding to the power spectrum of the audio data to be processed;

[0045] Filter the cepstrum through a low-pass filter based on a preset cut-off frequency to obtain the spectral envelope corresponding to the audio data to be processed.

[0046] In a second aspect, an embodiment of the present application provides an audio processing device, and the device includes:

[0047] An acquisition module, configured to acquire a fundamental frequency sequence and a spectral envelope of audio data to be processed, and acquire a target note sequence corresponding to the audio data to be processed;

[0048] A processing module, configured to determine a voiceless / voiced vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence;

[0049] The processing module is configured to input the spectral envelope, the vuv sequence, and the target note sequence into a pre-trained fundamental frequency prediction model to determine a predicted fundamental frequency sequence corresponding to the audio data to be processed.

[0050] Combined with the second aspect, in a possible implementation manner,

[0051] The acquisition module is further configured to acquire a training sample set and acquire a sample fundamental frequency sequence and a sample spectral envelope corresponding to each sample audio data in the training sample set, and determine a sample vuv sequence and a sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence;

[0052] The processing module is further configured to input the sample spectral envelope, the sample vuv sequence, and the sample note sequence into an initial fundamental frequency prediction model to obtain a sample predicted fundamental frequency sequence output by the initial fundamental frequency prediction model;

[0053] The processing module is further configured to train the initial fundamental frequency prediction model according to the sample fundamental frequency sequence and the sample predicted fundamental frequency sequence corresponding to each sample audio data in the training sample set to obtain the fundamental frequency prediction model.

[0054] Combined with the second aspect, in a possible implementation manner,

[0055] The acquisition module is further configured to acquire a set of candidate sample audio data, where the candidate sample audio data includes m candidate sample audio data, and m is an integer greater than 0;

[0056] The acquisition module is further configured to acquire a pitch distribution vector corresponding to each candidate sample audio data in the m candidate sample audio data;

[0057] The processing module is further configured to determine a sample sampling probability corresponding to each candidate sample audio data in the m candidate sample audio data according to the m pitch distribution vectors and the m;

[0058] The processing module is further configured to determine the sample audio data included in the training sample set from the m candidate sample audio data according to the sample sampling probability corresponding to each candidate sample audio data.

[0059] Combined with the second aspect, in a possible implementation manner,

[0060] The processing module is further configured to calculate the absolute value of the difference between each adjacent sample fundamental frequency value in the sample fundamental frequency sequence;

[0061] The processing module is further configured to determine at least one fundamental frequency segmentation boundary of the sample fundamental frequency sequence if there are multiple consecutive absolute values of differences greater than or equal to a preset threshold;

[0062] The processing module is further configured to segment the sample fundamental frequency sequence according to the at least one fundamental frequency segmentation boundary to obtain multiple fundamental frequency segmented sequences;

[0063] The processing module is further configured to perform smoothing processing on the multiple fundamental frequency segmented sequences and splice the sequences obtained after the smoothing processing as the sample note sequence corresponding to the sample audio data.

[0064] Combined with the second aspect, in a possible implementation manner,

[0065] The obtaining module is further configured to obtain a set of candidate note sequences corresponding to the audio data to be processed, where the set of candidate note sequences includes at least two candidate note sequences;

[0066] The processing module is further configured to determine a first parameter value based on each fundamental frequency value included in the fundamental frequency sequence;

[0067] The processing module is further configured to determine a target note sequence from the at least two candidate note sequences according to the first parameter value.

[0068] Combined with the second aspect, in a possible implementation manner,

[0069] The processing module is further configured to determine the note average value corresponding to each candidate note sequence in the at least two candidate note sequences, and determine a first parameter judgment value according to the note average values corresponding to the at least two candidate note sequences;

[0070] The processing module is further configured to determine the voice category of the audio data to be processed according to the first parameter judgment value and the first parameter value, where the voice category includes female voice or male voice;

[0071] The processing module is further configured to determine at least one candidate note sequence corresponding to the voice category from the at least two candidate note sequences according to the voice category of the audio data to be processed;

[0072] The processing module is further configured to determine one candidate note sequence from the at least one candidate note sequence corresponding to the voice category as the target note sequence.

[0073] In combination with the second aspect, in a possible implementation manner,

[0074] The processing module is further configured to traverse each fundamental frequency value included in the fundamental frequency sequence;

[0075] The processing module is further configured to set the vuv value corresponding to the fundamental frequency value equal to 0 to a first preset value, and set the vuv value corresponding to the fundamental frequency value not equal to 0 to a second preset value, so as to obtain a vuv sequence corresponding to the audio data to be processed.

[0076] In combination with the second aspect, in a possible implementation manner,

[0077] The processing module is further configured to perform frame addition and windowing processing on the audio data to be processed, so as to obtain N first framed signals, where N is an integer greater than 0;

[0078] The processing module is further configured to filter the first framed signals through M low-pass filters with different cut-off frequencies respectively, so as to obtain M filtered signals corresponding to the first framed signals, where M is an integer greater than 1;

[0079] The processing module is further configured to determine a cut-off frequency from the M cut-off frequencies as the fundamental frequency value corresponding to the first framed signal according to the period information of the M filtered signals;

[0080] The processing module generates a fundamental frequency sequence corresponding to the audio data to be processed according to the fundamental frequency values corresponding to the N first framed signals.

[0081] In combination with the second aspect, in a possible implementation manner,

[0082] The obtaining module is further configured to obtain the power spectrum of the audio data to be processed;

[0083] The processing module is further configured to perform inverse Fourier transform on the power spectrum of the audio data to be processed, so as to obtain the cepstrum corresponding to the power spectrum of the audio data to be processed;

[0084] The processing module is further configured to filter the cepstrum through a low-pass filter with a preset cut-off frequency, so as to obtain the spectral envelope corresponding to the audio data to be processed.

[0085] In a third aspect, an embodiment of the present application provides a terminal device, which includes a processor and a memory, and the processor and the memory are connected to each other. The memory is used to store a computer program for supporting the terminal device to execute the method provided in the first aspect and / or any possible implementation manner of the first aspect. The computer program includes program instructions, and the processor is configured to call the above program instructions to execute the method provided in the first aspect and / or any possible implementation manner of the first aspect.

[0086] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, where the computer program includes program instructions that, when executed by a processor, cause the processor to execute the method provided in the above first aspect and / or any possible implementation manner of the first aspect.

[0087] In an embodiment of the present application, a fundamental frequency sequence and a spectral envelope of the audio data to be processed are obtained, and a target note sequence corresponding to the audio data to be processed is obtained, where the target note sequence is a pre-stored note sequence associated with the audio data to be processed. Then, a voiceless / voiced vuv sequence corresponding to the audio data to be processed can be determined according to the fundamental frequency sequence. Further, according to the spectral envelope, the vuv sequence, and the target note sequence, in combination with a fundamental frequency prediction model, a predicted fundamental frequency sequence corresponding to the audio data to be processed can be determined. The fundamental frequency prediction model is trained by a sample fundamental frequency sequence, a sample spectral envelope, a sample vuv sequence converted from the sample fundamental frequency sequence, and a sample note sequence corresponding to the sample audio data. In an embodiment of the present application, by extracting the sound features of the audio data to be processed and then putting a preset note sequence (i.e., the target note sequence) together with the extracted sound features into a prediction model to obtain a predicted fundamental frequency for pitch correction purposes, not only the complexity of pitch correction is reduced, but also the pitch-corrected work is more in line with the auditory perception law of the human auditory sense while retaining the original singer's voice characteristics, and it sounds smoother and more natural. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments.

[0089] Figure 1 is a schematic diagram of an application scenario of the fundamental frequency prediction method provided by an embodiment of the present application;

[0090] Figure 2 is a flowchart of the fundamental frequency prediction method provided by an embodiment of the present application;

[0091] Figure 3 is a waveform diagram of a sine wave signal provided by an embodiment of the present application;

[0092] Figure 4 is a schematic diagram of a spectral envelope provided by an embodiment of the present application;

[0093] Figure 5 is a schematic diagram of obtaining a note sequence provided by an embodiment of the present application;

[0094] Figure 6 is another flowchart of the fundamental frequency prediction method provided by an embodiment of the present application;

[0095] Figure 7 It is a schematic flowchart of the model training process in the fundamental frequency prediction method provided by an embodiment of the present application;

[0096] Figure 8 It is a schematic structural diagram of the fundamental frequency prediction model provided by an embodiment of the present application;

[0097] Figure 9 It is a schematic diagram of the relationships of various sequences provided by an embodiment of the present application;

[0098] Figure 10 It is a schematic structural diagram of the fundamental frequency prediction device provided by an embodiment of the present application;

[0099] Figure 11 It is a schematic structural diagram of the terminal device provided by an embodiment of the present application. Detailed implementation manners

[0100] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0101] The embodiments of the present application relate to artificial intelligence (AI) and machine learning (ML). Among them, AI uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science, which mainly produces a new intelligent machine that can react in a way similar to human intelligence by understanding the essence of intelligence, so that the intelligent machine has multiple functions such as perception, reasoning, and decision-making.

[0102] AI technology is an interdisciplinary subject, which mainly includes several major directions such as computer vision technology (CV), speech processing technology, natural language processing technology, and machine learning (ML) or deep learning. Among them, computer vision technology is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, tracking, and measurement on targets, and further perform graphics processing to make the computer process images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data; it usually includes technologies such as audio data processing, video processing, video semantic understanding, video content, and video behavior recognition.

[0103] Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of AI and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning / deep learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0104] Based on the computer vision technology and machine learning technology in AI technology, the embodiment of this application provides a fundamental frequency prediction method. This method obtains the works sung by the user as the audio data to be processed, then extracts the sound features (including the fundamental frequency sequence and the unvoiced sound / voiced sound sequence, that is, the vuv sequence) from the audio data to be processed, and selects the target note sequence from multiple candidate note sequences. Finally, the sound features and the target note sequence are used as parameters and put into the model, and the predicted audio output by the model can be obtained. This application achieves the effect of retaining the singing characteristics of the user by extracting the user's sound features, and achieves the purpose of making the tuned works sound more natural by using the sound features and the target note sequence as parameters and putting them into the model, and the implementation complexity of the solution is low.

[0105] Specifically, in the application stage, firstly, the fundamental frequency sequence and spectrum envelope of the audio data to be processed are obtained, and the target note sequence corresponding to the audio data to be processed is obtained, wherein the target note sequence is determined from a plurality of candidate note sequences according to the fundamental frequency sequence extracted from the audio data to be processed, and the target note sequence is associated with the audio data to be processed (for example, the audio data to be processed and the target note sequence are the same song and both the audio data to be processed and the target note sequence meet the characteristics of male / female singing). Then, the unvoiced and voiced vuv sequence corresponding to the audio data to be processed is determined according to the fundamental frequency sequence, wherein the unvoiced and voiced vuv sequence may include a first preset value and a second preset value (in the embodiment of the present application, the first preset value may be 0 and the second preset value may be 1); finally, according to the spectrum envelope, the vuv sequence and the target note sequence, the predicted fundamental frequency sequence corresponding to the audio data to be processed is determined in combination with the fundamental frequency prediction model. Among them, the fundamental frequency prediction model is obtained by training the sample fundamental frequency sequence corresponding to the sample audio data, the sample spectrum envelope, and the sample vuv sequence and the sample note sequence converted from the sample fundamental frequency sequence, wherein the sample audio data is determined from the candidate sample audio data collection according to the sampling probability. By adopting the embodiments of the present application, the complexity of fundamental frequency adjustment can be reduced, and the accuracy of the acquired fundamental frequency can be improved, so that the tuned work sounds more natural and has the user's personal singing characteristics.

[0106] See also Figure 1 , Figure 1 : is a schematic diagram of an application scenario of the fundamental frequency prediction method provided in an embodiment of the present application. Figure 1 101 in represents the base frequency sequence before adjustment, and 102 represents the base frequency sequence after adjustment. Figure 1 As shown, the user selects the work that needs to be tuned and enters Figure 1 The interface shown in the figure allows users to select a certain section or multiple sections or the entire work to be tuned according to their needs. The tuning process includes the user turning on the tuning button and selecting the strength of the tuning effect according to their needs (if no selection is made, the default value will be automatically selected), and the current tuning section will be displayed on the screen, such as Figure 1 As shown, 101 is the baseband sequence before adjustment, 102 is the adjusted baseband sequence obtained after the audio is adjusted according to the adjustment strength selected by the user, and 101 is displayed as 102 after adjustment.

[0107] It should be noted that Figure 1 The interface in the application scenario of the fundamental frequency prediction method shown and the specific operation steps described above are only examples and are not limited here.

[0108] It should be noted that the fundamental frequency prediction method provided by the embodiments of the present application can be widely applied to various computer devices. Among them, the above-mentioned computer devices include, but are not limited to, servers, smart phones, tablet computers, laptop computers, desktop computers, etc., and are not limited here. For the convenience of description, the following will take the terminal device as an example for illustrative explanation.

[0109] Next, the methods and related devices provided by the embodiments of the present application will be described in detail respectively. Figures 2 to 11

[0110] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the fundamental frequency prediction method provided by the embodiments of the present application. The method provided by the embodiments of the present application may include the following steps S201 to S203.

[0111] S201. Obtain the fundamental frequency sequence and spectral envelope of the audio data to be processed, and obtain the target note sequence corresponding to the audio data to be processed, where the target note sequence is a pre-stored note sequence associated with the audio data to be processed.

[0112] In some feasible embodiments, the pre-stored audio can be obtained from the local storage of the terminal device or from an external memory connected to the terminal device as the audio data to be processed. Or the audio recorded by the microphone of the terminal device can also be obtained in real time as the audio data to be processed. It can be understood that the above-mentioned audio data to be processed can be a work sung by the user before, a work being sung at present, a work sung by others, etc., and is not limited here. Furthermore, the terminal device can perform audio feature extraction on the obtained audio data to be processed to obtain the fundamental frequency sequence and spectral envelope of the audio data to be processed, and obtain the target note sequence corresponding to the audio data to be processed. Among them, the target note sequence is a pre-stored note sequence associated with the audio data to be processed. It should be understood that the audio data to be processed involved in the embodiments of the present application is dry sound, that is, pure human voice without accompaniment or music.

[0113] Specifically, obtaining the fundamental frequency sequence of the audio data to be processed includes: performing frame addition and windowing processing on the audio data to be processed to obtain N first sub-frame signals, where N is an integer greater than 0; filtering each first sub-frame signal through M low-pass filters with different cut-off frequencies to obtain M filtered signals corresponding to each first sub-frame signal, where M is an integer greater than 1. For the M filtered signals corresponding to any one first sub-frame signal, according to the period information of the M filtered signals corresponding to the first sub-frame signal, a cut-off frequency can be determined from the M cut-off frequencies as the fundamental frequency value corresponding to the first sub-frame signal. Finally, according to the N fundamental frequency values corresponding to the determined N first sub-frame signals, the fundamental frequency sequence corresponding to the audio data to be processed is generated.

[0114] Among them, the processing of each of the N first sub-frame signals is the same. Therefore, for ease of understanding, the following embodiments of the present application mainly take the processing of one of the N first sub-frame signals as an example for description. Among them, for ease of description, this one first sub-frame signal can be described as the target first sub-frame signal.

[0115] Specifically, by respectively filtering the target first sub-frame signal through M low-pass filters with different cut-off frequencies, M filtered signals corresponding to the target first sub-frame signal can be obtained. The target first sub-frame signal is one of the N first sub-frame signals, and M is an integer greater than 1. One cut-off frequency is determined from the M cut-off frequencies as the fundamental frequency value corresponding to the target first sub-frame signal according to the period information of the M filtered signals. A fundamental frequency sequence corresponding to the audio data to be processed is generated according to the N fundamental frequency values corresponding to the N first sub-frame signals.

[0116] Among them, the window function used in the windowing operation includes any one of the following window functions: rectangular window, Hamming window, Hanning window, and the selection of the window function is not limited here. Generally speaking, the frame length of the framing operation can be selected in the range of 10 milliseconds to 30 milliseconds, and the frame shift is selected in the range of 0 to 1 / 2 of the ratio to the frame length according to the actual situation. For example, if the frame length is selected as 20 milliseconds, then the frame shift can be 8 milliseconds, etc., and it is not limited here.

[0117] Among the M filtering signals corresponding to the target first sub-frame signal, the determination rule for determining a cut-off frequency from the M cut-off frequencies as the fundamental frequency value corresponding to the target first sub-frame signal according to the period information of the M filtering signals can be: for each of the M filtering signals on the coordinate axis with the abscissa being time and the ordinate being amplitude, obtain 4 time points when each filtering signal crosses the abscissa (i.e., the amplitude is 0). Suppose the M filtering signals are the filtering signal M1 corresponding to the cut-off frequency 1, the filtering signal M2 corresponding to the cut-off frequency 2, the filtering signal M3 corresponding to the cut-off frequency 3, and the filtering signal M4 corresponding to the cut-off frequency 4. For ease of understanding, here, the processing of one of the M filtering signals (for example, the filtering signal M1) is taken as an example for illustrative explanation. Among them, the four time points corresponding to the filtering signal M1 can be respectively described as the first time point t1, the second time point t2, the third time point t3, and the fourth time point t4. Then, calculate the absolute value of the difference △T1 between the second time point t2 and the first time point t1 among these four time points (i.e., △T1 = |t2 - t1|), and the absolute value of the difference △T2 between the fourth time point t4 and the third time point (i.e., △T2 = |t4 - t3|). Further, calculate the reciprocal △T1' of the absolute value of the difference between △T1 and △T2 (i.e., △T1' = 1 / |△T2 - △T1|) as the confidence value of the cut-off frequency 1. By analogy, the confidence value △T2' corresponding to the cut-off frequency 2, the confidence value △T3' corresponding to the cut-off frequency 3, and the confidence value △T4' corresponding to the cut-off frequency 4 can be obtained respectively. Compare each confidence value with a preset judgment threshold. When the confidence values of the signals filtered by multiple cut-off frequencies are all greater than the judgment threshold, select the one with the smallest cut-off frequency among the multiple cut-off frequencies as the fundamental frequency value corresponding to the signal.

[0118] For example, for the first sub-frame signal 1, it is now filtered by two cut-off frequencies: cut-off frequency A and cut-off frequency B respectively. Among them, cut-off frequency A is less than cut-off frequency B. In the first case, a judgment threshold is preset to be 1 / 3. For a section of the waveform obtained by filtering with cut-off frequency A on the coordinate axis with the abscissa being time and the ordinate being amplitude, the points where it crosses the abscissa (i.e., the amplitude is 0) are successively: t1 = 1, t2 = 4, t3 = 8, t4 = 13. The absolute value of the difference between the abscissa of the second point of this waveform (i.e., t2 = 4) and the abscissa of the first point (i.e., t1 = 1) is △T1 = (4 - 1) = 3, and the absolute value of the difference between the abscissa of the fourth point (i.e., t4 = 13) and the abscissa of the third point (i.e., t3 = 8) is △T2 = (13 - 8) = 5. The reciprocal of the absolute value of the difference between these two absolute values is △T1’ = 1 / (5 - 3) = 1 / 2, that is, the confidence value of cut-off frequency A as the fundamental frequency value of the first sub-frame signal 1; in another section of the waveform obtained by filtering with cut-off frequency B, the points where it crosses the abscissa (i.e., the amplitude is 0) are successively: t5 = 1, t6 = 3, t7 = 10, t8 = 16, and the absolute value of the difference between the abscissa of the second point (i.e., t6 = 3) and the abscissa of the first point (i.e., t5 = 1) is △T3 = (3 - 1) = 2, and the absolute value of the difference between the abscissa of the fourth point (i.e., t8 = 16) and the abscissa of the third point (i.e., t7 = 10) is △T4 = (16 - 10) = 6. The reciprocal of the absolute value of the difference between these two absolute values is △T2’ = 1 / (6 - 2) = 1 / 4, that is, the confidence value of cut-off frequency B as the fundamental frequency value of the first sub-frame signal 1. Since the confidence value △ (1 / 2) of cut-off frequency A is greater than the judgment threshold (1 / 3), and the confidence value △T2’ (1 / 4) of cut-off frequency B is less than the judgment threshold (1 / 3). Therefore, cut-off frequency A is selected between cut-off frequency A and cut-off frequency B. In the second case, a judgment threshold is preset to be 1 / 5. The confidence value △T3’ of the waveform filtered by cut-off frequency A is 1 / 2, and the confidence value △T4’ of the waveform filtered by cut-off frequency B is 1 / 4. Although both of them (T3’ = 1 / 2 and △T4’ = 1 / 4) are greater than the judgment threshold (judgment threshold = 1 / 5) at this time, because cut-off frequency A is greater than cut-off frequency B, cut-off frequency B with a smaller cut-off frequency is selected between cut-off frequency A and cut-off frequency B as the fundamental frequency value corresponding to the first sub-frame signal 1.

[0119] Optionally, in addition to the above-mentioned difference method of passing through the horizontal coordinate point (referred to as Rule 1), the rule for obtaining the confidence of the waveform can also be the following Rule 2. Specifically, Rule 2 is: after filtering the target first sub-frame signal using M cutoff frequencies to obtain M filtered signals distributed on the coordinate axis with the horizontal coordinate being time and the vertical coordinate being frequency, respectively calculate the reciprocal of the average value of the absolute value of the difference between multiple adjacent signal periods of each filtered signal, and use the reciprocal of the average value as the first judgment value; then calculate the reciprocal of the average value of the absolute value of the difference between the widths of multiple adjacent partial waveforms when the filtered signal is at the same height from the horizontal coordinate, and use the reciprocal of the average value as the second judgment value. The sum of the first judgment value and the second judgment value is used as the confidence of the waveform.

[0120] Exemplarily, when M is equal to 4, four low-pass filters with different cutoff frequencies are used to filter the target first sub-frame signal, and the above four different cutoff frequencies are respectively M1 cutoff frequency, M2 cutoff frequency, M3 cutoff frequency and M4 cutoff frequency, wherein the size relationship between the four cutoff frequencies can be: M1 cutoff frequency>M2 cutoff frequency>M3 cutoff frequency>M4 cutoff frequency, or: M3 cutoff frequency>M4 cutoff frequency>M2 cutoff frequency>M1 cutoff frequency, and the size relationship between the four cutoff frequencies is not limited here. For ease of understanding, the following embodiments of the present application are all schematically explained by taking M1 cutoff frequency>M2 cutoff frequency>M3 cutoff frequency>M4 cutoff frequency as an example. The target first sub-frame signal is filtered using the M1 cutoff frequency to obtain the M1 filtered signal, the target first sub-frame signal is filtered using the M2 cutoff frequency to obtain the M2 filtered signal, the target first sub-frame signal is filtered using the M3 cutoff frequency to obtain the M3 filtered signal, and the target first sub-frame signal is filtered using the M4 cutoff frequency to obtain the M4 filtered signal. It should be noted that, for the convenience of subsequent description, in this embodiment, M1, M2, M3 and M4 are collectively referred to as Mn, that is, for the processing of each of the M1 filter signal, the M2 filter signal, the M3 filter signal and the M4 filter signal, the following embodiments of this application are all described by taking the processing of the Mn filter signal as an example. Figure 3 , Figure 3 : is a waveform diagram of a sine wave signal provided in an embodiment of the present application. The target first frame signal is filtered using the Mn cutoff frequency to obtain the following: Figure 3 The Mn filtered signal shown, Figure 3 The four marked t0, t1, t2 and t3 are all four points that cross the horizontal coordinate (the vertical coordinate amplitude is 0). The Mn filter signal includes three signal periods, namely T1, T2 and T3, where T1 = (t1-t0), T2 = (t2-t1) and T3 = (t3-t2). Because the signal period of the sine wave signal satisfies that each signal period is the same, that is, ifFigure 3 If the signal is a sine wave signal, then the three periods T1, T2, and T3 satisfy T1 = T2 = T3. Therefore, the closer the absolute value of the difference between two adjacent periods among the three periods T1, T2, and T3 of the Mn filtered signal is to 0, that is, the larger the reciprocal of this absolute value is, the more similar the Mn filtered signal is to the sine signal. Thus, the reciprocal of the average value of the absolute values of the period differences of the filtered signal is used as the first evaluation value for judging the similarity between the filtered signal and the sine signal. That is to say, the calculation formula for the first evaluation value of the Mn filtered signal is:

[0121]

[0122] where T1, T2, and T3 represent three consecutive signal periods of the Mn filtered signal.

[0123] Please also refer to Figure 3 . Take two horizontal lines with the same height h1 as the abscissa on the waveform of the Mn filtered signal, such as Figure 3 the two horizontal lines marked d1 and d2 shown. Their widths are d1 and d2 respectively. In the sine wave signal, each width with the same height as the abscissa is the same. That is to say, if Figure 3 the waveform shown is a sine wave, then d1 = d2 or it can be expressed as d2 - d1 = 0. Therefore, the closer the absolute value of the difference between the two widths d1 and d2 of the Mn filtered signal is to 0, that is, the larger the reciprocal of this absolute value is, the more similar the Mn filtered signal is to the sine signal. Thus, the reciprocal of the average value of the absolute values of the width differences of the filtered signal is used as the second evaluation value for judging the similarity between the filtered signal and the sine signal. That is to say, the calculation formula for the second evaluation value of the Mn filtered signal is:

[0124]

[0125] where d1, d2, and d3 represent three widths formed by taking four adjacent points with the same height as the abscissa but not exceeding the peak of the waveform of the Mn filtered signal, pairwise.

[0126] Because the sum of the first evaluation value and the second evaluation value is the evaluation value, that is, the first evaluation value + the second evaluation value = the confidence level of the waveform. In the embodiments of the present application, it is expressed as:

[0127]

[0128] where T1, T2, and T3 represent three signal periods included in the Mn filtered signal, and d1, d2, and d3 represent the waveform widths when the three waveforms are at y = y1.

[0129] Thus, the confidence of the waveform obtained by filtering with the cut-off frequencies of M1, M2, M3, and M4 can be obtained by the above method.

[0130] By setting the evaluation criteria for the filtered signal, the filtered signal with the highest standard can be found more accurately to obtain the appropriate cut-off frequency, so that the obtained fundamental frequency and fundamental frequency sequence are more accurate, and finally the predicted fundamental frequency is more accurate.

[0131] Among them, obtaining the spectral envelope of the audio data to be processed includes: obtaining the power spectrum of the audio data to be processed; performing inverse Fourier transform on the power spectrum of the audio data to be processed to obtain the cepstrum corresponding to the power spectrum of the audio data to be processed; filtering the cepstrum based on a low-pass filter with a preset cut-off frequency to obtain the spectral envelope corresponding to the audio data to be processed. Among them, the frequency envelope determines the timbre in popular terms. Specifically, first, the audio data to be processed needs to be framed and windowed to obtain at least one second framed signal. By performing short-time Fourier transform on each second framed signal, the sub-spectrum corresponding to each second framed signal can be obtained. It can be understood that the first framed signal and the second framed signal can be the same. Among them, by taking the absolute value of the sub-spectrum corresponding to each second framed signal, the power spectrum of each second framed signal can be obtained. Furthermore, by taking the logarithm of the power spectrum corresponding to each second framed signal, performing phase unwrapping, and then performing inverse Fourier transform, the cepstrum of the power spectrum corresponding to each second framed signal can be obtained. Finally, filtering the cepstrum of the power spectrum corresponding to each second framed signal based on a low-pass filter can obtain multiple sub-spectrum envelopes corresponding to each second framed signal. Finally, splicing the multiple sub-spectrum envelopes can obtain the spectral envelope corresponding to the audio data to be processed. Exemplarily, please refer to Figure 4 , Figure 4 is the schematic diagram of the spectral envelope provided by the embodiment of the present application. As Figure 4 shown, on the coordinate axis where the abscissa represents frequency and the ordinate represents amplitude, a continuous spectral envelope can be composed of splicing multiple sub-spectrum envelopes.

[0132] Among them, obtaining the target note sequence corresponding to the audio data to be processed includes: obtaining a set of candidate note sequences corresponding to the audio data to be processed; determining a first parameter value based on each fundamental frequency value included in the fundamental frequency sequence; and determining the target note sequence from at least two candidate note sequences according to the first parameter value. Among them, the set of candidate note sequences includes at least two candidate note sequences, and the at least two candidate note sequences include at least one candidate note sequence corresponding to a female voice and at least one candidate note sequence corresponding to a male voice. It can be understood that the note sequence corresponding to a song sung by a professional singer or a singer, etc. can be used as the candidate note sequence included in the set of candidate note sequences corresponding to the song, and there is no limitation here.

[0133] Among them, at least two candidate note sequences include a candidate note sequence corresponding to at least one female voice and a candidate note sequence corresponding to at least one male voice; determining a target note sequence from at least two candidate note sequences according to the first parameter value includes: determining the average note value corresponding to each candidate note sequence among at least two candidate note sequences, and determining a first parameter judgment value according to the average note values corresponding to at least two candidate note sequences; determining the voice category of the audio data to be processed according to the first parameter judgment value and the first parameter value, where the voice category includes female voice or male voice; determining at least one candidate note sequence corresponding to the voice category from at least two candidate note sequences according to the voice category of the audio data to be processed; and determining one candidate note sequence from at least one candidate note sequence corresponding to the voice category as the target note sequence. It can be understood that the dry voice corresponding to the candidate note sequence can be obtained by a professional singer or a professional vocalist. And the dry voice corresponding to the audio data to be processed can be obtained by the user singing the same song (with an "off-key" performance).

[0134] Among them, the above-mentioned first parameter value can be the average fundamental frequency value of the fundamental frequency sequence of the audio data to be processed, or the first parameter value can also be the distribution probability of the fundamental frequency values of the fundamental frequency sequence of the audio data to be processed, or it can also be the median fundamental frequency value in the fundamental frequency sequence of the audio data to be processed.

[0135] Among them, the first parameter value is the average fundamental frequency value of the fundamental frequency sequence of the audio data to be processed. That is, the sum of all the fundamental frequency values in the fundamental frequency sequence of the audio data to be processed is divided by the number of fundamental frequency values included in the fundamental frequency sequence of the processed audio data. For example, if the fundamental frequency sequence of a piece of audio data to be processed is [120, 125, 120, 130, 131], then the average fundamental frequency value (120 + 125 + 120 + 130 + 131) / 5 = 125.2 of the fundamental frequency sequence [120, 125, 120, 130, 131] is used as the first parameter value.

[0136] Optionally, the first parameter value is the probability distribution of the fundamental frequency values of the fundamental frequency sequence of the audio data to be processed. It can be understood that the probability of the fundamental frequency values of the fundamental frequency sequence of the audio data to be processed distributed in each preset frequency range is calculated. For example, if the duration of the audio data to be processed is 3 minutes and the length of each frame is 18 milliseconds, then the audio data to be processed includes 10,000 frames (3 minutes is equal to 180,000 milliseconds, and if the length of one frame is 18 milliseconds, then 3 minutes, that is, 180,000 milliseconds, includes 180,000÷18 = 10,000 frames). Among them, the number of frames in the range of 698HZ to 1.1KHZ is 1,400, and the probability distribution of the fundamental frequency values in this frequency range, which is 14%, is taken as the first parameter value. Another example is that the fundamental frequency sequence of a piece of audio data to be processed is [120, 125, 120, 130, 131], and the fundamental frequency values distributed in the range of 60HZ to 123HZ (the human vocalization frequency is generally greater than 60HZ, so frequencies below 60HZ are usually not considered) are: 120 and 120. The probability of its distribution in the range of 60HZ to 123HZ, which is 2 / 5 = 40%, is taken as the first parameter value.

[0137] Optionally, the first parameter value is the median fundamental frequency in the fundamental frequency sequence of the audio data to be processed. It can be understood that after sorting the fundamental frequency values in the fundamental frequency sequence of the audio data to be processed from small to large or from large to small, the middle value of the sorted fundamental frequency sequence is taken as the first parameter value. For example, the fundamental frequency sequence of a piece of audio data to be processed is [120, 125, 120, 130, 131]. After sorting the fundamental frequency sequence from small to large, the obtained fundamental frequency sequence is [120, 120, 125, 130, 131], and the middle value 125 of the sorted fundamental frequency sequence is taken as the first parameter value.

[0138] Among them, when the first parameter value is the average fundamental frequency value of the fundamental frequency sequence of the audio data to be processed, the preset first parameter value judgment value is also the preset average fundamental frequency value. The voice category of the audio data to be processed is judged by judging the magnitude relationship between the first parameter value and the first parameter judgment value (the voice category includes female voice or male voice). It can be understood that if the first parameter value is less than the first parameter judgment value, the audio data to be processed is a male voice; if the first parameter value is greater than the first parameter judgment value, the audio data to be processed is a female voice. At least one candidate note sequence corresponding to the voice category is determined from at least two candidate note sequences according to the voice category of the audio data to be processed. For example, the fundamental frequency sequence of a piece of audio data to be processed is [120, 125, 120, 130, 131], then the average fundamental frequency value of this fundamental frequency sequence [120, 125, 120, 130, 131], (120 + 125 + 120 + 130 + 131) / 5 = 125.2 is used as the first parameter value. The preset first parameter value judgment value is 130, then the first parameter value (125.2) is less than the first parameter judgment value (130), and it is judged that the voice category of this audio data to be processed is a male voice.

[0139] Optionally, when the first parameter value is the distribution probability of the fundamental frequency values of the fundamental frequency sequence of the audio data to be processed, the preset first parameter value judgment value is also the preset distribution probability of the fundamental frequency values. The voice category of the audio data to be processed is judged by judging the magnitude relationship between the first parameter value and the first parameter judgment value (the voice category includes female voice or male voice). It can be understood that if the first parameter value is less than the first parameter judgment value, the audio data to be processed is a male voice; if the first parameter value is greater than the first parameter judgment value, the audio data to be processed is a female voice. At least one candidate note sequence corresponding to the voice category is determined from at least two candidate note sequences according to the voice category of the audio data to be processed. For example, the fundamental frequency sequence of a piece of audio data to be processed is [120, 125, 120, 130, 131], and the fundamental frequency values distributed in 60HZ - 123HZ (the vocal frequency of people is generally greater than 60HZ, so the frequencies below 60HZ are usually not considered) are: 120 and 120, and the probability of its distribution in 60HZ - 123HZ is 2 / 5 = 40% as the first parameter value. The preset first parameter value judgment value is 50%, then the first parameter value (40%) is less than the first parameter judgment value (50%), and it is judged that the voice category of this audio data to be processed is a male voice.

[0140] Optionally, when the first parameter value is the median fundamental frequency of the fundamental frequency sequence of the audio data to be processed, the preset first parameter value judgment value is also the preset median fundamental frequency. By judging the magnitude relationship between the first parameter value and the first parameter judgment value, the voice category of the audio data to be processed can be determined (the voice category includes female voice or male voice). It can be understood that if the first parameter value is less than the first parameter value judgment value, the audio data to be processed is a male voice; if the first parameter value is greater than the first parameter value judgment value, the audio data to be processed is a female voice. At least one candidate note sequence corresponding to the voice category is determined from at least two candidate note sequences according to the voice category of the audio data to be processed. For example, the fundamental frequency sequence of a piece of audio data to be processed is [120, 125, 120, 130, 131]. After sorting the fundamental frequency sequence from small to large, the obtained fundamental frequency sequence is [120, 120, 125, 130, 131]. The value 125 in the middle of the sorted fundamental frequency sequence is taken as the first parameter value. The preset first parameter value judgment value is 130. Then the first parameter value (125) is less than the first parameter value judgment value (130), and it is judged that the voice category of the audio data to be processed is a male voice.

[0141] Among them, if the voice category includes multiple candidate note sequences, one candidate note sequence is randomly selected from the multiple candidate note sequences in the voice category as the target note sequence.

[0142] S202. Determine the voiceless / voiced vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence.

[0143] In some feasible embodiments, determining the vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence includes: traversing each fundamental frequency value included in the fundamental frequency sequence; setting the vuv value corresponding to the fundamental frequency value equal to 0 as the first preset value, and setting the vuv value corresponding to the fundamental frequency value not equal to 0 as the second preset value, to obtain the vuv sequence corresponding to the audio data to be processed. Among them, the vuv sequence can be understood as being used to judge whether there is a fundamental frequency in the speech or sound, or to judge whether the speech or sound has vocal cord vibration. Signals with vocal cord vibration are generally the so-called voiced consonants or vowels (such as b, d, g, v in Chinese pinyin are all voiced consonants), and signals without vocal cord vibration are generally the so-called voiceless consonants or silences (such as p, t, g, f in Chinese pinyin are all voiceless consonants).

[0144] In the embodiments of the present application, the first preset value may specifically be 0, and the second preset value may specifically be 1. It can be understood that when the fundamental frequency value is equal to 0, the corresponding VUV value is 0, and when the fundamental frequency value is not equal to 0, the corresponding VUV value is 1. For example, for a fundamental frequency sequence [120, 121, 121, 0, 0, 0, 163], its first fundamental frequency value 120 is a number not equal to 0, so the VUV value corresponding to the first fundamental frequency value 120 is 1, and its fourth fundamental frequency value is 0, so the VUV value corresponding to the fourth fundamental frequency value 0 is 0. Therefore, the fundamental frequency sequence [120, 121, 121, 0, 0, 0, 163] corresponds to the VUV sequence [1, 1, 1, 0, 0, 0, 1]. For another example, for a fundamental frequency sequence [0, 0, 0, 0, 0, 0, 0], since each of its fundamental frequency values is 0, the VUV sequence corresponding to this fundamental frequency sequence is [0, 0, 0, 0, 0, 0, 0].

[0145] S203. Determine a predicted fundamental frequency sequence corresponding to the audio data to be processed by combining the spectral envelope, the VUV sequence, and the target note sequence with the fundamental frequency prediction model.

[0146] In some feasible implementation manners, a predicted fundamental frequency sequence corresponding to the audio data to be processed is determined by combining the spectral envelope, the VUV sequence, and the target note sequence with the fundamental frequency prediction model. Among them, the fundamental frequency prediction model may include, for example Figure 8 the fundamental frequency prediction model shown, where 801 is a schematic diagram of the prediction model, and 802 is a schematic diagram of the structure of the self-attention mechanism module. The spectral envelope (for example, the dimension of this spectral envelope may be 1025-dimensional), the note sequence (for example, the dimension of this note sequence may be 256-dimensional), and the VUV sequence (for example, the dimension of this VUV sequence may be 1-dimensional) are concatenated to obtain a vector with a dimension of 1025 + 256 + 1 = 1282 (the dimension of the spectral envelope + the dimension of the note sequence + the dimension of the VUV sequence). This vector passes through linear layer 1 to obtain a 256-dimensional fused feature vector. The 256-dimensional fused feature vector passes through the self-attention mechanism module to obtain a 768-dimensional feature vector. Finally, this 768-dimensional feature vector is input into linear layer 2 to obtain a 1-dimensional predicted fundamental frequency. By calculating the predicted fundamental frequency and the extracted fundamental frequency with the mean squared error as the loss function, this fundamental frequency prediction model is optimized. The optimization method may be Adam, and the learning rate is 1e-3. Among them, as Figure 8FIG. 802 is a schematic structural diagram of the self-attention mechanism module. The self-attention mechanism module includes linear layer 3, linear layer 4, and linear layer 5, an Attention layer, a splicing layer, and linear layer 6. Among them, the fused feature vector output by linear layer 1 is sequentially passed through linear layer 3, linear layer 4, and linear layer 5, the Attention layer, the splicing layer, and linear layer 6 of the self-attention mechanism module, and the output of the second linear layer (i.e., linear layer 6) can be used as the output of the self-attention mechanism module and input to linear layer 2.

[0147] Exemplarily, Figure 9 FIG. is a schematic diagram of the relationship between various sequences provided by an embodiment of the present application. Among them Figure 9 line 1 therein is the vuv sequence, line 2 is the predicted fundamental frequency sequence, line 3 is the note sequence, and line 4 is the fundamental frequency sequence. Combining the above description of Figure 8 it can be understood that after putting line 1 (vuv sequence), line 3 (spectral envelope), and line 4 (true fundamental frequency sequence) into the model as parameters, line 2 (predicted fundamental frequency sequence) can be obtained. The relationship among line 1 (vuv sequence), line 3 (target note sequence), line 4 (true fundamental frequency sequence, that is, the fundamental frequency sequence corresponding to the audio data to be processed extracted), and line 2 (predicted fundamental frequency sequence) is as Figure 9 shown.

[0148] Among them, the fundamental frequency prediction model can be trained by the sample fundamental frequency sequence, sample spectral envelope, and sample vuv sequence and sample note sequence converted from the sample fundamental frequency sequence corresponding to the sample audio data. Specifically, the training process of the fundamental frequency prediction model includes: obtaining a training sample set and the sample fundamental frequency sequence and sample spectral envelope corresponding to each sample audio data in the training sample set, and determining the sample vuv sequence and sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence; inputting the sample spectral envelope, sample vuv sequence, and sample note sequence into the initial fundamental frequency prediction model to obtain the sample predicted fundamental frequency sequence output by the initial fundamental frequency prediction model; training the initial fundamental frequency prediction model according to the sample fundamental frequency sequence and sample predicted fundamental frequency sequence corresponding to each sample audio data in the training sample set to obtain the fundamental frequency prediction model. The training sample set includes multiple sample audio data corresponding to the first pitch distribution range, multiple sample audio data corresponding to the second pitch distribution range, and multiple sample audio data corresponding to the third pitch distribution range.

[0149] Among them, the training sample set includes multiple sample audio data corresponding to the first pitch distribution range, which can be understood as audio data in the range of 60HZ to 120HZ (commonly known as bass), multiple sample audio data corresponding to the second pitch distribution range, which can be understood as audio data in the range of 120HZ to 480HZ (commonly known as midrange), and multiple sample audio data corresponding to the third pitch distribution range, which can be understood as audio data in the range of 120HZ to 480HZ (commonly known as treble). Generally speaking, since the vocal frequency of a person is between 60Hz and 1200Hz, for a fundamental frequency prediction model, if it is desired that the fundamental frequency prediction model obtained by training can obtain accurate fundamental frequency prediction values (that is, the model training effect is good), then its training samples need to cover audio data in the range of 60Hz to 1200Hz. That is to say, the training sample set for the fundamental frequency prediction model learning in the embodiments of the present application needs to include multiple training samples covering multiple pitch distribution ranges, and the training samples in each pitch distribution range are relatively evenly distributed. For example, the training sample set includes 12 training samples, among which 4 samples correspond to the treble distribution range, 4 samples correspond to the bass distribution range, and 4 samples correspond to the midrange distribution range. The present application is not limited to only three pitch distribution ranges, and there can be multiple pitch distribution ranges.

[0150] Among them, obtaining the training sample set includes: obtaining a set of candidate sample audio data, where the candidate sample audio data includes m candidate sample audio data, and m is an integer greater than 0; obtaining the pitch distribution vector corresponding to each candidate sample audio data among the m candidate sample audio data; and determining the sample sampling probability corresponding to each candidate sample audio data among the m candidate sample audio data according to the m pitch distribution vectors and m. Finally, according to the sample sampling probability corresponding to each candidate sample audio data, the sample audio data included in the training sample set is determined from the m candidate sample audio data. Among them, the sample sampling probability can be understood as the probability value of each sample being used as a training sample. Generally speaking, the more the number of candidate sample audio data in a certain pitch distribution range, the smaller the sample sampling probability corresponding to the candidate sample audio data in that pitch distribution range. Correspondingly, the fewer the number of candidate sample audio data, the greater the sample sampling probability corresponding to the candidate sample audio data in that pitch distribution range. The advantage of doing this is that it can make the pitch distribution of the sample audio data included in the training sample set for the fundamental frequency prediction model training balanced, thereby making the model training effect better and more universal.

[0151] Exemplarily, since the occurrence frequency of humans is between 60HZ and 1200HZ, the training sample set needs to cover 60HZ to 1200HZ to make the training data pitch uniform, so that the finally trained predicted fundamental frequency model can accurately predict the fundamental frequency sequence. For each of the m candidate sample audio data, each candidate sample audio data is analyzed to obtain a pitch distribution vector n corresponding to each candidate sample audio data. The dimension of the pitch distribution vector n can be 1200 dimensions (generally speaking, the choice of dimension is based on the maximum frequency to be covered). Among them, the pitch distribution vector n corresponding to any candidate sample audio data is used to represent the number of frames of the candidate sample audio data at each pitch, or understood as the distribution at each pitch. Taking a 1600-millisecond candidate sample audio data as an example, assuming that the length of each frame is 16 milliseconds, then the 1600-millisecond candidate sample audio data includes 1000 frames. If there are 0 frames distributed at 60HZ among these 1000 frames, the value of the pitch distribution vector n of the 1600-millisecond candidate sample audio data at 60HZ is 0. If there are 10 frames distributed at 300HZ among these 1000 frames, the value of the pitch distribution vector n of the 1600-millisecond candidate sample audio data at 300HZ is 10, etc. The corresponding values for other pitches will not be illustrated one by one here. Therefore, for the m candidate sample audio data, an m*n pitch distribution matrix A can be obtained. To achieve pitch distribution balance, solve the sample sampling probability corresponding to the m audio sample distributions:

[0152] A T X = B, s.t. X i > 0

[0153] B i = sum(A) / m

[0154] Among them, A is an m*n pitch distribution matrix, X is the sample sampling probability corresponding to the audio sample distribution, B is a vector with dimension n*1, and each element of the vector is equal.

[0155] Among them, obtain the sample fundamental frequency sequence and sample spectral envelope corresponding to each sample audio data in the training sample set, and determine the sample vuv sequence and sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence. The method of obtaining the sample fundamental frequency sequence and sample spectral envelope corresponding to each sample audio data is the same as the method described above. The method of determining the sample vuv sequence corresponding to the sample audio data according to the sample fundamental frequency sequence is also the same as the method of converting the fundamental frequency sequence into the vuv sequence described above, and will not be elaborated here.

[0156] Among them, determining the sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence includes: calculating the absolute value of the difference between each adjacent sample fundamental frequency value in the sample fundamental frequency sequence; if there are multiple consecutive absolute values of the differences greater than or equal to a preset threshold, determining at least one fundamental frequency segmentation boundary of the sample fundamental frequency sequence; segmenting the sample fundamental frequency sequence according to at least one fundamental frequency segmentation boundary to obtain multiple fundamental frequency segmented sequences; performing smoothing processing on the multiple fundamental frequency segmented sequences, and splicing the sequences obtained after the smoothing processing to be used as the sample note sequence corresponding to the sample audio data. That is to say, the specific method of determining the sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence can be: first, perform differencing on the consecutive sample fundamental frequency values in the sample fundamental frequency sequence. Among them, when the absolute value of multiple consecutive difference values is greater than or equal to a certain preset threshold, the multiple sample fundamental frequency values corresponding to the multiple consecutive difference values are determined to be the transition stage of the fundamental frequency sequence, and the middle position of the transition stage is defined as the fundamental frequency segmentation boundary. Then, segment the sample fundamental frequency sequence according to the obtained at least one fundamental frequency segmentation boundary to obtain multiple fundamental frequency segmented sequences. Finally, perform mean processing on the values of each fundamental frequency segmented sequence after removing the transition stage to be used as the sequence obtained after the smoothing processing of the fundamental frequency segmented sequence. Finally, splice the sequences obtained after the smoothing processing of each fundamental frequency segmented sequence to be used as the sample note sequence corresponding to the sample audio data.

[0157] In some feasible implementation manners, starting from the leftmost end of the fundamental frequency sequence, that is, starting from the first fundamental frequency value of the fundamental frequency sequence. Calculate the absolute value of the difference between the second fundamental frequency value and the first fundamental frequency value, then calculate the absolute value of the difference between the third fundamental frequency value and the second fundamental frequency value, and so on, calculate the absolute value of the difference between each adjacent fundamental frequency value included in the fundamental frequency sequence. When the absolute value of the difference between a certain fundamental frequency value and the adjacent fundamental frequency value is greater than the first threshold, the midpoint of this consecutive fundamental frequency value is determined as the fundamental frequency segmentation boundary.

[0158] Exemplarily, as Figure 5 shown, Figure 5501 is a continuous fundamental frequency sequence. Starting from the leftmost end (point e) of 501 and going up to point a in the fundamental frequency sequence 501, the absolute value of the difference between adjacent fundamental frequency values is less than the first threshold. Starting from point a, the absolute value of the difference between the fundamental frequency value corresponding to a + 1 and the fundamental frequency value corresponding to a is greater than the first threshold, and the absolute value of the difference between the fundamental frequency value corresponding to a and the fundamental frequency value corresponding to a - 1 is less than the first threshold. This continues until the absolute value of the difference between the fundamental frequency value corresponding to b and the fundamental frequency value corresponding to b - 1 is greater than the first threshold, and the absolute value of the difference between the fundamental frequency value corresponding to b + 1 and the fundamental frequency value corresponding to b is less than the first threshold. Thus, the ab segment between point a and point b is defined as the first transition stage of the fundamental frequency sequence 501 of the fundamental frequency sequence, and the middle position i (i = (a + b) / 2) of this first transition stage is defined as the fundamental frequency segmentation boundary. Similarly, the cd segment is the second transition stage of the fundamental frequency sequence 501 of the fundamental frequency sequence, and the middle position j (j = (c + d) / 2) of this second transition stage is defined as another fundamental frequency segmentation boundary of the fundamental frequency sequence 501.

[0159] Based on the two fundamental frequency segmentation boundaries of the fundamental frequency sequence 501 obtained from the above description, the fundamental frequency sequence 501 is further segmented and smoothed to obtain the corresponding note sequence 502 of the fundamental frequency sequence 501. Among them, it can also be known from the above description that i and j are the two fundamental frequency segmentation boundaries of the note sequence 502, the ab segment is the first transition stage, and the cd segment is the second transition stage. According to the two fundamental frequency segmentation boundaries (i and j) of the note sequence 502, the note sequence 502 is divided into three segments: the ei segment, the ij segment, and the jf segment, thus completing the segmentation process. After completing the segmentation process, the note sequence 502 is smoothed. The smoothing process includes two steps: removing the transition stage and taking the average value of the values in each fundamental frequency segmented sequence after removing the transition stage. The step of removing the transition stage includes: deleting the fundamental frequency values in the ab segment (the ab segment includes: the ai segment and the ib segment) and the cd segment (the cd segment includes: the cj segment and the jd segment). The step of taking the average value includes: respectively calculating the average value of the fundamental frequency values in the ea segment, the bc segment, and the df segment, namely the ea fundamental frequency average value, the bc fundamental frequency average value, and the df fundamental frequency average value. Use the ea fundamental frequency average value to replace the missing values in the ai segment, use the bc fundamental frequency average value to replace the missing values in the ib segment and the cj segment, and use the df fundamental frequency average value to replace the missing values in the jd segment. Thus, three sequences ei, ij, and jf are obtained, and finally these three sequences are concatenated to obtain the note sequence 502.

[0160] Exemplarily, for a clearer understanding of the fundamental frequency prediction method, the following will be further combined with Figure 6 , such as Figure 6It is another flowchart of the fundamental frequency prediction method provided by the embodiments of the present application. After feature extraction is performed on the audio data to be processed, two feature sequences are obtained: the fundamental frequency sequence and the spectral envelope. The extraction of the fundamental frequency sequence can adopt one of algorithms such as the harvest algorithm, pYin algorithm, or DIO algorithm of the world vocoder, which is not limited here. The fundamental frequency is the frequency of the first peak of the spectrum of the audio data to be processed and is the human perception of the fundamental frequency of the note. The spectral envelope is the resonance generated when the sound wave generated by vocal cord vibration passes through the vocal tract composed of the oral cavity, nasal cavity, etc. The spectral envelope is what we commonly call timbre. Therefore, by extracting the fundamental frequency sequence and the spectral envelope, it can be obtained whether the singer of the audio data to be processed is in tune (that is, commonly known as out of tune), and the timbre characteristics of the singer of the audio data to be processed are also obtained. Then, feature transformation is performed on the fundamental frequency sequence to obtain the vuv sequence, where the vuv sequence determines which parts of the audio data to be processed have a fundamental frequency and which parts do not. Ensure that the positions of the fundamental frequency values included in the predicted fundamental frequency sequence finally predicted are the same as those of the fundamental frequency sequence of the audio data to be processed, and thus it can be said that the rhythm of the predicted fundamental frequency sequence finally predicted is the same as that of the fundamental frequency sequence of the audio data to be processed. Finally, the target note sequence (the singer of the target note sequence is the same female / male voice as the singer of the audio data to be processed and is the same song) is put into the fundamental frequency prediction model together with the vuv sequence and the spectral envelope sequence to obtain the predicted fundamental frequency sequence.

[0161] Exemplarily, for a clearer understanding of the fundamental frequency prediction method, the following will be further combined with Figure 6 , refer to Figure 7 , Figure 7 It is a flowchart of the model training process in the fundamental frequency prediction method provided by the embodiments of the present application. Compare Figure 6 , Figure 7 Similarly, two features are also extracted from the sample audio data. However, compared with Figure 6 , Figure 7 the sample note sequence in Figure 6 is obtained through two steps of feature extraction and feature transformation from the sample audio data. As Figure 7 that is to say, the note sequence is an externally added note sequence when the present application is applied, while Figure 7 that is to say, the sample note sequence during the training of the present application is the note sequence of the sample audio data. By inputting the sample note sequence, the sample vuv sequence, and the sample spectral envelope as parameters into the initial fundamental frequency prediction model, the sample fundamental frequency prediction sequence can be obtained. The sample fundamental frequency prediction sequence output by the initial fundamental frequency prediction model is compared with the sample fundamental frequency sequence obtained by performing feature extraction on the sample audio data. The more similar the sample fundamental frequency prediction sequence output by the initial fundamental frequency prediction model is to the sample fundamental frequency sequence obtained by performing feature extraction on the sample audio data, the better the prediction model is, and the more accurate the fundamental frequency predicted by the prediction model is.

[0162] In the embodiment of the present application, the obtained sample fundamental frequency prediction sequence is compared with the sample fundamental frequency sequence obtained by feature extraction from the sample audio data, that is, there is no need to have any requirements on the song singing standard corresponding to the sample fundamental frequency sequence, nor does it require a large number of complex algorithms. Therefore, by adopting the embodiment of the present application, the complexity of fundamental frequency adjustment can be reduced, and the accuracy of the obtained fundamental frequency can be improved, so that the tuned work sounds more natural and has more personal singing characteristics of the user.

[0163] See also Figure 10 , Figure 10 is a structural diagram of a fundamental frequency prediction device provided in an embodiment of the present application. The audio processing device provided in an embodiment of the present application includes:

[0164] Among them Figure 10 1001 is an acquisition module 1001, which is used to acquire the fundamental frequency sequence and spectrum envelope of the audio data to be processed, and to acquire the target note sequence corresponding to the audio data to be processed. 1002 is a processing module 1002, which is used to determine the unvoiced and voiced vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence; and to determine the predicted fundamental frequency sequence corresponding to the audio data to be processed according to the spectrum envelope, vuv sequence and target note sequence in combination with the fundamental frequency prediction model.

[0165] In another implementation, the acquisition module 1001 is also used to acquire a training sample set and to acquire a sample fundamental frequency sequence and a sample spectrum envelope corresponding to each sample audio data in the training sample set; the processing module 1002 is also used to determine a sample vuv sequence and a sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence; the processing module 1002 is also used to input the sample spectrum envelope, the sample vuv sequence and the sample note sequence into an initial fundamental frequency prediction model to obtain a sample predicted fundamental frequency sequence output by the initial fundamental frequency prediction model; the processing module 1002 is also used to train the initial fundamental frequency prediction model according to the sample fundamental frequency sequence and the sample predicted fundamental frequency sequence corresponding to each sample audio data in the training sample set to obtain a fundamental frequency prediction model.

[0166] In another implementation, the acquisition module 1001 is also used to obtain a set of candidate sample audio data, the candidate sample audio data including m candidate sample audio data, where m is an integer greater than 0; the acquisition module 1001 is also used to obtain a pitch distribution vector corresponding to each candidate sample audio data in the m candidate sample audio data; the processing module 1002 is also used to determine the sample sampling probability corresponding to each candidate sample audio data in the m candidate sample audio data based on the m pitch distribution vectors and m; the processing module 1002 is also used to determine the sample audio data included in the training sample set from the m candidate sample audio data based on the sample sampling probability corresponding to each candidate sample audio data.

[0167] In another implementation, the processing module 1002 is further configured to calculate the absolute value of the difference between adjacent sample fundamental frequency values in the sample fundamental frequency sequence; the processing module 1002 is further configured to determine at least one fundamental frequency segmentation boundary of the sample fundamental frequency sequence if there are multiple consecutive absolute values of differences greater than or equal to a preset threshold; the processing module 1002 is further configured to segment the sample fundamental frequency sequence according to at least one fundamental frequency segmentation boundary to obtain multiple fundamental frequency segmented sequences; the processing module 1002 is further configured to smooth the multiple fundamental frequency segmented sequences and splice the sequences obtained after the smoothing process to be used as the sample note sequence corresponding to the sample audio data.

[0168] In another implementation, the obtaining module 1001 is further configured to obtain a set of candidate note sequences corresponding to the audio data to be processed, where the set of candidate note sequences includes at least two candidate note sequences; the processing module 1002 is further configured to determine a first parameter value based on each fundamental frequency value included in the fundamental frequency sequence; the processing module 1002 is further configured to determine a target note sequence from at least two candidate note sequences according to the first parameter value.

[0169] In another implementation, the processing module 1002 is further configured to determine the note average value corresponding to each candidate note sequence in at least two candidate note sequences, and determine a first parameter judgment value according to the note average values corresponding to at least two candidate note sequences; the processing module 1002 is further configured to determine the voice category of the audio data to be processed according to the first parameter judgment value and the first parameter value, where the voice category includes female voice or male voice; the processing module 1002 is further configured to determine at least one candidate note sequence corresponding to the voice category from at least two candidate note sequences according to the voice category of the audio data to be processed; the processing module 1002 is further configured to determine one candidate note sequence from at least one candidate note sequence corresponding to the voice category as the target note sequence.

[0170] In another implementation, the processing module 1002 is further configured to traverse each fundamental frequency value in the multiple fundamental frequency values included in the fundamental frequency sequence; the processing module 1002 is further configured to set the vuv value corresponding to the fundamental frequency value equal to 0 to a first preset value, and set the vuv value corresponding to the fundamental frequency value not equal to 0 to a second preset value to obtain the vuv sequence corresponding to the audio data to be processed.

[0171] In another implementation, the processing module 1002 is further configured to perform frame segmentation and windowing processing on the audio data to be processed, obtaining N first segmented signals, where N is an integer greater than 0; the processing module 1002 is further configured to respectively perform filtering processing on the first segmented signals through M low-pass filters with different cut-off frequencies, obtaining M filtered signals corresponding to the first segmented signals, where M is an integer greater than 1; the processing module 1002 is further configured to determine one cut-off frequency from the M cut-off frequencies as the fundamental frequency value corresponding to the first segmented signal according to the period information of the M filtered signals; the processing module 1002 is further configured to generate a fundamental frequency sequence corresponding to the audio data to be processed according to the fundamental frequency values corresponding to the N first segmented signals.

[0172] In another implementation, the acquisition module 1001 is further configured to acquire the power spectrum of the audio data to be processed; the processing module 1002 is further configured to perform inverse Fourier transform on the power spectrum of the audio data to be processed, obtaining the cepstrum corresponding to the power spectrum of the audio data to be processed; the processing module 1002 is further configured to perform filtering processing on the cepstrum based on a low-pass filter with a preset cut-off frequency to obtain the spectral envelope corresponding to the audio data to be processed.

[0173] In the embodiments of the present application, the fundamental frequency prediction device puts the acquired fundamental frequency sequence, spectral envelope, target note sequence, and vuv sequence into a prediction model to obtain a predicted fundamental frequency sequence, thereby achieving the effect of pitch correction. By adopting the embodiments of the present application, the complexity of fundamental frequency adjustment can be reduced, and the accuracy of the acquired fundamental frequency can be improved. Furthermore, the pitch-corrected work sounds more natural and has more personal singing characteristics of the user.

[0174] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of the terminal device provided by the embodiments of the present application. As Figure 11 shown, the terminal device in this embodiment may include: one or more processors 1101, one or more memories 1102, and one or more transceivers 1103. The above-mentioned processors 1101, memories 1102, and transceivers 1103 are connected through a bus 1104. The memory 1102 is used to store a computer program, and the computer program includes program instructions. The processor 1101 is configured to execute the program instructions stored in the memory 1102 to execute the processes described in steps S201 to S203 in the above-mentioned embodiments, and perform the following operations:

[0175] In one implementation, acquire the fundamental frequency sequence and spectral envelope of the audio data to be processed, and acquire the target note sequence corresponding to the audio data to be processed, where the target note sequence is a pre-stored note sequence associated with the audio data to be processed;

[0176] Determine the vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence;

[0177] According to the spectral envelope, vuv sequence, and target note sequence, a predicted fundamental frequency sequence corresponding to the audio data to be processed is determined in combination with a fundamental frequency prediction model, where the fundamental frequency prediction model is trained by a sample fundamental frequency sequence corresponding to sample audio data, a sample spectral envelope, and a sample vuv sequence and a sample note sequence obtained by converting the sample fundamental frequency sequence.

[0178] It should be understood that in some feasible embodiments, the above-mentioned processor 501 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc. The memory 502 may include a read-only memory and a random access memory, and provide instructions and data to the processor 501. A part of the memory 502 may also include a non-volatile random access memory. For example, the memory 502 may also store information about the device type.

[0179] In specific implementation, the above-mentioned terminal device may execute the implementation manners provided in each step as described above through its built-in various functional modules. Specifically, reference may be made to the implementation manners provided in each step as described above, which will not be elaborated herein. Figures 1 to 9 in each step as described above, specifically, reference may be made to the implementation manners provided in each step as described above, which will not be elaborated herein.

[0180] In the embodiments of the present application, the terminal device puts the obtained fundamental frequency sequence, spectral envelope, target note sequence, and vuv sequence into the prediction model to obtain a predicted fundamental frequency sequence, thereby achieving the effect of pitch correction. By adopting the embodiments of the present application, the complexity of fundamental frequency adjustment can be reduced, and the accuracy of the obtained fundamental frequency can be improved, so that the pitch-corrected work sounds more natural and has more personal singing characteristics of the user.

[0181] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, they implement Figure 2 the fundamental frequency prediction method provided in each step as described above. Specifically, reference may be made to the implementation manners provided in each step as described above, which will not be elaborated herein.

[0182] The above computer-readable storage medium may be the internal storage unit of the recommendation model training device provided in any of the foregoing embodiments or the above terminal device, such as the hard disk or memory of an electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0183] The terms "first", "second", "third", "fourth", etc. in the claims, the description and the drawings of this application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0184] Reference to "embodiments" in this text means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of this application. The phrase is presented at various places in the specification not necessarily all referring to the same embodiment, nor are they independent or alternative embodiments mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the specification and claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0185] The methods and related devices provided by the embodiments of the present application are described with reference to the method flowcharts and / or structural schematic diagrams provided by the embodiments of the present application. Specifically, each process and / or block of the method flowchart and / or structural schematic diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in the process Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in the process Figure 1 one process or multiple processes and / or structural schematic Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in the process Figure 1 one process or multiple processes and / or structural schematic one block or multiple blocks.

Claims

1. A fundamental frequency prediction method, characterized in that, Including: Obtaining a fundamental frequency sequence and a spectral envelope of the audio data to be processed, and obtaining a target note sequence corresponding to the audio data to be processed, where the target note sequence is a pre-stored note sequence associated with the audio data to be processed; Determining a voiceless / voiced (V / UV) sequence corresponding to the audio data to be processed according to the fundamental frequency sequence; Inputting the spectral envelope, the V / UV sequence, and the target note sequence into a pre-trained fundamental frequency prediction model to determine a predicted fundamental frequency sequence corresponding to the audio data to be processed; Among them, training the fundamental frequency prediction model includes: Obtaining a training sample set, a sample fundamental frequency sequence and a sample spectral envelope corresponding to each sample audio data in the training sample set, and determining a sample V / UV sequence and a sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence; Inputting the sample spectral envelope, the sample V / UV sequence, and the sample note sequence into an initial fundamental frequency prediction model to obtain a sample predicted fundamental frequency sequence output by the initial fundamental frequency prediction model; Training the initial fundamental frequency prediction model according to the sample fundamental frequency sequence and the sample predicted fundamental frequency sequence corresponding to each sample audio data in the training sample set to obtain the fundamental frequency prediction model.

2. The method according to claim 1, characterized in that, The obtaining of the training sample set includes: Obtaining a set of candidate sample audio data, where the candidate sample audio data includes m candidate sample audio data, and m is an integer greater than 0; Obtaining a pitch distribution vector corresponding to each candidate sample audio data in the m candidate sample audio data; Determining a sample sampling probability corresponding to each candidate sample audio data in the m candidate sample audio data according to the m pitch distribution vectors and the m; Determining the sample audio data included in the training sample set from the m candidate sample audio data according to the sample sampling probabilities corresponding to the respective candidate sample audio data.

3. The method according to claim 1, characterized in that, The determining of the sample note sequence corresponding to the sample audio data according to the sample fundamental frequency sequence includes: Calculating the absolute value of the difference between adjacent sample fundamental frequency values in the sample fundamental frequency sequence; If there are continuously multiple absolute values of differences greater than or equal to a preset threshold, determining at least one fundamental frequency segmentation boundary of the sample fundamental frequency sequence; Segmenting the sample fundamental frequency sequence according to the at least one fundamental frequency segmentation boundary to obtain a plurality of fundamental frequency segmented sequences; Smoothing the plurality of fundamental frequency segmented sequences and splicing the sequences obtained after smoothing to be used as the sample note sequence corresponding to the sample audio data.

4. The method according to claim 1, characterized in that, The obtaining of the target note sequence corresponding to the audio data to be processed includes: Obtaining a set of candidate note sequences corresponding to the audio data to be processed, where the set of candidate note sequences includes at least two candidate note sequences; Determining a first parameter value based on each fundamental frequency value included in the fundamental frequency sequence; Determining a target note sequence from the at least two candidate note sequences according to the first parameter value.

5. The method according to claim 4, characterized in that, The at least two candidate note sequences include at least one candidate note sequence corresponding to a female voice and at least one candidate note sequence corresponding to a male voice; Determining a target note sequence from the at least two candidate note sequences according to the first parameter value includes: Determining the note average value corresponding to each candidate note sequence in the at least two candidate note sequences, and determining a first parameter judgment value according to the note average values corresponding to the at least two candidate note sequences; Determining the voice category of the audio data to be processed according to the first parameter judgment value and the first parameter value, where the voice category includes female voice or male voice; Determining at least one candidate note sequence corresponding to the voice category from the at least two candidate note sequences according to the voice category of the audio data to be processed; Determining one candidate note sequence from the at least one candidate note sequence corresponding to the voice category as the target note sequence.

6. The method according to claim 1, characterized in that, Determining the vuv sequence corresponding to the audio data to be processed according to the fundamental frequency sequence includes: Traversing each fundamental frequency value included in the fundamental frequency sequence; Setting the vuv value corresponding to the fundamental frequency value equal to 0 to a first preset value, and setting the vuv value corresponding to the fundamental frequency value not equal to 0 to a second preset value to obtain the vuv sequence corresponding to the audio data to be processed.

7. The method according to any one of claims 1-6, characterized in that, Obtaining the fundamental frequency sequence of the audio data to be processed includes: Performing frame addition and windowing processing on the audio data to be processed to obtain N first frame signals, where N is an integer greater than 0; Filtering the first frame signals respectively through M low-pass filters with different cut-off frequencies to obtain M filtered signals corresponding to the first frame signals, where M is an integer greater than 1; Determining one cut-off frequency from the M cut-off frequencies as the fundamental frequency value corresponding to the first frame signal according to the period information of the M filtered signals; Generating the fundamental frequency sequence corresponding to the audio data to be processed according to the fundamental frequency values corresponding to the N first frame signals.

8. The method according to any one of claims 1-6, characterized in that, Obtaining the spectral envelope of the audio data to be processed includes: Obtaining the power spectrum of the audio data to be processed; Performing inverse Fourier transform on the power spectrum of the audio data to be processed to obtain the cepstrum corresponding to the power spectrum of the audio data to be processed; Filtering the cepstrum through a low-pass filter with a preset cut-off frequency to obtain the spectral envelope corresponding to the audio data to be processed.

9. A terminal device, characterized in that, Including a processor and a memory, the processor and the memory are connected to each other; The memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by the processor, the processor is caused to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Audio file processing method and device

    CN104091599A

  • Parameter speech synthesis method and system

    WO2013020329A1