Vibrato Detection Method, Computer Device, Storage Medium, and Computer Program Product
The method enhances tremolo detection in vocal signals by segmenting lyrics and filtering noise, improving accuracy through base frequency analysis.
Patent Information
- Application Number
- CN202211178090.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-23
AI Technical Summary
In traditional methods, by detecting the zero-crossing rate of the vocal signal, the accuracy of tremolo is low and is easily disturbed by noise, resulting in inaccurate detection.
The lyrics are sliced to the fundamental frequency sequence of the voice to be detected, filtered to reduce noise, determine the fluctuation amplitude and frequency, and determine whether it is a vibrato based on the preset conditions.
It improves the accuracy and robustness of vibrato detection, can avoid the influence of vocal differences and melody trend noise, and achieves finer-grained vibrato judgment.
Smart Images

Figure CN115565553B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of audio signal processing, and in particular, to a vibrato detection method, a computer device, a storage medium, and a computer program product. Background Art
[0002] Vibrato is an important vocal performance technique in the field of music, and it is usually demonstrated using musical instruments or through singing. Due to individual differences, the vibrato demonstrated by each person in singing is also different.
[0003] In traditional techniques, the vibrato in the human voice signal is often judged by detecting the zero-crossing rate of the human voice signal. However, there are often noise points in the actual human voice signal, which are likely to interfere with the signal fluctuation of the human voice signal, resulting in a low accuracy of the method for detecting vibrato based on the zero-crossing rate of the human voice signal. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a vibrato detection method, a computer device, a computer-readable storage medium, and a computer program product that can improve the accuracy of vibrato detection of the user's dry voice.
[0005] In a first aspect, this application provides a vibrato detection method. The method includes:
[0006] Performing lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected;
[0007] Performing filtering processing on each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment;
[0008] Determining the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment;
[0009] Obtaining the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, where the vibrato detection result of each noise-reduced fundamental frequency segment is the target vibrato detection result of the dry voice to be detected.
[0010] In one embodiment, performing filtering processing on each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment includes:
[0011] Performing simulation processing on the filter according to the preset bandwidth range and preset filter order to obtain the polynomial coefficients of the filter;
[0012] Performing band-pass filtering processing on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
[0013] In one embodiment, performing band-pass filtering on each of the fundamental frequency segments according to the polynomial coefficients of the filter to obtain a noise-reduced fundamental frequency segment corresponding to each of the fundamental frequency segments, includes:
[0014] Performing forward filtering on each of the fundamental frequency segments according to the polynomial coefficients of the filter to obtain a forward noise-reduced fundamental frequency segment corresponding to each of the fundamental frequency segments;
[0015] Performing backward filtering on each forward noise-reduced fundamental frequency segment to obtain a backward noise-reduced fundamental frequency segment corresponding to each forward noise-reduced fundamental frequency segment;
[0016] Performing a flipping process on each backward noise-reduced fundamental frequency segment to obtain a noise-reduced fundamental frequency segment corresponding to each of the fundamental frequency segments.
[0017] In one embodiment, performing band-pass filtering on each of the fundamental frequency segments according to the polynomial coefficients of the filter to obtain a noise-reduced fundamental frequency segment corresponding to each of the fundamental frequency segments, includes:
[0018] Performing backward filtering on each of the fundamental frequency segments according to the polynomial coefficients of the filter to obtain a backward noise-reduced fundamental frequency segment corresponding to each of the fundamental frequency segments;
[0019] Performing backward filtering on each backward noise-reduced fundamental frequency segment again to obtain a noise-reduced fundamental frequency segment corresponding to each of the fundamental frequency segments.
[0020] In one embodiment, performing lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain fundamental frequency segments corresponding to respective lyrics in the dry voice to be detected, includes:
[0021] Performing speech recognition on the dry voice to be detected to obtain respective lyrics corresponding to the dry voice to be detected;
[0022] Segmenting the fundamental frequency sequence of the dry voice to be detected according to the respective lyrics corresponding to the dry voice to be detected to obtain fundamental frequency segments corresponding to the respective lyrics.
[0023] In one embodiment, before performing lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain fundamental frequency segments corresponding to respective lyrics in the dry voice to be detected, further includes:
[0024] Obtaining the fundamental frequency sampling rate of the dry voice to be detected according to the frame interval;
[0025] Performing fundamental frequency extraction processing on the dry voice to be detected according to the fundamental frequency sampling rate to obtain an initial fundamental frequency sequence of the dry voice to be detected; the unit of the initial fundamental frequency sequence is frequency;
[0026] According to the mapping relationship between frequency and musical notes, perform unit conversion processing on the initial fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency sequence of the dry voice to be detected; the unit of the fundamental frequency sequence is musical notes.
[0027] In one embodiment, determining the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment includes:
[0028] According to the peak spacing and valley spacing in each noise-reduced fundamental frequency segment, obtain the fluctuation period of each noise-reduced fundamental frequency segment;
[0029] According to the fluctuation period of each noise-reduced fundamental frequency segment and the fundamental frequency sampling rate of the dry voice to be detected, obtain the fluctuation frequency of each noise-reduced fundamental frequency segment.
[0030] In one embodiment, obtaining the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition includes:
[0031] According to the peaks and valleys in the fluctuation amplitude of each noise-reduced fundamental frequency segment, obtain the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within the preset amplitude range;
[0032] According to whether the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within the preset amplitude range meet the preset vibrato condition, and whether the fluctuation frequency meets the preset vibrato condition, obtain the vibrato detection result of each noise-reduced fundamental frequency segment.
[0033] In one embodiment, obtaining the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within the preset amplitude range meet the preset vibrato condition, and whether the fluctuation frequency meets the preset vibrato condition includes:
[0034] For each noise-reduced fundamental frequency segment, when the number of peaks within the preset amplitude range of the noise-reduced fundamental frequency segment meets the preset peak number condition, the number of valleys meets the preset valley number condition, and the fluctuation frequency meets the preset fluctuation frequency condition, confirm that the noise-reduced fundamental frequency segment belongs to vibrato.
[0035] In a second aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0036] Perform lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected;
[0037] Filter each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment;
[0038] Determine the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment;
[0039] Obtain the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, where the vibrato detection result of each noise-reduced fundamental frequency segment is the vibrato detection result of the dry voice to be detected.
[0040] In a third aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0041] Perform lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected;
[0042] Filter each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment;
[0043] Determine the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment;
[0044] Obtain the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, where the vibrato detection result of each noise-reduced fundamental frequency segment is the vibrato detection result of the dry voice to be detected.
[0045] In a fourth aspect, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:
[0046] Perform lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected;
[0047] Filter each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment;
[0048] Determine the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment;
[0049] Obtain the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, where the vibrato detection result of each noise-reduced fundamental frequency segment is the vibrato detection result of the dry voice to be detected.
[0050] The above-mentioned vibrato detection method, computer device, storage medium, and computer program product perform lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected, obtaining fundamental frequency segments corresponding to each lyric in the dry voice to be detected; perform filtering processing on each fundamental frequency segment to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; determine the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment; and obtain the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, where the vibrato detection result of each noise-reduced fundamental frequency segment is the vibrato detection result of the dry voice to be detected, realizing the vibrato detection of the dry voice to be detected. By performing vibrato detection on the fundamental frequency sequence of the dry voice to be detected, the problem of vibrato detection failure caused by human voice differences can be avoided, and by performing lyric segmentation on the fundamental frequency sequence of the dry voice to be detected, a finer-grained judgment of the vibrato of the dry voice to be detected can be made to improve the accuracy of vibrato detection. Additionally, by performing filtering processing on the fundamental frequency segments, the influence of melody trends and noise on vibrato detection can be avoided, greatly improving the accuracy and robustness of vibrato detection. Description of the Drawings
[0051] Figure 1 It is an application environment diagram of the vibrato detection method in an embodiment;
[0052] Figure 2 It is a flowchart of the vibrato detection method in an embodiment;
[0053] Figure 3 It is a flowchart of the steps to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment in an embodiment;
[0054] Figure 4 It is a schematic diagram of the comparison result of the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment in an embodiment;
[0055] Figure 5 It is a schematic diagram of the fundamental frequency segments corresponding to the three lyrics "zài", "yī", and "qǐ" in an embodiment;
[0056] Figure 6 It is a flowchart of the steps to confirm whether the noise-reduced fundamental frequency segment belongs to vibrato in an embodiment;
[0057] Figure 7 It is a flowchart of the vibrato detection method in another embodiment;
[0058] Figure 8 It is a flowchart of the vibrato detection method in yet another embodiment;
[0059] Figure 9 It is an internal structure diagram of a computer device in an embodiment. Detailed Implementation Manner
[0060] In order to make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0061] The vibrato detection method provided by the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 101 communicates with the server 102 through the network. The terminal 101 can obtain the dry voice of the user and can also provide relevant services for vibrato detection; the server 102 generally refers to a background system that provides relevant services for vibrato detection. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or can be placed in the cloud or on other network servers. Among them, the terminal 101 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 102 can be implemented by an independent server or a server cluster composed of multiple servers.
[0062] In one implementation, the terminal 101 obtains the dry voice to be detected, and the terminal 101 performs lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected; filters each fundamental frequency segment to obtain the denoised fundamental frequency segment corresponding to each fundamental frequency segment; determines the fluctuation amplitude and fluctuation frequency of each denoised fundamental frequency segment; and obtains the vibrato detection result of each denoised fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each denoised fundamental frequency segment meet the preset vibrato condition, where the vibrato detection result of each denoised fundamental frequency segment is the vibrato detection result of the dry voice to be detected. Therefore, the execution subject of the above vibrato detection method can be the terminal 101.
[0063] In one implementation, the above vibrato detection method can also be independently implemented based on the server 102. For example, the server 102 can obtain the dry voice to be detected from the background database and obtain the target vibrato detection result of the dry voice to be detected by executing the above vibrato detection method.
[0064] In one implementation, the above vibrato detection method can also be implemented based on the interaction between the terminal 101 and the server. For example, after obtaining the dry voice to be detected, the terminal 101 sends the dry voice to be detected to the server 102, and the server 102 obtains the target vibrato detection result of the dry voice to be detected by executing the above vibrato detection method. For another example, the server can obtain the dry voice from the background database and send the dry voice to be detected to the terminal 101. The terminal 101 obtains the target vibrato detection result of the dry voice to be detected by executing the above vibrato detection method.
[0065] As can be seen from the above, in this exemplary implementation, the execution subject of the above vibrato detection method can be the above terminal 101 or the server 102, and can also be applied to a system including the terminal 101 and the server 102 and implemented through the interaction between the terminal 101 and the server 102. The present disclosure does not limit this.
[0066] In one embodiment, as Figure 2 shown, a vibrato detection method is provided. Taking the case where the method is applied to the Figure 1 terminal as an example, the method includes the following steps:
[0067] Step S201: Perform lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected.
[0068] Wherein, the dry voice to be detected refers to the dry voice that needs to be subjected to vibrato detection. The dry voice refers to the pure human voice in the audio field that has not undergone any spatial property processing, post-processing, or processing. The fundamental frequency refers to the lowest frequency in the dry voice to be detected, and the fundamental frequency is used to determine the pitch of the sound. The fundamental frequency sequence refers to the data composed of multiple fundamental frequencies. The fundamental frequency segment refers to a segment of the fundamental frequency sequence.
[0069] Specifically, the terminal obtains the dry voice to be detected of the user; according to the frame interval, performs fundamental frequency extraction processing on the dry voice to be detected to obtain the fundamental frequency sequence of the dry voice to be detected; for example, according to the frame interval, performs fundamental frequency extraction processing on the dry voice to be detected through the probabilistic YIN (pYIN) technology to obtain the fundamental frequency sequence of the dry voice to be detected; for another example, according to the frame interval, performs fundamental frequency extraction processing on the dry voice to be detected through the CREPE (Convolutional Representation for Pitch Estimation) technology to obtain the fundamental frequency sequence of the dry voice to be detected; for yet another example, according to the frame interval, performs fundamental frequency extraction processing on the dry voice to be detected through the Harvest fundamental frequency extractor to obtain the fundamental frequency sequence of the dry voice to be detected. After the terminal obtains the fundamental frequency sequence of the dry voice to be detected, since the vibrato of the human voice often adheres to a certain word in the lyrics, the terminal can perform lyrics segmentation on the fundamental frequency sequence to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected, so as to more accurately detect the vibrato part in the dry voice to be detected.
[0070] It should be noted that the granularity of the lyrics for lyrics segmentation processing can be a single word, two words, or three words, and no specific limitation is made here. For example, the terminal obtains the dry voice of a song "The Little Road" sung by a user, then converts the dry voice into a fundamental frequency sequence, and according to each word of the lyrics in the song "The Little Road", segments the fundamental frequency sequence to obtain the fundamental frequency segments corresponding to each word of the lyrics.
[0071] Step S202, perform filtering processing on each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Among them, the noise-reduced fundamental frequency segment refers to the fundamental frequency segment obtained after filtering and noise reduction processing.
[0072] Step S203, determine the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment.
[0073] Step S204, according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, obtain the vibrato detection result of each noise-reduced fundamental frequency segment, where the vibrato detection result of each noise-reduced fundamental frequency segment is the vibrato detection result of the dry voice to be detected.
[0074] Among them, the fluctuation amplitude refers to the width of the change in the pitch curve of the noise-reduced fundamental frequency segment.
[0075] The vibrato detection result is used to indicate the vibrato judgment result of the dry voice to be detected under each lyric. For example, the lyrics of the dry voice to be detected include the three lyrics "at", "one", and "together", and the target vibrato detection result will indicate whether the dry voice corresponding to the lyric "at" is vibrato, whether the dry voice corresponding to the lyric "one" is vibrato, and whether the dry voice corresponding to the lyric "together" is vibrato.
[0076] Specifically, the fundamental frequency fluctuation range of human voice is generally between 4 Hz and 8 Hz. To avoid the interference of high-frequency noise and low-frequency melody trends on the fluctuation of the fundamental frequency sequence, the terminal inputs each fundamental frequency segment of the fundamental frequency sequence into a band-pass filter for filtering, and obtains a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment output by the band-pass filter; the fundamental frequency fluctuation amplitude of human voice is generally between 30 and 150 cents. The terminal can calculate the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment. For each noise-reduced fundamental frequency segment, according to the signal fluctuation of the noise-reduced fundamental frequency segment under the preset fluctuation frequency condition and within the preset amplitude range, the vibrato detection result of each noise-reduced fundamental frequency segment is obtained; according to the vibrato detection results of all noise-reduced fundamental frequency segments, the vibrato detection result of the dry voice to be detected is obtained.
[0077] In the above vibrato detection method, the fundamental frequency sequence of the dry voice to be detected is subjected to lyric segmentation processing to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected; each fundamental frequency segment is subjected to filtering processing to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment are determined; according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, the vibrato detection result of each noise-reduced fundamental frequency segment is obtained, where the vibrato detection result of each noise-reduced fundamental frequency segment is the vibrato detection result of the dry voice to be detected, realizing the vibrato detection of the dry voice to be detected. By performing vibrato detection on the fundamental frequency sequence of the dry voice to be detected, the problem of vibrato detection failure caused by human voice differences can be avoided, and by performing lyric segmentation on the fundamental frequency sequence of the dry voice to be detected, a finer-grained judgment of the vibrato of the dry voice to be detected can be made to improve the accuracy of vibrato detection. In addition, by filtering the fundamental frequency segments, the influence of melody trends and noise on vibrato detection can be avoided, greatly improving the accuracy and robustness of vibrato detection.
[0078] In one embodiment, as Figure 3 shown, step S202 above, performing filtering processing on each fundamental frequency segment to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment, specifically includes the following contents:
[0079] Step S301, according to the preset bandwidth range and preset filter order, perform simulation processing on the filter to obtain the polynomial coefficients of the filter.
[0080] Among them, the preset bandwidth range can be set to [4, 20] Hz. The filter refers to a filtering circuit used to eliminate frequencies outside a specific frequency range. The filter can be a Butterworth filter, and of course, it can also be other band-pass filters, such as Chebyshev filters and elliptic filters, etc. The preset filter order is used to determine the structure of the filter. The preset filter order can be the first order, the second order, the third order, etc. The larger the value of the preset filter order, the closer the filter structure is to the target bandwidth.
[0081] Specifically, the terminal constructs a filter with a band-pass structure according to the preset bandwidth range and the preset filter order; performs simulation processing on the filter to obtain the polynomial coefficients of the filter; among them, the polynomial coefficients include the numerator polynomial coefficients and the denominator polynomial coefficients.
[0082] Taking the Butterworth filter as an example, the preset bandwidth range is set to [4, 20] Hz, and the preset filter order is set to the second order. Then a second-order band-pass structure Butterworth filter is constructed, and its expression H(z) is as follows:
[0083]
[0084] Among them, z represents the frequency domain of the second-order band-pass structure Butterworth filter; the superscripts -1 and -2 of z respectively represent the order; b0, b1, and b2 are all numerator polynomial coefficients; a0, a1, and a2 are all denominator polynomial coefficients.
[0085] By performing simulation processing on the second-order band-pass structure Butterworth filter, the coefficients of the filter are obtained: b0 = 0.1122, b1 = 0, b2 = -0.1122, a0 = 1.0000, a1 = -1.7581, a2 = 0.7757.
[0086] Step S302, perform band-pass filtering on each base frequency segment according to the polynomial coefficients of the filter to obtain a noise-reduced base frequency segment corresponding to each base frequency segment.
[0087] Among them, the polynomial coefficients refer to the coefficients in the expression of the filter. The polynomial coefficients can also be regarded as the weights of the delay modules in the filter.
[0088] Specifically, after the terminal obtains the molecular polynomial coefficients, the denominator polynomial coefficients, and the filter, it performs band-pass filtering on each fundamental frequency segment through the filter to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Among them, the band-pass filtering includes zero-phase filtering. By performing zero-phase filtering on each fundamental frequency segment, the phase of the obtained noise-reduced fundamental frequency segment is the same as that of the original fundamental frequency segment, so that the noise-reduced fundamental frequency segment can maintain the change trend of the original fundamental frequency segment.
[0089] Taking the dry voice of a user's song "The Path" as an example, from 31.5s to 34.6s of the user's dry voice, the fundamental frequency sequences corresponding to the three lyrics "zài", "yī", and "qǐ" can be extracted. Figure 4 It is a schematic diagram of the comparison result of performing band-pass filtering on each fundamental frequency segment through the filter to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Figure 4 The three left-side graphs in it are the fundamental frequency sequences corresponding to the three lyrics "zài", "yī", and "qǐ". Figure 4 The three right-side graphs in it are the noise-reduced fundamental frequency segments corresponding to the fundamental frequency sequences of the three lyrics "zài", "yī", and "qǐ", and the ordinate represents the cents of the dry voice to be detected. It can be seen that Figure 4 the phases of the fundamental frequency sequences corresponding to the three lyrics "zài", "yī", and "qǐ" are the same as those of the noise-reduced fundamental frequency sequences. And due to the band-pass filtering, the noise in the noise-reduced fundamental frequency sequences is significantly less than that in the fundamental frequency sequences. For example, the noise between the abscissas of 120 and 140 in the fundamental frequency sequence of the lyric "zài" is filtered. Another example is that the noise between the abscissas of 100 and 140 in the fundamental frequency sequence of the lyric "yī" is filtered. And for another example, the noise between the abscissas of 0 and 50 in the fundamental frequency sequence of the lyric "qǐ" is filtered, making the audio curve on the right side smoother than that on the left side.
[0090] In this embodiment, the terminal performs simulation processing on the filter according to the preset bandwidth range and the preset filter order to obtain the polynomial coefficients of the filter, and performs band-pass filtering on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment, which can reduce the interference of low-frequency and high-frequency noises in each fundamental frequency segment on the fundamental frequency segment, so as to improve the accuracy of vibrato detection for the fundamental frequency segment.
[0091] In one embodiment, according to the polynomial coefficients of the filter, each fundamental frequency segment is subjected to band-pass filtering to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Specifically, it includes the following content: According to the polynomial coefficients of the filter, each fundamental frequency segment is subjected to forward filtering to obtain a forward noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; each forward noise-reduced fundamental frequency segment is subjected to reverse filtering to obtain a reverse noise-reduced fundamental frequency segment corresponding to each forward noise-reduced fundamental frequency segment; each reverse noise-reduced fundamental frequency segment is subjected to a flipping process to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
[0092] Among them, forward filtering refers to filtering the fundamental frequency segment in sequence; reverse filtering refers to flipping the fundamental frequency segment and then filtering it.
[0093] In practical applications, the terminal can perform FRR filtering on each fundamental frequency segment according to the polynomial coefficients of the filter. Specifically, the terminal filters each fundamental frequency segment in sequence according to the polynomial coefficients of the filter to obtain a forward noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; then the forward noise-reduced fundamental frequency segment is flipped to obtain a flipped noise-reduced fundamental frequency segment; then the flipped noise-reduced fundamental frequency segment is input into the filter for filtering to obtain a reverse noise-reduced fundamental frequency segment corresponding to each flipped noise-reduced fundamental frequency segment; each reverse noise-reduced fundamental frequency segment is flipped again to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Among them, the phase of each noise-reduced fundamental frequency segment is the same as the phase of the corresponding fundamental frequency segment.
[0094] For example, after the terminal obtains the molecular polynomial coefficients b0, b1, and b2, and the denominator polynomial coefficients a0, a1, and a2, the implementation process of performing FRR filtering on each fundamental frequency segment according to the polynomial coefficients of the filter can be represented by the following formula:
[0095]
[0096] Among them, y 1 (n) represents the forward noise-reduced fundamental frequency segment; y 2 (n) represents the flipped noise-reduced fundamental frequency segment; y 3 (n) represents the reverse noise-reduced fundamental frequency segment; y(n) represents the noise-reduced fundamental frequency segment.
[0097] In this embodiment, after the terminal obtains the polynomial coefficients of the filter, according to the polynomial coefficients of the filter, forward filtering is performed on each fundamental frequency segment to obtain a forward noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; reverse filtering is performed on each forward noise-reduced fundamental frequency segment to obtain a reverse noise-reduced fundamental frequency segment corresponding to each forward noise-reduced fundamental frequency segment; flipping processing is performed on each reverse noise-reduced fundamental frequency segment to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; so that the obtained noise-reduced fundamental frequency segment can maintain the phase of the original fundamental frequency segment while reducing the interference of low-frequency and high-frequency noises in each fundamental frequency segment to the fundamental frequency segment, thereby improving the accuracy of vibrato detection for the fundamental frequency segment.
[0098] In one embodiment, performing band-pass filtering on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment includes: performing reverse filtering on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain a reverse noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; performing reverse filtering on each reverse noise-reduced fundamental frequency segment again to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
[0099] In practical applications, the terminal can perform RRF filtering on each fundamental frequency segment according to the polynomial coefficients of the filter. Specifically, the terminal performs flipping processing on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain a first flipped fundamental frequency segment corresponding to each fundamental frequency segment; then each first flipped fundamental frequency segment is input into the filter for filtering to obtain a reverse noise-reduced fundamental frequency segment; flipping processing is performed on each reverse noise-reduced fundamental frequency segment again to obtain a second flipped fundamental frequency segment corresponding to each reverse noise-reduced fundamental frequency segment; the second flipped fundamental frequency segment is input into the filter for filtering again to obtain a noise-reduced fundamental frequency segment corresponding to each second flipped fundamental frequency segment, and this noise-reduced fundamental frequency segment is used as the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Among them, the phase of each noise-reduced fundamental frequency segment is the same as the phase of the corresponding fundamental frequency segment.
[0100] In this embodiment, after the terminal obtains the polynomial coefficients of the filter, according to the polynomial coefficients of the filter, reverse filtering is performed on each fundamental frequency segment to obtain a reverse noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; reverse filtering is performed on each reverse noise-reduced fundamental frequency segment again to obtain a noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; so that the obtained noise-reduced fundamental frequency segment can maintain the phase of the original fundamental frequency segment while reducing the interference of low-frequency and high-frequency noises in each fundamental frequency segment to the fundamental frequency segment, thereby improving the accuracy of vibrato detection for the fundamental frequency segment.
[0101] In one embodiment, in step S201 above, the fundamental frequency sequence of the dry voice to be detected is subjected to lyric segmentation processing to obtain fundamental frequency segments corresponding to each lyric in the dry voice to be detected, which specifically includes the following: performing speech recognition on the dry voice to be detected to obtain each lyric corresponding to the dry voice to be detected; and segmenting the fundamental frequency sequence of the dry voice to be detected according to each lyric corresponding to the dry voice to be detected to obtain fundamental frequency segments corresponding to each lyric.
[0102] Specifically, the terminal can perform speech recognition processing on the fundamental frequency sequence of the dry voice to be detected through Automatic Speech Recognition (ASR) technology to obtain each lyric corresponding to the dry voice to be detected; and then segment the fundamental frequency sequence according to a preset lyric granularity to obtain fundamental frequency segments corresponding to each lyric.
[0103] Taking a song "The Path" sung by a user as an example, the terminal performs speech recognition processing on the user's dry voice through Automatic Speech Recognition technology, and obtains that the user's dry voice contains the lyrics "zài", "yī", "qǐ" in the segment from 31.5 s to 34.6 s; assuming that the preset lyric granularity is a single character, segment the fundamental frequency sequence to obtain the fundamental frequency segment corresponding to the lyric "zài" in the fundamental frequency sequence, the fundamental frequency segment corresponding to the lyric "yī", and the fundamental frequency segment corresponding to the lyric "qǐ"; among them, the fundamental frequency segments corresponding to the three lyrics "zài", "yī", "qǐ" are as Figure 5 shown.
[0104] When the computing power and processing power of the terminal are limited, the terminal can also perform effective energy judgment on the fundamental frequency sequence of the dry voice to be detected through Voice Activity Detection (VAD) technology with lower computing power requirements to obtain the effective energy in the fundamental frequency sequence, and then use the qrc format lyric file of the dry voice to be detected to predict the start and end times of the lyrics in the fundamental frequency sequence carrying the effective energy to obtain the start and end times of the lyrics in the fundamental frequency sequence; and then segment the fundamental frequency sequence according to the preset lyric granularity and the start and end times of the lyrics in the fundamental frequency sequence to obtain fundamental frequency segments corresponding to each lyric; where the start and end times of the lyrics refer to the start time and end time of the user's voice when singing the lyric.
[0105] In this embodiment, by performing speech recognition on the dry voice to be detected to obtain each lyric corresponding to the dry voice to be detected, and then segmenting the fundamental frequency sequence of the dry voice to be detected according to each lyric corresponding to the dry voice to be detected to obtain fundamental frequency segments corresponding to each lyric, it is possible to perform more fine-grained vibrato detection on the vibrato of the dry voice to be detected, so as to improve the accuracy of vibrato detection.
[0106] In one embodiment, before performing lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected, it further includes: obtaining the fundamental frequency sampling rate of the dry voice to be detected according to the frame interval; performing fundamental frequency extraction processing on the dry voice to be detected according to the fundamental frequency sampling rate to obtain the initial fundamental frequency sequence of the dry voice to be detected; the unit of the initial fundamental frequency sequence is frequency; performing unit conversion processing on the initial fundamental frequency sequence of the dry voice to be detected according to the mapping relationship between frequency and notes to obtain the fundamental frequency sequence of the dry voice to be detected; the unit of the fundamental frequency sequence is notes.
[0107] Wherein, the frame interval refers to the time interval between the transmissions of audio frames in the dry voice to be detected. The fundamental frequency sampling rate refers to the number of samples extracted from the dry voice to be detected per second and composed into the fundamental frequency.
[0108] Specifically, the terminal obtains the frame interval for the dry voice to be detected, and then takes the reciprocal of this frame interval as the fundamental frequency sampling rate; furthermore, the terminal performs fundamental frequency extraction processing on the dry voice to be detected according to the frame interval and the fundamental frequency sampling rate through a fundamental frequency extraction tool (such as, probabilistic YIN technology, CREPE technology, and Harvest fundamental frequency extractor) to obtain the initial fundamental frequency sequence of the dry voice to be detected, wherein the unit of the initial fundamental frequency sequence is frequency.
[0109] Furthermore, to conform to the user's auditory habit, the frequency unit of the initial fundamental frequency sequence can be converted to the note unit. The terminal performs unit conversion processing on the initial fundamental frequency sequence of the dry voice to be detected according to the mapping relationship (or conversion relationship) between frequency and notes to obtain the fundamental frequency sequence of the dry voice to be detected, wherein the unit of the fundamental frequency sequence is notes. The mapping relationship between frequency and notes can be expressed by the following formula:
[0110]
[0111] Wherein, c represents the note number; f represents the frequency.
[0112] In practical applications, the frame interval can be set to 5 ms, then the fundamental frequency sampling rate of the fundamental frequency sequence is 1 / 5 ms = 1 / 0.005 s = 200 Hz, and the fundamental frequency sequence of the dry voice to be detected can be expressed as X(n), where n = 0, 1, 2, 3, …, N - 1, and N represents the total number of frames of the fundamental frequency sequence.
[0113] In this embodiment, the terminal obtains the fundamental frequency sampling rate of the dry sound to be detected according to the frame interval. On the one hand, the fundamental frequency sequence can be extracted from the dry sound to be detected using the fundamental frequency sampling rate. On the other hand, the fluctuation frequency of the fundamental frequency sequence can be obtained through the fundamental frequency sampling rate. Then, based on the fundamental frequency sampling rate, fundamental frequency extraction processing is performed on the dry sound to be detected to obtain the initial fundamental frequency sequence of the dry sound to be detected. Furthermore, according to the mapping relationship between frequency and notes, unit conversion processing is performed on the initial fundamental frequency sequence of the dry sound to be detected to obtain the fundamental frequency sequence of the dry sound to be detected, realizing the reasonable extraction of the fundamental frequency sequence of the dry sound to be detected and performing reasonable conversion on the unit of the fundamental frequency sequence, making the subsequent analysis of the fundamental frequency sequence more in line with the user's auditory effect.
[0114] In one embodiment, in step S203 above, determining the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment specifically includes the following: According to the peak-to-peak distance and valley-to-valley distance in each noise-reduced fundamental frequency segment, the fluctuation period of each noise-reduced fundamental frequency segment is obtained; according to the fluctuation period of each noise-reduced fundamental frequency segment and the fundamental frequency sampling rate of the dry sound to be detected, the fluctuation frequency of each noise-reduced fundamental frequency segment is obtained.
[0115] Among them, the peak-to-peak distance refers to the distance between two adjacent peaks; a peak refers to the maximum value of the wave amplitude within one wavelength range of the noise-reduced fundamental frequency segment. The valley-to-valley distance refers to the distance between two adjacent valleys; a valley refers to the minimum value of the wave amplitude within one wavelength range of the noise-reduced fundamental frequency segment.
[0116] Among them, the fluctuation frequency refers to the number of times the signal undergoes periodic changes per unit time.
[0117] Specifically, the terminal obtains the peak sequence in each noise-reduced fundamental frequency segment according to the amplitude values of each peak in each noise-reduced fundamental frequency segment, and obtains the valley sequence in each noise-reduced fundamental frequency segment according to the amplitude values of each valley in each noise-reduced fundamental frequency segment. Then, according to the peak sequence in each noise-reduced fundamental frequency segment, the peak-to-peak distance between two adjacent peaks and the valley-to-valley distance between two adjacent valleys in each noise-reduced fundamental frequency segment are determined. According to the average value of the sum of all peak-to-peak distances and valley-to-valley distances in each noise-reduced fundamental frequency segment, the fluctuation period of each noise-reduced fundamental frequency segment is obtained. In practical applications, the fluctuation period can be expressed by the following formula:
[0118]
[0119] Among them, is denoted as the fluctuation period; I represents the total number of peaks of the noise-reduced fundamental frequency segments; J represents the total number of valleys of the noise-reduced fundamental frequency segments; Δp(i) represents the spacing between the position of the i-th peak and the position of the (i - 1)-th peak in the peak sequence, that is, Δp(i) = pks(i) - pks(i - 1), where pks(i) represents the peak sequence; Δv(j) represents the spacing between the position of the j-th valley and the position of the (j - 1)-th valley in the valley sequence, that is, Δv(j) = vly(j) - vly(j - 1), where vly(j) represents the valley sequence.
[0120] Then, the terminal calculates the fluctuation frequency of each noise-reduced fundamental frequency segment by using the obtained fundamental frequency sampling rate of the dry sound to be detected (i.e., the fundamental frequency sampling rate of the fundamental frequency sequence) and the fluctuation period of each noise-reduced fundamental frequency segment. In practical applications, the fluctuation frequency can be represented by the following formula:
[0121]
[0122] where F represents the fluctuation frequency; fs represents the fundamental frequency sampling rate, and the fundamental frequency sampling rate of the fundamental frequency sequence can be 200 Hz.
[0123] In this embodiment, the terminal obtains the fluctuation period of each noise-reduced fundamental frequency segment according to the peak spacing and valley spacing in each noise-reduced fundamental frequency segment; and then obtains the fluctuation frequency of each noise-reduced fundamental frequency segment according to the fluctuation period of each noise-reduced fundamental frequency segment and the fundamental frequency sampling rate of the dry sound to be detected, realizing the reasonable acquisition of the fluctuation frequency of each noise-reduced fundamental frequency segment, so as to perform vibrato detection through the fluctuation frequency of each noise-reduced fundamental frequency segment in the subsequent steps.
[0124] In one embodiment, in the above step S203, according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition, the vibrato detection result of each noise-reduced fundamental frequency segment is obtained, which specifically includes the following content: according to the peaks and valleys in the fluctuation amplitude of each noise-reduced fundamental frequency segment, the number of peaks and valleys of each noise-reduced fundamental frequency segment within the preset amplitude range is obtained; according to whether the number of peaks and valleys of each noise-reduced fundamental frequency segment within the preset amplitude range meets the preset vibrato condition, and whether the fluctuation frequency meets the preset vibrato condition, the vibrato detection result of each noise-reduced fundamental frequency segment is obtained.
[0125] Among them, the preset amplitude range refers to the amplitude range of the tone curve of the noise-reduced fundamental frequency segment set.
[0126] Specifically, the terminal sets a preset amplitude range according to the fluctuation amplitude of the noise-reduced fundamental frequency segments; based on the values of the peaks and valleys in the fluctuation amplitude of each noise-reduced fundamental frequency segment, it determines whether each peak and each valley in the fluctuation amplitude of each noise-reduced fundamental frequency segment are within the preset amplitude range, and obtains the number of peaks and the number of valleys within the preset amplitude range for each noise-reduced fundamental frequency segment. It should be noted that the amplitude range of the conventional fundamental frequency is generally between 30 and 150 cents, that is, 0.3 semitones to 1.5 semitones. Since the fundamental frequency segments are filtered in this method, the preset amplitude range needs to be set according to the fluctuation amplitude of the fundamental frequency segments after band-pass filtering, that is, the preset amplitude range is set according to the fluctuation amplitude of the noise-reduced fundamental frequency segments.
[0127] For example, as Figure 4 shown, the preset amplitude range can be set to [30, 210] cents, that is, 0.3 semitones to 2.1 semitones, that is Figure 4 in the [0.3, 2.1] semitone range and [-2.1, -0.3] semitone range of the vertical axis in Figure 4 in the three figures on the right in Figure 4 "◆" represents valleys and "*" represents peaks, then Figure 4 the first figure on the right in Figure 4 contains 4 peaks and 4 valleys in the [0.3, 2.1] semitone range and [-2.1, -0.3] semitone range,
[0128] Furthermore, the terminal determines whether the number of peaks of each noise-reduced fundamental frequency segment within the preset amplitude range meets the preset peak number condition, whether the number of valleys meets the preset valley number condition, and whether the fluctuation frequency meets the preset fluctuation frequency condition. Then, the terminal determines the vibrato detection result corresponding to each noise-reduced fundamental frequency segment according to whether each noise-reduced fundamental frequency segment meets the above three conditions.
[0129] In this embodiment, the terminal obtains the number of peaks and valleys of each noise reduction fundamental frequency segment within a preset amplitude range based on the peaks and valleys in the fluctuation amplitude of each noise reduction fundamental frequency segment; and obtains the vibrato detection result of each noise reduction fundamental frequency segment according to whether the number of peaks and valleys of each noise reduction fundamental frequency segment within the preset amplitude range meets the preset vibrato condition and whether the fluctuation frequency meets the preset vibrato condition. By using this method, the preset amplitude range can be set according to the filtering degree of the noise reduction fundamental frequency segment, and then the vibrato detection result corresponding to the noise reduction fundamental frequency segment can be determined based on the number of peaks, valleys and fluctuation frequency of the noise reduction fundamental frequency segment within the preset amplitude range, realizing reasonable vibrato detection of the noise reduction fundamental frequency segment, rather than setting the preset amplitude range according to the conventional fundamental frequency fluctuation amplitude, thereby improving the accuracy of vibrato detection.
[0130] In one embodiment, obtaining the vibrato detection result of each noise reduction fundamental frequency segment according to whether the number of peaks and valleys of each noise reduction fundamental frequency segment within the preset amplitude range meets the preset vibrato condition and whether the fluctuation frequency meets the preset vibrato condition specifically includes the following content: for each noise reduction fundamental frequency segment, when the number of peaks of the noise reduction fundamental frequency segment within the preset amplitude range meets the preset peak number condition, the number of valleys meets the preset valley number condition, and the fluctuation frequency meets the preset fluctuation frequency condition, it is confirmed that the noise reduction fundamental frequency segment belongs to vibrato.
[0131] Among them, the preset peak number condition and the preset valley number condition respectively refer to the peak number range and valley number range of the fundamental frequency segment belonging to vibrato. The preset fluctuation frequency condition refers to the range of the fluctuation frequency of the fundamental frequency segment belonging to vibrato. It can be seen that the preset vibrato condition (including the preset peak number condition, the preset valley number condition and the preset fluctuation frequency condition) is a condition that conforms to the characteristics of vibrato.
[0132] Specifically, for each noise-reduced fundamental frequency segment, the terminal obtains the peak sequence, valley sequence, and fluctuation frequency in the fundamental frequency segment. First, according to a preset amplitude range, from the peak sequence and valley sequence of the noise-reduced fundamental frequency sequence, the peak sequence and valley sequence that meet the preset amplitude range are screened out, and the fundamental frequency fluctuations outside this preset amplitude range in the noise-reduced fundamental frequency sequence are regarded as invalid fluctuations. In the case where there is no peak sequence and valley sequence that meet the preset amplitude range in the peak sequence and valley sequence of the noise-reduced fundamental frequency sequence, it is confirmed that this noise-reduced fundamental frequency segment does not belong to a trill. In the case where there are peak sequences and valley sequences that meet the preset amplitude range in the peak sequence and valley sequence of the noise-reduced fundamental frequency sequence, it is continued to judge whether the number of peaks and the number of valleys within the preset amplitude range of this noise-reduced fundamental frequency sequence both meet the preset peak number condition and the preset valley number condition; in the case where at least one of the number of peaks and the number of valleys within the preset amplitude range of the noise-reduced fundamental frequency segment does not meet the preset peak number condition and the preset valley number condition, it is confirmed that this noise-reduced fundamental frequency segment does not belong to a trill; in the case where the number of peaks and the number of valleys within the preset amplitude range of the noise-reduced fundamental frequency segment both meet the preset peak number condition and the preset valley number condition, it is continued to judge whether the fluctuation frequency of this noise-reduced fundamental frequency sequence meets the preset fluctuation frequency condition; in the case where the fluctuation frequency of this noise-reduced fundamental frequency sequence does not meet the preset fluctuation frequency condition, it is confirmed that this noise-reduced fundamental frequency segment does not belong to a trill; in the case where the fluctuation frequency of this noise-reduced fundamental frequency sequence meets the preset fluctuation frequency condition, it is confirmed that this noise-reduced fundamental frequency segment belongs to a trill.
[0133] For example, Figure 6 is a schematic flowchart for confirming whether a noise-reduced fundamental frequency segment belongs to a trill. As Figure 6 shown, for each noise-reduced fundamental frequency segment, the terminal obtains the peak sequence pks(i), valley sequence vly(j), and fluctuation frequency F in the fundamental frequency segment; the preset amplitude range can be set to the [0.3, 2.1] semitone range and the [-2.1, -0.3] semitone range, and then it is judged whether the peak sequence pks(i) of the noise-reduced fundamental frequency sequence is greater than or equal to 0.3 semitones and less than or equal to 2.1 semitones, and whether the valley sequence vly(j) of the noise-reduced fundamental frequency sequence is greater than or equal to -2.1 semitones and less than or equal to -0.3 semitones; furthermore, it is judged whether the number of peaks (marked as np) of the peak sequence pks(i) of the noise-reduced fundamental frequency segment within the [0.3, 2.1] semitone range is greater than or equal to 2, and whether the number of valleys (marked as nv) of the valley sequence vly(j) of this noise-reduced fundamental frequency segment within the [-2.1, -0.3] semitone range is greater than or equal to 2; the preset fluctuation frequency condition can be set to the [4, 8] Hz range, and finally it is judged whether the fluctuation frequency F of the noise-reduced fundamental frequency sequence is within the [4, 8] Hz range; in the case where the results of the above three judgments are all yes, it is confirmed that this noise-reduced fundamental frequency segment belongs to a trill, otherwise, it is confirmed that this noise-reduced fundamental frequency segment does not belong to a trill.
[0134] For example, as can be seen from the first figure on the right in Figure 4 , the denoised fundamental frequency segment of the lyric "zài" contains 4 peaks and 4 valleys within the ranges of [0.3, 2.1] semitones and [-2.1, -0.3] semitones, and its fluctuation frequency F is 5.5. After judgment through the Figure 6 process schematic diagram, it is obtained that the lyric "zài" belongs to vibrato; as can be seen from the second figure on the right in Figure 4 , the denoised fundamental frequency segment of the lyric "yī" contains 1 peak and 3 valleys within the ranges of [0.3, 2.1] semitones and [-2.1, -0.3] semitones. After judgment through the Figure 6 process schematic diagram, it is obtained that the lyric "yī" does not belong to vibrato; Figure 4 The third figure on the right in Figure 6 contains 5 peaks and 6 valleys within the ranges of [0.3, 2.1] semitones and [-2.1, -0.3] semitones, and its fluctuation frequency F is 4.8. After judgment through the
[0135] process schematic diagram, it is obtained that the lyric "qǐ" belongs to vibrato.
[0136] In one embodiment, as shown in Figure 7 , another vibrato detection method is provided. Taking the application of this method to a terminal as an example for illustration, it includes the following steps:
[0137] Step S701: Obtain the fundamental frequency sampling rate of the dry voice to be detected according to the frame interval; perform fundamental frequency extraction processing on the dry voice to be detected according to the fundamental frequency sampling rate to obtain the initial fundamental frequency sequence of the dry voice to be detected; the unit of the initial fundamental frequency sequence is frequency.
[0138] Step S702: Perform unit conversion processing on the initial fundamental frequency sequence of the dry voice to be detected according to the mapping relationship between frequency and notes to obtain the fundamental frequency sequence of the dry voice to be detected; the unit of the fundamental frequency sequence is notes.
[0139] Step S703: Perform speech recognition on the dry voice to be detected to obtain each lyric corresponding to the dry voice to be detected; perform segmentation on the fundamental frequency sequence of the dry voice to be detected according to each lyric corresponding to the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric.
[0140] Step S704: According to the preset bandwidth range and the preset filter order, perform simulation processing on the filter to obtain the polynomial coefficients of the filter; according to the polynomial coefficients of the filter, perform band-pass filtering on each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
[0141] Among them, step S704 can be specifically implemented through the following steps: Step S704-11: According to the polynomial coefficients of the filter, perform forward filtering on each fundamental frequency segment to obtain the forward noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Step S704-12: Perform reverse filtering on each forward noise-reduced fundamental frequency segment to obtain the reverse noise-reduced fundamental frequency segment corresponding to each forward noise-reduced fundamental frequency segment; perform flipping processing on each reverse noise-reduced fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
[0142] Among them, step S704 can also be specifically implemented through the following steps: Step S704-21: According to the polynomial coefficients of the filter, perform reverse filtering on each fundamental frequency segment to obtain the reverse noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment. Step S704-22: Perform reverse filtering on each reverse noise-reduced fundamental frequency segment again to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
[0143] Step S705: According to the peak-to-peak distance and valley-to-valley distance in each noise-reduced fundamental frequency segment, obtain the fluctuation period of each noise-reduced fundamental frequency segment; according to the fluctuation period of each noise-reduced fundamental frequency segment and the fundamental frequency sampling rate of the dry sound to be detected, obtain the fluctuation frequency of each noise-reduced fundamental frequency segment.
[0144] Step S706: According to the peaks and valleys in the fluctuation amplitude of each noise-reduced fundamental frequency segment, obtain the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within the preset amplitude range.
[0145] Step S707: For each noise-reduced fundamental frequency segment, when the number of peaks within the preset amplitude range of the noise-reduced fundamental frequency segment meets the preset vibrato condition, the number of valleys meets the preset valley number condition, and the fluctuation frequency meets the preset fluctuation frequency condition, confirm that the noise-reduced fundamental frequency segment belongs to vibrato.
[0146] It should be noted that steps S704-11, S704-12 and steps S704-21, S704-22 are two execution methods of the example of step S704, and they are in a parallel relationship. In practical applications, either one can be selected for execution, or both can be executed in parallel and the results of both can be combined to obtain the final noise-reduced fundamental frequency segment.
[0147] The above vibrato detection method can achieve the following beneficial effects: performing lyric segmentation on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected; performing filtering processing on each fundamental frequency segment to obtain the denoised fundamental frequency segment corresponding to each fundamental frequency segment; obtaining the vibrato detection result of each denoised fundamental frequency segment according to the fluctuation amplitude and fluctuation frequency of each denoised fundamental frequency segment, and using it as the target vibrato detection result of the dry voice to be detected, realizing the vibrato detection of the dry voice to be detected. By performing vibrato detection on the fundamental frequency sequence of the dry voice to be detected, the problem of vibrato detection failure caused by human voice differences can be avoided, and by performing lyric segmentation on the fundamental frequency sequence of the dry voice to be detected, a finer-grained judgment of the vibrato of the dry voice to be detected can be made to improve the accuracy of vibrato detection. In addition, by performing filtering processing on the fundamental frequency segments, the influence of melody trends and noise on vibrato detection can be avoided, greatly improving the accuracy and robustness of vibrato detection.
[0148] To more clearly illustrate the vibrato detection method provided by the embodiments of the present disclosure, the above vibrato detection method will be specifically described below with a specific embodiment. As Figure 8 shown, another vibrato detection method is provided, which can be applied to Figure 1 the terminal in, and specifically includes the following content:
[0149] (1) Fundamental frequency extraction: The terminal obtains the high-quality dry voice to be detected of the user, and then the terminal sets the frame interval to 5 ms, so the fundamental frequency sampling rate of the fundamental frequency sequence is 1 / 5 ms = 1 / 0.005 s = 200 Hz. Through the fundamental frequency extraction tool, an effective initial fundamental frequency sequence X(n) is extracted from the dry voice to be detected, where n = 0, 1, 2, 3,..., N - 1, and N represents the total number of frames of the fundamental frequency sequence. The fundamental frequency extraction tool includes: probabilistic YIN technology, CREPE technology, and Harvest fundamental frequency extractor. Then, according to the mapping relationship between frequency and notes shown below, unit conversion processing is performed on the initial fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency sequence of the dry voice to be detected.
[0150]
[0151] (2) Lyric segmentation: Since the vibrato of the human voice often adheres to a single word in the lyrics, the terminal can perform speech recognition processing on the fundamental frequency sequence of the dry voice to be detected through Automatic Speech Recognition (ASR) technology to obtain each lyric in the dry voice to be detected; then, the fundamental frequency sequence is segmented according to the lyric granularity of a single word to obtain the fundamental frequency segments corresponding to each lyric.
[0152] (3) Fundamental frequency filtering: To avoid the interference of high-frequency noise and low-frequency melody trends on the fluctuations of the fundamental frequency sequence, the terminal sets the preset bandwidth range to [4, 20] Hz and the preset filter order to the second order, and then constructs a Butterworth filter with a second-order band-pass structure. Its expression H(z) is as follows:
[0153]
[0154] Among them, z represents the frequency domain of the Butterworth filter with a second-order band-pass structure; the superscripts -1 and -2 of z represent the order respectively; b0, b1, and b2 are all numerator polynomial coefficients; a0, a1, and a2 are all denominator polynomial coefficients.
[0155] By simulating the Butterworth filter with a second-order band-pass structure, the coefficients of the filter are obtained: b0 = 0.1122, b1 = 0, b2 = -0.1122, a0 = 1.0000, a1 = -1.7581, a2 = 0.7757. Then, according to the numerator polynomial coefficients b0, b1, and b2, and the denominator polynomial coefficients a0, a1, and a2, each fundamental frequency segment is subjected to zero-phase band-pass filtering through this filter to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
[0156] (4) Fundamental frequency fluctuation period estimation: The terminal can obtain the fluctuation period of each noise-reduced fundamental frequency segment according to the average value of all peak spacings and valley spacings in each noise-reduced fundamental frequency segment; the fluctuation period can be expressed by the following formula:
[0157]
[0158] Among them, represents the fluctuation period; I represents the total number of peaks in the noise-reduced fundamental frequency segment; J represents the total number of valleys in the noise-reduced fundamental frequency segment; Δp(i) represents the spacing between the i-th peak position and the (i - 1)-th peak position in the peak sequence, that is, Δp(i) = pks(i) - pks(i - 1), where pks(i) represents the peak sequence; Δv(j) represents the spacing between the j-th valley position and the (j - 1)-th valley position in the valley sequence, that is, Δv(j) = vly(j) - vly(j - 1), where vly(j) represents the valley sequence.
[0159] The terminal obtains the fluctuation frequency of each noise-reduced fundamental frequency segment according to the fluctuation period of each noise-reduced fundamental frequency segment and the fundamental frequency sampling rate of the dry sound to be detected. The fluctuation frequency can be expressed by the following formula:
[0160]
[0161] Among them, F represents the fluctuation frequency; fs represents the fundamental frequency sampling rate. It can be known from step (1) that fs = 200 Hz.
[0162] (5) Vibrato judgment: The terminal can set the preset amplitude range to [0.3, 2.1] and [-2.1, -0.3], and then judge whether the peak sequence pks(i) of the noise-reduced fundamental frequency sequence is greater than or equal to 0.3 and less than or equal to 2.1, and whether the valley sequence vly(j) of the noise-reduced fundamental frequency sequence is greater than or equal to -2.1 and less than or equal to -0.3; furthermore, judge whether the number of peaks (marked as np) of the peak sequence pks(i) of the noise-reduced fundamental frequency segment within the range of [0.3, 2.1] is greater than or equal to 2, and whether the number of valleys (marked as nv) of the valley sequence vly(j) of the noise-reduced fundamental frequency segment within the range of [-2.1, -0.3] is greater than or equal to 2; the preset fluctuation frequency condition can be set to [4, 8] Hz, and finally judge whether the fluctuation frequency F of the noise-reduced fundamental frequency sequence is within the range of [4, 8] Hz; when the results of the above three judgments are all yes, confirm that the noise-reduced fundamental frequency segment belongs to vibrato, otherwise, confirm that the noise-reduced fundamental frequency segment does not belong to vibrato; finally, obtain the target vibrato detection result of the dry voice to be detected by the user.
[0163] In this embodiment, it is possible to more accurately detect the vibrato part in the dry voice to be detected, and it is also possible to avoid the vibrato detection failure caused by the influence of the pitch trend of the dry voice to be detected by the user. At the same time, by setting the preset bandwidth range in the fundamental frequency wave, the influence of unreasonable frequencies can be effectively avoided, and it has higher detection accuracy and robustness.
[0164] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0165] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 9As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a tremor detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0166] Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0167] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0168] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0169] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0170] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0171] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0172] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0173] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A tremolo detection method, characterized in that, The method includes: Performing lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected; Performing filtering processing on each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; Determining the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment; Obtaining the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within a preset amplitude range according to the peaks and valleys in the fluctuation amplitude of each noise-reduced fundamental frequency segment; wherein the vibrato detection result of each noise-reduced fundamental frequency segment is the vibrato detection result of the dry voice to be detected; For each noise-reduced fundamental frequency segment, when the number of peaks within the preset amplitude range of the noise-reduced fundamental frequency segment meets the preset peak number condition, the number of valleys meets the preset valley number condition, and the fluctuation frequency meets the preset fluctuation frequency condition, it is confirmed that the noise-reduced fundamental frequency segment belongs to vibrato; the preset peak number condition is used to characterize the range of the number of peaks of the fundamental frequency segment belonging to vibrato; the preset valley number condition is used to characterize the range of the number of valleys of the fundamental frequency segment belonging to vibrato; the preset fluctuation frequency condition is used to characterize the range of the fluctuation frequency of the fundamental frequency segment belonging to vibrato.
2. The method according to claim 1, wherein The performing filtering processing on each fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment includes: Performing simulation processing on the filter according to a preset bandwidth range and a preset filter order to obtain the polynomial coefficients of the filter; Performing band-pass filtering processing on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
3. The method according to claim 2, wherein The performing band-pass filtering processing on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment includes: Performing forward filtering processing on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain the forward noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; Performing reverse filtering processing on each forward noise-reduced fundamental frequency segment to obtain the reverse noise-reduced fundamental frequency segment corresponding to each forward noise-reduced fundamental frequency segment; Performing flipping processing on each reverse noise-reduced fundamental frequency segment to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
4. The method according to claim 2, wherein The performing band-pass filtering processing on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment includes: Performing reverse filtering processing on each fundamental frequency segment according to the polynomial coefficients of the filter to obtain the reverse noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment; Performing reverse filtering processing on each reverse noise-reduced fundamental frequency segment again to obtain the noise-reduced fundamental frequency segment corresponding to each fundamental frequency segment.
5. The method according to claim 1, wherein The performing lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected includes: Performing speech recognition on the dry voice to be detected to obtain each lyric corresponding to the dry voice to be detected; Segment the fundamental frequency sequence of the dry voice to be detected according to each lyric corresponding to the dry voice to be detected, so as to obtain the fundamental frequency segments corresponding to each lyric.
6. The method according to claim 1, wherein Before performing lyric segmentation processing on the fundamental frequency sequence of the dry voice to be detected to obtain the fundamental frequency segments corresponding to each lyric in the dry voice to be detected, it further includes: Obtain the fundamental frequency sampling rate of the dry voice to be detected according to the frame interval. Perform fundamental frequency extraction processing on the dry voice to be detected according to the fundamental frequency sampling rate to obtain the initial fundamental frequency sequence of the dry voice to be detected; the unit of the initial fundamental frequency sequence is frequency. Perform unit conversion processing on the initial fundamental frequency sequence of the dry voice to be detected according to the mapping relationship between frequency and notes to obtain the fundamental frequency sequence of the dry voice to be detected; the unit of the fundamental frequency sequence is notes.
7. The method according to claim 1, characterized in that, The determining the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment includes: Obtain the fluctuation period of each noise-reduced fundamental frequency segment according to the peak spacing and valley spacing in each noise-reduced fundamental frequency segment. Obtain the fluctuation frequency of each noise-reduced fundamental frequency segment according to the fluctuation period of each noise-reduced fundamental frequency segment and the fundamental frequency sampling rate of the dry voice to be detected.
8. The method according to claim 1, wherein The obtaining the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the fluctuation amplitude and fluctuation frequency of each noise-reduced fundamental frequency segment meet the preset vibrato condition includes: Obtain the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within the preset amplitude range according to the peaks and valleys in the fluctuation amplitude of each noise-reduced fundamental frequency segment. Obtain the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within the preset amplitude range meet the preset vibrato condition and whether the fluctuation frequency meets the preset vibrato condition.
9. The method according to claim 8, wherein The obtaining the vibrato detection result of each noise-reduced fundamental frequency segment according to whether the number of peaks and the number of valleys of each noise-reduced fundamental frequency segment within the preset amplitude range meet the preset vibrato condition and whether the fluctuation frequency meets the preset vibrato condition includes: For each noise-reduced fundamental frequency segment, when the number of peaks within the preset amplitude range of the noise-reduced fundamental frequency segment meets the preset peak number condition, the number of valleys meets the preset valley number condition, and the fluctuation frequency meets the preset fluctuation frequency condition, confirm that the noise-reduced fundamental frequency segment belongs to vibrato.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Trill modeling method, device, computer equipment and storage medium
CN109817191A
Vibrato identification method and device
CN110827859A