Real-time Voice-to-Singing Conversion Technology

By extracting specific information of speech frames and combining these information to generate singing frames, the problem that speech characteristics cannot be maintained when voice conversion into singing in the prior art is solved, and a natural singing conversion effect is achieved.

CN114765029BActive Publication Date: 2025-06-24AGORA LAB INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110608545.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-01-14
Filing Date
2021-06-01
Publication Date
2025-06-24
Estimated Expiration
2041-06-01

AI Technical Summary

Technical Problem

The prior art is difficult to maintain the speaker's voice characteristics when converting voice into singing, resulting in unnatural output sound.

Method used

By extracting the pitch value, formant information, non-periodic information, lead pitch and chord pitch of the speech frame, combining these information to generate singing frames, thereby realizing the conversion of speech to singing.

Benefits of technology

It realizes the conversion of voice into natural singing voice on the basis of maintaining the speaker's voice characteristics, solving the problem of unnatural sound in traditional technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114765029B_ABST
    Figure CN114765029B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for converting a sample voice frame into a singing voice frame, including: obtaining the pitch value of an audio frame; obtaining the formant information of the frame using the pitch value; obtaining the aperiodic information of the frame using the pitch value; obtaining the fundamental pitch and chord pitches; obtaining the singing voice frame using the formant information, aperiodic information, fundamental pitch, and chord pitches; and outputting or saving the singing voice frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Patent Application No. 17 / 149,224, filed on January 14, 2021, with the title "Real - Time Conversion Technology from Speech to Singing", the entire content of which is incorporated herein by reference. Field of the Invention

[0003] The present invention generally relates to the field of speech enhancement. More specifically, the present invention relates to the technology of converting spoken speech into singing voice in real - time applications. Background Art

[0004] Interactive communication often occurs online through different media types in different communication channels. For example, real - time communication (RTC) transmitted using video conferencing or video streaming. The video may contain audio and video content. A user (i.e., the sender user) can send user - generated content (such as video) to one or more recipient users. For example, a concert can be live - streamed to many viewers. Also, a teacher can live - stream a class to students. Additionally, users can also conduct real - time chats that include real - time video.

[0005] In real - time communication, some users may want to add filters, masks, and other visual effects to add fun to the communication. For example, a user can select a pair of sunglasses filter, which is digitally added to the user's face by the communication application. Similarly, users may want to change their voices. More specifically, a user may want to transform their own voice into the effect of singing voice according to a reference sample. Summary of the Invention

[0006] On the one hand, the present invention proposes a method for converting a speech sample frame into a singing voice frame. The method includes obtaining the pitch value of an audio frame; obtaining the formant information of the frame using the pitch value; obtaining the aperiodic information of the frame using the pitch value; obtaining the tonic pitch and chord pitches; obtaining the singing voice frame using the formant information, aperiodic information, tonic pitch, and chord pitches; and outputting or saving the singing voice frame.

[0007] On the other hand, the present invention proposes a device for converting a speech sample frame into a singing voice frame. The device includes a processor configured to obtain the pitch value of an audio frame; obtain the formant information of the frame using the pitch value; obtain the aperiodic information of the frame using the pitch value; obtain the tonic pitch and chord pitches; obtain the singing voice frame using the formant information, aperiodic information, tonic pitch, and chord pitches; and outputting or saving the singing voice frame.

[0008] In a third aspect, the present invention provides a non-transitory computer-readable storage medium containing instructions executable by a processor, and the operations executable by the instructions include: obtaining the pitch value of an audio frame; obtaining the formant information of the frame using the pitch value; obtaining the aperiodic information of the frame using the pitch value; obtaining the fundamental pitch and the chord pitch; obtaining a singing voice frame using the formant information, the aperiodic information, the fundamental pitch, and the chord pitch; and outputting or saving the singing voice frame.

[0009] The above aspects can be implemented in various different embodiments. For example, the above aspects can be implemented by a suitable computer program, which can be implemented on a suitable carrier medium. The suitable carrier medium can be a tangible carrier medium (such as a disk) or an intangible carrier medium (such as a communication signal). The functions of each aspect can also be implemented using a suitable device, which can take the form of a programmable computer running a computer program configured to implement the methods and / or technologies described in the present invention. The above aspects can also be combined to enable the functions described in one aspect of the technology to be implemented in another aspect of the technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The descriptions herein refer to the accompanying drawings, in which the same reference numerals denote the same components in the respective drawings.

[0011] Figure 1 is an example diagram of a system for converting speech to singing voice drawn according to an embodiment of the present invention.

[0012] Figure 2A is a technical flow chart of a feature extraction module drawn according to an embodiment of the present invention.

[0013] Figure 2B is a technical flow chart of calculating the pitch value drawn according to an embodiment of the present invention.

[0014] Figure 2C is a technical flow chart of calculating the aperiodic information drawn according to an embodiment of the present invention.

[0015] Figure 2D is a technical flow chart of extracting the formant information drawn according to an embodiment of the present invention.

[0016] Figure 3A is a technical flow chart of generating singing voice features in a static mode drawn according to an embodiment of the present invention.

[0017] Figure 3B is a technical flow chart of generating singing voice features in a dynamic mode drawn according to an embodiment of the present invention.

[0018] Figure 3CShows a visual view of an example MIDI file.

[0019] Figure 3D Shows a visual view of an example pitch track file.

[0020] Figure 3E Shows a visual view of the perfect fifth rule.

[0021] Figure 4 Is a technical flow chart of singing synthesis drawn according to an embodiment of the present invention.

[0022] Figure 5 Is a technical example flow chart of voice-to-singing conversion drawn according to an embodiment of the present invention.

[0023] Figure 6 Is an example block diagram of a computing device drawn according to an embodiment of the present invention. Detailed implementation

[0024] As described above, a user may wish to convert his / her voice (i.e., speech) into singing according to a reference sample. That is, when the user speaks in his / her normal voice (i.e., the source speech sample), the remote recipient may hear the user's speech sung according to the reference sample. That is, the pitch of the speaker is modified (such as through tuning) to sing the melody of the reference sample, which can be a song, a tune, a musical work, etc.

[0025] Although traditional pitch adjustment techniques, such as the phase vocoder or pitch synchronous overlap and add (PSOLA), etc., can modify the pitch of speech, since the energy distribution of the entire frequency band may be evenly spread or squeezed, which may also change the formants of the speech, the output (such as the effect) of this technique is speech (such as speech) that is not similar to the speaker's speech, and may sound like the voice of another person, or become unnatural (such as a robotic voice, etc.). That is, traditional techniques tend to lose the speech characteristics of the original speaker.

[0026] We hope to maintain the speech characteristics of the speaker when converting a speech sample into singing according to a reference sample. The speech characteristics of the speaker (such as the unique characteristics of the speaker's speech) can be embedded (such as through encoding, etc.) in the formant information. Formants are the concentration of acoustic energy near a specific frequency in a sound wave. When a vowel is pronounced, the formants represent the resonant characteristics of the vocal tract. Each cavity in the vocal tract can resonate at a corresponding frequency. These resonance characteristics can be used to identify a person's voice quality.

[0027] For a reference sample, the pitch trajectory and chords of the reference sample will be applied to the speech sample. Pitch refers to the starting and ending notes of a musical scale used in music composition. We define a note as the first degree of a diatonic scale, a pitch center, and / or a final resolved pitch. For example, referring to a reference sample (such as a musical piece) as being in the key of "C" means that the reference sample harmonically centers around the note C, and the first note or pitch of the major scale is C. We define the tonic pitch in the reference sample as the one that produces the maximum amplitude. A pitch trajectory refers to the sequence of pitches in the reference sample. A chord refers to a string of notes separated by intervals. A chord can be a group of notes played together.

[0028] Traditional singing voice generation techniques can generate multiple tracks of chords based on a pitch trajectory and then mix the chord tracks with the pitch track to generate a singing voice signal. Such techniques result in an increased computational cost and have the drawback that they cannot be implemented on portable devices (such as mobile phones).

[0029] The embodiments described in the present invention can convert a speech sample (such as a spoken speech sample) into a singing voice according to a reference sample. The voice conversion to singing voice technology described herein can modify the pitch trajectory of the original speech according to the reference pitch of a given melody without changing the characteristics of the speaker. The conversion process can be implemented in real time. The conversion can be based on a static reference sample or a dynamic reference sample. In the case of using a static reference sample, the preset tonic pitch and the trajectory of the chord pitch can be recycled. In the case of using a dynamic reference sample (i.e., the dynamic mode), the tonic pitch and the chord pitch signals can be received (such as calculated, extracted, analyzed, etc.) in real time from an input device (or a virtual device) (such as a keyboard or a touch screen, etc.). For example, when a user is speaking, the performance of an instrument may be playing in the background, and thus the user's voice can be modified according to the pitch and chords of the music being played.

[0030] Figure 1 FIG. 10 is an example diagram of a system for converting speech to singing voice drawn according to an embodiment of the present invention. Device 100 can convert the received audio sample into a singing voice. Device 100 can be a sending device of the sender, can be implemented in the sending device, or can be a part of the sending device. Device 100 can be a receiving device of the receiver, can be implemented in the receiving device, or can be a part of the receiving device.

[0031] Device 100 can receive audio samples (such as speech) from the sending user. For example, the audio sample can be the speech of the sending user, such as in the scenario of an audio or video conference call with one or more receiving users. In one example, the sending device of the sending user can convert the speech of the sending user into singing voice and then send the singing voice to the receiving user. In another example, the speech of the sending user can be sent to the receiving user as it is, and the receiving device of the receiving user can convert the received speech into singing voice before outputting the singing voice to the receiving user, such as using the microphone of the receiving device. The singing audio can be output to a storage medium for later playback.

[0032] Device 100 receives the source speech in the form of frames, such as source audio frame 108. In another example, device 100 can divide the received audio signal into frames, including source audio frame 108. Device 100 processes the source speech frame by frame. One frame can be m milliseconds of audio. In one example, m can be 20 milliseconds. Of course, m can also be other values. Device 100 outputs (such as generates, obtains, produces, calculates, etc.) singing audio frame 112. Source audio frame 108 is the original speech of the sending user, and singing audio frame 112 is the singing audio frame converted according to reference signal 110.

[0033] Device 100 includes a feature extraction module 102, a singing feature generation module 104, and a singing synthesis module 106. The feature extraction module 102 can estimate the pitch and formant information of each received audio frame (i.e., source audio frame 108). In the present invention, "estimate" can mean calculating, obtaining, identifying, selecting, constructing, deriving, forming, producing, or other forms of estimation in any way. The singing feature generation module 104 can obtain the main pitch and chord pitch from the reference signal 110 and apply them to each frame. The singing synthesis module 106 uses the information provided by the feature extraction module 102 and the singing feature generation module 104 to generate the singing signal (i.e., singing audio frame 112) frame by frame.

[0034] To summarize the above content and give an example, when the speaker is speaking, the feature extraction module 102 extracts the features of the real-time speech signal; at the same time, the singing feature generation module 104 generates singing information such as the main pitch and chord pitch; then the singing synthesis module 106 generates the singing signal according to the speech and singing features.

[0035] The following refers to Figures 2A - 2D 、 Figures 3A - 3D and Figure 4 to further describe the feature extraction module 102, the singing feature generation module 104, and the singing synthesis module 106.

[0036] Each module of device 100 can be implemented by a computing device (such asFigure 6 It can be implemented by a computing device (such as computing device 600) in []. Technology 600 can be implemented as a software program executed by a computing device (such as computing device 600). The software program can include machine-readable instructions that can be stored in a memory (such as memory 604 or auxiliary memory 614), and when run by a processor (such as processor 602), can cause the computing device to execute technology 600. Technology 600 can be implemented using dedicated hardware or firmware. Multiple processors and / or multiple memories can also be used.

[0037] Figures 2A - 2D It is a detailed example diagram of extracting features from an audio frame drawn according to an embodiment of the present invention.

[0038] Figure 2A It is a flowchart of technology 200 for a feature extraction module drawn according to an embodiment of the present invention. Technology 200 can be implemented by Figure 1 feature extraction module 102. Technology 200 includes a pitch detection module (detecting pitch through the autocorrelation technique of autocorrelation module 204); and an aperiodicity estimation module 208 for extracting the aperiodic features of the source audio frame 108. The formant extraction module 210 can adopt a spectral smoothing technique to extract formant information, as described below.

[0039] The pitch detection module (i.e., formant extraction module 210) can calculate the pitch value (F0) for each source audio frame 108 of the speech signal. The pitch value can be used to determine the window length of the fast Fourier transform (FFT) 206, which is used by both the formant extraction module 210 and the aperiodicity estimation module 208. The FFT 206 can also be used to obtain the length of the audio signal required to perform the FFT. As described below, the lengths obtained from aperiodicity estimation and formant extraction can be 2*T0 and 3*T0 respectively, where T0 is determined by the pitch F0 (such as T0 = 1 / F0). For example, the feature extraction module 102 can search for the pitch value (F0) within the pitch search range. Another example is that the pitch search range can be from 75 Hz to 800 Hz, covering the normal range of human pitch. The autocorrelation module 204 can obtain the pitch value (F0), and the autocorrelation module 204 performs an autocorrelation operation on a part of the signal stored in the signal buffer 202. The length of the signal buffer 202 can be at least 40 ms, which is obtained from the lowest pitch (75 Hz) of the pitch detection range. The signal buffer 202 can include the sampled data of at least 2 frames in the source audio signal. The signal buffer 202 can be used to store audio frames of a specific total length (such as 40 ms).

[0040] The feature extraction module 102 can provide formants (i.e., spectral envelopes) and aperiodicity information to the singing synthesis module 106 through the concatenation module 212, as shown in Figure 2.

[0041] Figure 2B is a technical flowchart for calculating pitch values drawn according to an embodiment of the present invention. The pitch value (F0) can be obtained through the autocorrelation module 204 in FIG. 2, thereby implementing technology 220. More specifically, the autocorrelation technique (i.e., technology 220) can be used to calculate (such as detect, select, identify, choose, etc.) the pitch value (F0).

[0042] At 222, technology 220 calculates the autocorrelation information of the signal in the signal buffer. Autocorrelation calculation can be used to identify patterns in data (such as time series data). The autocorrelation function can be used to identify the correlation between a pair of values within a specific delay time. For example, the lag-1 autocorrelation calculation can measure the correlation between directly adjacent data points. The lag-2 autocorrelation calculation can measure the correlation between a pair of values separated by 2 time periods (i.e., 2 time distances). Equation (1) can be used to calculate the autocorrelation value:

[0043] r n = r(nΔτ) (1)

[0044] In Equation (1), r() is the autocorrelation function used to calculate the autocorrelation value between values with different time delays (such as nΔτ); Δτ is the sampling time. For example, given the sampling frequency f s of the source audio frame 108 is 10K, then Δτ is 0.1 millisecond (ms); n can be in the range of [12, 134], corresponding to the pitch search range.

[0045] At 224, technology 220 obtains (such as calculates, determines, obtains, etc.) the local maximum in the autocorrelation calculation. For example, the local maximum of the autocorrelation can be obtained in each pair of (m - 1)Δτ and (m + 1)Δτ, where m has the same range as n. That is, among all the calculated r n ’s, the local maximum r m ’s can be obtained. Each local maximum r m makes:

[0046] r m > r m+1 and r m > r m-1 (2)

[0047] At 226, for each local maximum r m , the in-frame corresponding time position of the local maximum (τ max ) and the interpolation value of the autocorrelation local maximum (r max ) are calculated using Equations (3) and (4) respectively. τ max can be the one with the maximum autocorrelation (r maxThe delay time of (). Of course, other methods can also be used to obtain τ max and r max .

[0048]

[0049]

[0050] At 228, the technique 220 sets (such as calculates, selects, identifies, etc.) the pitch value (F0). For example, if there is a local maximum with r max > 0.5, the pitch value can be calculated using the τ with the maximum r through Equation (5) max and the flag Pitch_flag is set to true; otherwise (that is, if there is no local maximum r max > 0.5), then F0 can be set to a predetermined value and Pitch_flag is set to false. The predetermined value can be a value within the pitch detection range, such as the middle value within this range. Another example is that the predetermined value can be 75, which is the lowest pitch value within the pitch detection range. max

[0051]

[0052] Figure 2C is a technical flowchart for calculating non-periodic information drawn according to an embodiment of the present invention. The non-periodicity is calculated based on the group delay. Through Figure 2A the non-periodicity estimation module 208 of obtains the band non-periodic information (that is, the non-periodicity of at least some frequency sub-bands) of the source audio frame 108, thereby implementing the technique 240.

[0053] At 242, the technique 240 calculates the group delay. The group delay indicates (such as describes, etc.) how the spectral envelope changes at different time points or over time. Therefore, the following method can be used to calculate the group delay of the source audio frame 108.

[0054] For each frame, the group delay τ is calculated using the signal s(t) with a length of (2*T0) D , where T0 = 1 / F0. The group delay is defined by Equation (6):

[0055]

[0056] In Equation (6), and respectively represent the real part and the imaginary part of the complex value; S(ω) represents the spectrum of the signal s(t), and S′(ω) is the weighted spectrum calculated using Equation (7), where represents the Fourier transform:

[0057] ​

[0058] At 244, the technique 240 uses group delay to calculate the aperiodicity of each sub - band. The entire vocal frequency range (i.e., [0 - 15] kHz) can be divided into a predetermined number of frequency bands. For example, the predetermined number of frequency bands can be 5. Of course, it can also be divided into other numbers. Thus, in one example, the frequency bands can be sub - bands [0 - 3 kHz], [3 kHz - 6 kHz], [6 kHz - 9 kHz], [9 kHz - 12 kHz], and [12 kHz - 15 kHz]. Of course, different divisions of the vocal frequency range can also be adopted. The aperiodicity of the sub - bands can be calculated using equations 8 - 10

[0059]

[0060]

[0061]

[0062] In equations 8 - 10, where is the center frequency of the i - th sub - band. w(w) is a window function; w l is the window length (which can be equal to twice the sub - frequency bandwidth); is the inverse Fourier transform. Thus, the waveform can be calculated using the inverse Fourier transform In the parameter (Equation (9)), p s (t, ω c ) represents a parameter calculated by sorting the power waveform in descending order on the time axis. In equation (10), w bw represents the main - lobe bandwidth of the window function w(w), which has a time dimension. Since the main - lobe bandwidth can be defined as the shortest frequency range from 0 Hz to the frequency at which the amplitude is 0, 2w bw can be used.

[0063] In one example, a window function with low side - lobes can be used to prevent data from being aliased (or replicated) in the frequency domain. For example, the Nuttall window can be used because the side - lobes of this window function are low. In another example, the Blackman window can also be used.

[0064] Figure 2D is a technical flowchart for extracting formant information drawn according to an embodiment of the present invention. Through Figure 2AThe formant extraction module 210 obtains the formant information of the source audio frame 108, thereby implementing the technology 260. The formant information can be represented by a spectral envelope (such as a smoothed spectrum). A filtering function can be applied to the cepstrum of the windowed signal to achieve smoothing of the magnitude spectrum. Since human speech or voice signals can have sidebands, the cepstrum can be used in speech processing to understand (such as analyze) the differences in pronunciation and different words. The cepstrum is a technique by which a set of sidebands from a signal source can be aggregated into one parameter. Of course, other methods can also be used to extract formant information.

[0065] At 262, the technology 260 calculates the power cepstrum from the windowed signal. As is well known, the cepstrum of a signal is the inverse Fourier transform of the Fourier transform of the signal and the logarithm of its Fourier transform. As described above, the length of the window can be 3*T0, where T0 = 1 / F0. Since the cepstrum is obtained using the inverse Fourier transform, the cepstrum is in the time domain. The Hamming window w(t) can be used in Equation (11) to calculate the power cepstrum:

[0066] p s (t) = F -1 [log(|F{s(t)*w(t)}| 2 )] (11)

[0067] At 264, the technology 260 calculates the smoothed spectrum (i.e., formant) from the cepstrum using Equation (12):

[0068]

[0069] The constants 1.18 and 0.18 are derived empirically to obtain a smoothed formant. Of course, other values can also be used.

[0070] Now look at Figure 1 the singing voice feature generation module 104. As described above, the singing voice feature generation module 104 can operate in either a static mode or a dynamic mode. The singing voice feature generation module 104 can obtain (such as use, calculate, derive, select, etc.) the fundamental pitch and chord pitches (such as zero or more chord pitches) for converting the source audio frame 108 into a singing voice audio frame 112.

[0071] Figure 3A is a flowchart of the technology 300 for generating singing voice features in the static mode drawn according to an embodiment of the present invention. The technology 300 can be implemented by Figure 1 the singing voice feature generation module 104. In the static mode, before performing real-time voice-to-singing voice conversion on the input speech signal, Figure 1The reference signal 110 (i.e., the reference sample 302) is sent to the singing voice feature generation module 104.

[0072] For example, the reference sample 302 can be a Musical Instrument Digital Interface (MIDI) file. A MIDI file can contain detailed information about all aspects from recording to performance (such as playing on a piano). A MIDI file can be regarded as containing a copy of the performance. For example, a MIDI file includes information such as the notes of the performance, the order of the notes, the length of each played note, whether the pedal is depressed (in the case of a piano), and so on. Figure 3C A visual view 360 of an example MIDI file is shown. For example, the channel 362 represents the playing position of the E2 note relative to other notes and the duration of each E2 note.

[0073] In one example, the reference sample 302 can be a pitch track file. Figure 3D A visual view 370 of a pitch track file is shown. The visual view 370 shows the pitch (vertical axis) information used for each frame (horizontal axis) of the audio file. The solid line 372 represents the main pitch; the dashed line 374 represents the first chord pitch; the dotted line 376 represents the second chord pitch.

[0074] In the static mode, the singing voice feature generation module 104 (such as the main pitch loop module 304 therein) repeatedly provides the main pitch for each frame according to a preset pitch track described at the reference sample 302 (such as configuration, recording, setting, etc.). When all the pitches of the reference sample 302 are exhausted, the main pitch loop module 304 will start cycling from the first frame of the reference sample 302. In one example, the reference sample 302 (such as a MIDI file) can also include chord pitch information. Therefore, the chord pitch generation module 306 can also obtain the chord pitch (such as one or more chord pitches) for each frame by referring to the reference sample 302. In another example, the chord pitch generation module 306 can obtain (such as derive, calculate, etc.) the chord pitch using chord rules (such as triads, perfect fifths, or some other rules). Figure 3E An example of chord pitch using perfect fifths is shown in. Figure 3E A visual view 380 of perfect fifths is shown. The dashed line 382 represents the main pitch; the dashed line 384 represents the first chord pitch; the dotted-dashed-dotted line 386 represents the second chord pitch.

[0075] For each frame of the source audio frame 108, the concatenation module 308 concatenates the main pitch and the chord pitch and provides it to Figure 1 the singing voice synthesis module 106.

[0076] Figure 3Bis a technical flow chart for generating singing voice features in dynamic mode drawn according to an embodiment of the present invention. It can be achieved by Figure 1 's singing voice feature generation module 104 to implement technology 350 in dynamic mode. In dynamic mode, virtual musical instruments (such as virtual keyboards, virtual guitars or other virtual musical instruments) played on a portal device (such as a smartphone touch screen) or a digital musical instrument (such as an electric guitar, etc.) can provide the fundamental pitch and chord pitch information in real time. Also, for example, when a user is speaking, a background music work may be playing in the background. In this way, the user may "play" his / her voice with any melody of the musical instrument he / she plays. The signal conversion module 354 can extract the fundamental pitch and chord pitch frame by frame from the played music in real time to provide to Figure 1 's singing voice synthesis module 106. In one example, a media stream (such as a MIDI stream) containing pitch and volume information can be obtained through the signal conversion module 354, and the fundamental pitch and chord pitch are extracted frame by frame from it. For example, the played musical instrument or the software for playing music (such as musical instrument software) can support and send a MIDI stream containing pitch and volume information.

[0077] It should be noted that the pitch distribution of normal people is from 55Hz to 880Hz. Therefore, in one example,

[0078] the fundamental pitch and chord pitch can be distributed within the pitch range of normal people in order to obtain natural singing voice. That is to say, the fundamental pitch and / or chord pitch can be limited within the range of [55, 880]. For example, if the pitch is less than 55Hz, it can be set (such as clipped) to 55Hz; if it is greater than 880, it can be set (such as clipped) to 880. In another example, since clipping may produce inharmonious sounds, pitches beyond this range will not be generated.

[0079] Figure 4 is a technical flow chart for singing voice synthesis drawn according to an embodiment of the present invention. It can be achieved by Figure 1 's singing voice synthesis module 106 to implement technology 400. Technology 400 can receive the spectral envelope 402 (i.e., formants) and aperiodic information 404 at the input layer 412, which are obtained from the feature extraction module 102. Technology 400 can also receive the fundamental pitch 406 and zero or more chord pitches (such as the first chord pitch 408 and the second chord pitch 410), which are obtained from the singing voice feature generation module 104. Technology 400 uses these inputs to generate singing voice signals (i.e., singing voice audio frames 112) frame by frame.

[0080] Technology 400 can generate two types of sounds: a periodic sound generated from a pulse signal module (i.e., module 416), and white noise generated from a noise signal module (i.e., module 418). A pulse signal is a rapid transient change in signal amplitude followed by a return to the baseline value. For example, a clap inserted into the signal or inherent in the signal is an example of a pulse signal.

[0081] At module 416, a pre-prepared pulse signal is stored And at module 418, a pre-prepared (such as calculated, derived, etc.) white noise signal is stored for each frequency sub-band (such as the above five sub-bands) In this way, during real-time calculation, at least some (e.g., each) corresponding pulse signals and noise signals of the frequency sub-bands can be directly read to avoid repeated calculations.

[0082] Module 414 can use this pulse signal to generate a periodic response (i.e., a periodic sound).

[0083] Any known technique can be used to obtain the pulse signal For example, equations (13)-(14) can be used to calculate the pulse signal

[0084]

[0085]

[0086] In equation (13) for obtaining the frequency-domain pulse signal of each sub-band, the index i represents the sub-frequency band and the index j represents the frequency bin. The parameters a, b, and c can be constants derived empirically. For example, the constants a, b, and c can take the values 0.5, 3000, and 1500 respectively, which approximates the pulse signal of human speech. f(j) is the frequency of the j-th frequency point of the pulse signal spectrum, and the range of f(j) can be the entire frequency band (such as 0 - 24 kHz). For example, if the i-th frequency band is 150 - 440 Hz, then when f(j) is within 150 - 440 Hz, it will take a numerical value, and when f(j) is outside this range, then it is equal to 0. Equation (14) obtains the time-domain pulse signal of each frequency sub-band by performing an inverse Fourier transform. Therefore, for each frequency bin of the sub-frequency band, the respective pulse spectra will be obtained or acquired; then these pulse spectra are combined into a time-domain pulse signal.

[0087] Any known technique can be used to obtain the noise signal through module 420 For example, equations (15)-(17) can be used to calculate the noise signal

[0088]

[0089]

[0090]

[0091] The spectral noise (i.e., white noise) of the frequency bin (indexed by j) is obtained using Equation (15). where x1 and x2 are random number vectors starting from [0, 1], and the length is equal to half of the sampling frequency (0.5f s ). Equation (15) divides the spectral noise into respective sub-band noises. That is, Equation (15) divides the spectral noise into different sub-bands. Equation (17) obtains the noise wave signal from the spectral signal by performing an inverse Fourier transform.

[0092] Module 414 can calculate the positions within the source audio frame 108 where pulses (such as start, insert, etc.) need to be added. First, the pitch value of each sampling point of the source audio frame 108 is obtained. For the pitch value (i.e., timing index) of each sampling point j of frame k in the current source speech frame (i.e., frame k) (i.e., source audio frame 108), the interpolated pitch value F0 int (j) can be obtained using the pitch value of the previous frame. That is, F0 int (j) can be obtained by interpolating between F0(k) and F0(k - 1). The interpolation can be linear interpolation. For example, assume F0(k) = 100 and F0(k - 1) = 148, and there are 480 sampling points in each frame. Then the interpolated pitch value F0 int (j) of the k-th frame can be [147.9, 147.8, …, 100], where j = 1,.., 480.

[0093] Given a frame size of F size samples and a sampling frequency of f s , each sampling position can be a potential pulse position. By obtaining the phase shift at the sampling position j using Equation (18), the pulse position in the k-th frame can be obtained. This Equation (18) calculates the phase modulo (MOD) 2π. The phase can be in the range of [-π, π]. As shown in the pseudocode of Table I, if the phase difference between the current timing point (j) and the immediately following timing point (j + 1) is greater than π, the current timing point is identified as the pulse position. Therefore, depending on the pitch, there may be 0 or more positions in a frame where a pulse can be added. When the phase difference is large (such as greater than π), a pulse can be added to avoid phase discontinuity.

[0094]

[0095]

[0096] At module 422, an excitation signal is obtained by combining (such as mixing, etc.) the corresponding pulse and noise signals at each pulse position. The number of pulse signals and noise signals used depends on the aperiodicity of the signal. The aperiodicity in each subband can be used as the percentage allocation of the pulse-to-noise ratio in the excitation signal. The excitation signal can be obtained using Equation (19) where s represents the pulse position and k represents the current frame.

[0097]

[0098] Module 424 (i.e., the waveform generation module) can use the excitation signal to obtain the singing audio frame 112. As described above, the excitation signal can be combined with the cepstrum (calculated as described above) using Equations (20)-(22) to obtain the generated waveform signal S wav , which is the singing audio frame 112.

[0099]

[0100]

[0101]

[0102] Equation (20) obtains the Fourier transform of the smoothed spectrum (i.e., formants), which is calculated by the feature extraction module 102 as described above. In Equation (21), fft size is the size of the fast Fourier transform (FFT), which is the same as the FFT size used to calculate the smoothed spectrum. Equation (21) is an intermediate step in calculating S wav . In one example, fft size can be equal to 2048 to ensure sufficient frequency resolution. In Equation (22), w han refers to the Hanning window.

[0103] Figure 5 is a technical example flowchart of voice-to-singing conversion drawn according to an embodiment of the present invention. Technique 500 converts the frames of voice (speech) samples into singing frames. The frames of voice samples are as described for the source audio frame 108, and the singing frames can be Figure 1 the singing audio frame 112 in

[0104] Technique 500 can be implemented by a computing device (such as Figure 1 the computing device 100 in Figure 6The software program executed by the computing device 600). The software program may include machine-readable instructions that may be stored in a memory (such as memory 604 or auxiliary memory 614) and, when run by a processor (such as processor 602), may cause the computing device to execute technology 500. Technology 500 may be implemented using dedicated hardware or firmware. Multiple processors and / or multiple memories may also be used.

[0105] At 502, technology 500 obtains the pitch value of an audio frame. For the method of obtaining the pitch value, refer to the above description about F0. Thus, as described above, obtaining the pitch value of a frame may include calculating the autocorrelation value of the signal in the signal buffer; finding the local maximum value in the autocorrelation value; and obtaining the pitch value using the local maximum value.

[0106] At 504, technology 500 uses the pitch value to obtain the formant information of the frame. The method of obtaining the formant information is as described above. Thus, using the pitch value to obtain the formant information of the frame may include: obtaining the window length using the pitch value; calculating the power cepstrum of the frame using the window length; and obtaining the formant information from the cepstrum.

[0107] At 506, technology 500 uses the pitch value to obtain the aperiodicity information of the frame. The method of obtaining the aperiodicity information is as described above. Thus, obtaining the aperiodicity information may include: calculating the group delay using the pitch value; calculating the respective aperiodicity values for each frequency subband of the frame.

[0108] At 508, technology 500 obtains the fundamental pitch and chord pitches that need to be applied to (such as combined with) the audio frame. In one example, as described above, one or more fundamental pitches may be statically assigned according to a preset pitch trajectory. In another example, chord rules may be used to calculate the chord pitches. In yet another example, the fundamental pitch and chord pitches may be calculated in real time by referring to a sample. The reference sample may be a real or virtual instrument performance that is simultaneous with the speech.

[0109] At 510, technology 500 uses the formant information, aperiodicity information, fundamental pitch, and chord pitches to obtain a singing voice frame. The method of obtaining the singing voice frame is as described above. Thus, obtaining the singing voice frame may include: obtaining the corresponding impulse signals for each frequency subband of the frame; obtaining the corresponding noise signals for each frequency subband of the frame; obtaining the positions within the frame where the corresponding impulse signals and corresponding noise signals need to be inserted; obtaining the excitation signal; and obtaining the singing voice frame using the excitation signal.

[0110] At 512, the technique 500 outputs or saves the singing voice frames. For example, the singing voice frames can be converted into a savable format and stored for later playback. For example, the singing voice frames can be transmitted to the sending user or the receiving user. For another example, if the singing voice frames are generated using the sending user's device, then outputting the singing voice frames may mean sending (or sending through other devices) the singing voice frames to the receiving user. For yet another example, if the singing voice frames are generated using the receiving user's device, then outputting the singing voice frames may mean outputting the singing voice frames so that the receiving user can hear them.

[0111] Figure 6 FIG. 4 is an exemplary block diagram of a computing device drawn in accordance with an embodiment of the present invention. The computing device 600 can be a computing system including multiple computing devices, or a single computing device such as a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, and so on.

[0112] The processor 602 in the computing device 600 can be a conventional central processing unit. The processor 602 can also be other types of devices or multiple devices capable of manipulating or processing existing or future-developed information. For example, although in the examples herein a single processor as shown (such as the processor 602) can be used for implementation, advantages in terms of speed and efficiency can be realized if multiple processors are used.

[0113] In one implementation, the memory 604 in the computing device 600 can be a read-only memory (ROM) device or a random access memory (RAM) device. Other suitable types of storage devices can also be used as the memory 604. The memory 604 can contain code and data 606 accessed by the processor 602 using the bus 612. The memory 604 can also contain an operating system 608 and application programs 610, where the application programs 610 contain at least one program that allows the processor 602 to execute one or more of the techniques described herein. For example, the application programs 610 can include Application 1 to N, which contain programs and techniques available in the real-time voice-to-singing conversion application. For example, the application programs 610 can include one or more of the techniques 200, 220, 240, 250, 300, 350, 400, or 500. The computing device 600 can also include an auxiliary storage device 614, such as a memory card used with a mobile computing device.

[0114] The computing device 600 may also include one or more output devices, such as a display 618. For example, the display 618 may be a touch-sensitive display that combines a display with a touch-sensitive element for operable touch input. The display 618 may be coupled to the processor 602 via a bus 612. Other output devices that allow a user to program or use the computing device 600 may also be used as additional or alternative output devices in addition to or instead of the display 618. If the output device is a display or includes a display, the display may be implemented in various ways, including a liquid crystal display (LCD), a cathode ray tube (CRT) display, or a light-emitting diode (LED) display, such as an organic LED (OLED) display, etc.

[0115] The computing device 600 may also include an image sensing device 620 (such as a camera), or include any other image sensing device 620, existing or later developed, that can sense an image (such as an image of a user operating the computing device 600), or communicate with the above-mentioned image sensing device 620. The image sensing device 620 may be positioned to face the user operating the computing device 600. For example, the position and optical axis of the image sensing device 620 may be configured such that the field of view includes an area that is directly adjacent to and visible to the display 618.

[0116] The computing device 600 may also include a sound sensing device 622 (such as a microphone), or include any other sound sensing device 622, existing or later developed, that can sense sounds near the device 600, or communicate with the above-mentioned sound sensing device 622. The sound sensing device 622 may be positioned to face the user operating the computing device 600 and may be configured to receive sounds, and may be configured to receive sounds, such as sounds made by the user when operating the computing device 600, such as speech or other sounds. The computing device 400 may also include or communicate with a sound playback device 624, such as a speaker, headphones, or any other sound playback device, existing or later developed, that can play sounds according to instructions from the computing device 600.

[0117] Figure 6FIG. 0 depicts only the case where the processor 602 and the memory 604 of the computing device 600 are integrated into a single processing unit. Other configurations may also be adopted. The operations of the processor 602 may be distributed across multiple machines (each machine including one or more processors), and these machines may be directly coupled or coupled across a local area or other network. The memory 604 may be distributed across multiple machines, such as network-based memory or memory in multiple machines running the operations of the computing device 600. Only the case of a single bus is described herein. In addition, the bus 612 of the computing device 600 may also be composed of multiple buses. Further, the auxiliary memory 614 may be directly coupled to other components of the computing device 600, may also be accessed through a network, or may also include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, the computing device 600 may be implemented through a variety of configurations.

[0118] For simplicity of illustration, Figure 2A 、 Figure 2B 、 Figure 2C 、 Figure 2D 、 Figure 3A 、 Figure 3B 、 Figure 4 or Figure 5 the techniques 200, 220, 240, 250, 300, 350, 400 or 500 in

[0119] are respectively depicted by a series of modules, steps or operations. However, according to the present invention, these modules, steps or operations may occur in various sequences and / or simultaneously. Additionally, other steps or operations not mentioned and described herein may also be used. Further, the techniques designed according to the present invention may also be implemented without adopting all the steps or operations shown.

[0119] In this document, the word "example" is used to denote an example, instance or illustration. Any function or design described herein for "example" does not necessarily indicate that it is superior or better than other functions or designs. Instead, the word "example" is used to present concepts in a specific manner. The word "or" used herein is intended to mean an inclusive "or" rather than an exclusive "or". That is, "X includes A or B" is intended to represent any natural inclusive arrangement, unless otherwise stated or clearly judged from the context to be otherwise. In other words, if X includes A, X includes B, or X includes both A and B, then "X includes A or B" holds in any of the foregoing instances. In addition, in this application and the appended claims, "a", "an" should generally be construed to mean "one or more", unless otherwise stated or clearly indicated as singular from the context. Additionally, the two phrases "a function" or "one function" throughout this document do not mean the same embodiment or the same function, unless otherwise specifically stated.

[0120] Figure 6The computing device 600 shown and / or any component thereof, and Figure 1 any module or component shown (as well as the technologies, algorithms, methods, instructions, etc. stored thereon and / or executed thereby) can be implemented in hardware, software, or any combination thereof. Hardware includes, for example, intellectual property (IP) cores, application specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, firmware, microcontrollers, servers, microprocessors, digital signal processors, or any other suitable circuitry. In the present invention, the term "processor" should be understood to include a combination of one or more of any of the above. Terms such as "signal" and "data" may be used interchangeably.

[0121] In addition, on the one hand, the technology can be implemented using a general-purpose computer or processor having a computer program that, when run, can execute any of the corresponding technologies, algorithms, and / or instructions described herein. On the other hand, a special-purpose computer or processor can optionally be used, equipped with special hardware devices for executing any of the methods, algorithms, or instructions described herein.

[0122] Furthermore, all or part of the embodiments of the present invention may take the form of a computer program product that can be used by a computer or accessed by a computer-readable medium, etc. A computer-usable or computer-readable medium can be any device that can specifically contain, store, transmit, or transfer a program or data structure for use by or in conjunction with any processor. The medium can be electronic, magnetic, optical, electromagnetic, or semiconductor devices, etc., and can also include other suitable media.

[0123] Although the present invention has been described in connection with certain embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. On the other hand, the present invention is intended to cover various variations and equivalent arrangements within the scope of the claims, which scope should be given the broadest interpretation to cover all such variations and equivalent arrangements as are permitted by law.

Claims

1. A method for converting a speech frame into a singing voice frame, comprising: Obtaining the pitch value of an audio frame; Using the pitch value to obtain the formant information of the frame; Using the pitch value to obtain the aperiodicity information of the frame; Obtaining the fundamental pitch and chord pitches; Using the formant information, aperiodicity information, fundamental pitch, and chord pitches to obtain a singing voice frame; And Outputting or saving the singing voice frame; Wherein using the formant information, aperiodicity information, fundamental pitch, and chord pitches to obtain a singing voice frame includes: Obtaining the corresponding impulse signals of each frequency sub-band of the frame; Obtaining the corresponding noise signals of each frequency sub-band of the frame; After obtaining the corresponding noise signals of each frequency sub-band of the frame, obtaining the pitch value of each sampling point of the source audio frame; Obtain the interpolated pitch value F0 by combining the pitch value of the current frame and the pitch value of the previous frame int (j), based on Calculate the phase modulus and calculate the phase difference between the current sampling point j and the subsequent sampling point (j + 1) within the k-th frame; When the phase difference is greater than π, determining the current sampling point j as the impulse position, where the impulse position is the position in the frame where the corresponding impulse signal and corresponding noise signal need to be inserted; Obtaining an excitation signal by combining the corresponding impulse signal and noise signal at each impulse position, wherein the number of impulse signals and noise signals used depends on the aperiodicity; and Using the excitation signal to obtain a singing voice frame.

2. The method according to claim 1, wherein obtaining the pitch value of an audio frame includes: Calculating the autocorrelation value of the signal in the signal buffer; Finding the local maximum value in the autocorrelation value; And Using the local maximum value to obtain the pitch value.

3. The method according to claim 1, wherein using the pitch value to obtain the formant information of the frame includes: Using the pitch value to obtain the window length; Calculating the power cepstrum of the frame using the window length; And Obtaining the formant information from the cepstrum.

4. The method according to claim 1, wherein using the pitch value to obtain the aperiodicity information of the frame includes: Using the pitch value to calculate the group delay; And Calculating the respective aperiodicity values for each frequency sub-band of the frame.

5. The method according to claim 1, wherein the fundamental pitch is statically assigned according to a preset pitch trajectory.

6. The method according to claim 5, wherein the chord pitches are statically assigned.

7. The method according to claim 5, wherein the chord pitches are calculated using chord rules.

8. The method according to claim 1, wherein the fundamental pitch and chord pitches are calculated in real time by referring to a sample.

9. A device for converting a sample speech frame into a singing voice frame, comprising: A processor configured to perform the following operations: Obtaining the pitch value of an audio frame; Using the pitch value to obtain the formant information of the frame; Using the pitch value to obtain the aperiodicity information of the frame; Obtaining the fundamental pitch and chord pitches; Using the formant information, aperiodicity information, fundamental pitch, and chord pitches to obtain a singing voice frame; And Outputting or saving the singing voice frame; Wherein using the formant information, aperiodicity information, fundamental pitch, and chord pitches to obtain a singing voice frame includes: Obtaining the corresponding impulse signals of each frequency sub-band of the frame; Obtaining the corresponding noise signals of each frequency sub-band of the frame; After obtaining the corresponding noise signals of each frequency sub-band of the frame, obtaining the pitch value of each sampling point of the source audio frame; Obtain the interpolated pitch value F0 by combining the pitch value of the current frame and the pitch value of the previous frame int (j), based on Calculate the phase modulus and calculate the phase difference between the current sampling point j and the subsequent sampling point (j + 1) within the k-th frame; When the phase difference is greater than π, the current sampling point j is identified as the pulse position, where the pulse position is the position within the frame where the corresponding pulse signal and the corresponding noise signal need to be inserted; An excitation signal is obtained by combining the corresponding pulse signal and noise signal at each pulse position, where the number of pulse signals and noise signals used depends on the aperiodicity; and A singing frame is obtained using the excitation signal.

10. The apparatus according to claim 9, wherein obtaining the pitch value of the audio frame comprises: Calculating the autocorrelation value of the signal in the signal buffer; Finding the local maximum value in the autocorrelation value; And Obtaining the pitch value using the local maximum value.

11. The apparatus according to claim 9, wherein obtaining the formant information of the frame using the pitch value comprises: Obtaining the window length using the pitch value; Calculating the power cepstrum of the frame using the window length; And Obtaining the formant information from the cepstrum.

12. The apparatus according to claim 9, wherein obtaining the aperiodicity information of the frame using the pitch value comprises: Calculating the group delay using the pitch value; And Calculating the respective aperiodicity values for each frequency subband of the frame.

13. The apparatus according to claim 9, wherein the fundamental pitch is statically assigned according to a preset pitch trajectory.

14. The apparatus according to claim 13, wherein the chord pitch is statically assigned.

15. The apparatus according to claim 13, wherein the chord pitch is calculated using chord rules.

16. The apparatus according to claim 9, wherein the fundamental pitch and the chord pitch are calculated in real time by referring to a sample.

17. A non - transitory computer - readable storage medium, the storage medium contains instructions executed by a processor, and the operations that the instructions can perform include: Obtaining the pitch value of the audio frame; Obtaining the formant information of the frame using the pitch value; Obtaining the aperiodicity information of the frame using the pitch value; Obtaining the fundamental pitch and the chord pitch; Obtaining a singing frame using the formant information, aperiodicity information, fundamental pitch, and chord pitch; And Outputting or saving the singing frame; Wherein obtaining the singing frame using the formant information, aperiodicity information, fundamental pitch, and chord pitch comprises: Obtaining the corresponding pulse signal for each frequency subband of the frame; Obtaining the corresponding noise signal for each frequency subband of the frame; After obtaining the corresponding noise signal for each frequency subband of the frame, obtaining the pitch value of each sampling point of the source audio frame; Obtain the interpolated pitch value F0 by combining the pitch value of the current frame and the pitch value of the previous frame int (j), based on Calculate the phase modulus and calculate the phase difference between the current sampling point j and the subsequent sampling point (j + 1) within the k-th frame; When the phase difference is greater than π, the current sampling point j is identified as the pulse position, where the pulse position is the position within the frame where the corresponding pulse signal and the corresponding noise signal need to be inserted; An excitation signal is obtained by combining the corresponding pulse signal and noise signal at each pulse position, where the number of pulse signals and noise signals used depends on the aperiodicity; and Obtaining a singing frame using the excitation signal.

18. The non - transitory computer - readable storage medium according to claim 17, wherein the fundamental pitch is statically assigned according to a preset pitch trajectory, and the chord pitch is statically assigned, or the chord pitch is calculated using chord rules.

Citation Information

Patent Citations

  • Singing synthesis method and device, computer equipment and storage medium

    CN111402858A

  • Audio processing method and device

    CN111916093A