Real-time voice tone style transformation technology
By converting the voice signal to the time frequency domain and adjusting the equalizer parameters using the Bark domain gain, the complexity of tone change in real-time communication is solved, and a dynamic tone conversion effect is achieved.
Patent Information
- Application Number
- CN202110311790.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-15
- Filing Date
- 2021-03-24
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-03-24
AI Technical Summary
In real-time communication, it is difficult for users to simply change voice tone within the affordability of ordinary users, especially in real-time applications, where the prior art equalizers operate in complex and impractical.
By converting the voice signal to the time frequency domain, obtaining the amplitude average of the frequency bin and converting it to the Bark domain, using the gain of the frequency bin in the Bark domain to obtain the equalizer parameters, converting the voice to a reference tone, and dynamically update the equalizer parameters when a speech change is detected.
It realizes the simple and effective change of voice tone in real-time communication, adapts to voice changes, and provides dynamic tone conversion effect.
Smart Images

Figure CN114429763B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. patent application No. 17 / 071,454, filed on October 15, 2020, entitled “Real-time Transformation Technology for Voice Tone Style,” the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present invention generally relates to the field of speech enhancement, and more particularly to the field of speech timbre conversion technology in real-time applications. Background Art
[0004] Interactive communication often occurs online across various communication channels using various media types. For example, real-time communication (RTC) can be delivered using video conferencing or video streaming. Video can include both audio and video content. A user (the sender) can send user-generated content (such as a video) to one or more recipients. For example, a concert can be live-streamed to a large audience. Another example is a teacher live-streaming a class to students. Another example is a group of users engaging in a real-time chat that includes live video.
[0005] In real-time communication, some users may want to add filters, masks, and other visual effects to spice up the conversation. For example, a user might select a sunglasses filter, which the communication application digitally adds to the user's face. Similarly, a user might want to change their voice. More specifically, a user might want to modify the quality or timbre of their voice during an RTC session. Summary of the Invention
[0006] In one aspect, the present invention provides a method for converting a speaker's speech into a reference timbre. The method includes converting a first portion of the speaker's speech source signal into the time-frequency domain to obtain a time-frequency signal; obtaining the amplitude mean of frequency bins of the time-frequency signal over time; converting the amplitude mean of the frequency bins into the Bark domain to obtain a source frequency response curve (SR), where SR(i) corresponds to the amplitude mean of the i-th frequency bin; obtaining the gain of each frequency bin in the Bark domain corresponding to the reference frequency response curve (Rf); obtaining equalizer parameters using the corresponding gains of the frequency bins in the Bark domain; and converting the first portion of the speech into a reference timbre using the equalizer parameters.
[0007] In a second aspect, the present invention provides a device for converting a speaker's speech into a reference timbre. The device includes a processor configured to convert a first portion of the speaker's speech source signal into the time-frequency domain to obtain a time-frequency signal; obtain a time-varying frequency bin mean of the time-frequency signal; convert the amplitude mean of the frequency bin into the Bark domain to obtain a source frequency response curve (SR), where SR(i) corresponds to the amplitude mean of the i-th frequency bin; obtain the gain of each frequency bin in the Bark domain corresponding to the reference frequency response curve (Rf); obtain equalizer parameters using the corresponding gains of the frequency bins in the Bark domain; and use the equalizer parameters to convert the first portion of the speech into a reference timbre.
[0008] In a third aspect, the present invention proposes a non-transitory computer-readable storage medium, which contains instructions executed by a processor, and the operations that can be performed by the instructions include converting a first part of the speaker's speech source signal into the time-frequency domain to obtain a time-frequency signal; obtaining the frequency bin mean of the time-frequency signal changing with time; converting the amplitude mean of the frequency bin into the Bark domain to obtain a source frequency response curve (SR), where SR(i) corresponds to the amplitude mean of the i-th frequency bin; obtaining the gain of each frequency bin in the Bark domain corresponding to the reference frequency response curve (Rf); using the corresponding gains of the frequency bins in the Bark domain to obtain equalizer parameters; and using the equalizer parameters to convert the first part of the speech into a reference timbre.
[0009] Each of the above aspects can be implemented using a variety of different implementations. For example, each of the above aspects can be implemented using a suitable computer program, which can be implemented on a suitable carrier medium, which can be a tangible carrier medium (such as a disk) or an intangible carrier medium (such as a communication signal). Suitable devices can also be used to implement various aspects of the functions, which can take the form of a programmable computer running a computer program, which is configured to implement the methods and / or techniques described in the present invention. The above aspects can also be used in combination so that the functions described in one aspect of the technology can be implemented in another aspect of the technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The description herein refers to the drawings, wherein like numbers refer to like components throughout the various views.
[0011] Figure 1 3 is a technical example diagram of the preparation stage of timbre style conversion drawn according to an embodiment of the present invention.
[0012] Figure 2 is a Bark filter bank drawn according to an embodiment of the present invention.
[0013] Figure 33 is a technical example diagram of the real-time stage of timbre style conversion drawn according to an embodiment of the present invention.
[0014] Figure 4 is an example block diagram of a computing device drawn according to an embodiment of the present invention.
[0015] Figure 5 FIG. 4 is a flow chart of a technology for converting a speaker's speech into a target timbre according to an embodiment of the present invention. DETAILED DESCRIPTION
[0016] Timbre (also known as timbre quality) is the characteristic of sound that distinguishes one sound from another. For example, when two instruments (such as a piano and a violin) play the same note at the same frequency and amplitude, the sound we hear is different. There are many adjectives to describe timbre, such as sharp, mellow, shrill, brassy, high, magnetic, powerful, soft, flat, melodious, hoarse, breathy, gruff, and bright.
[0017] Different people and different musical styles have different timbres. Simply put, different people and different musical styles sound differently. Sometimes people may want to sound different from their usual voice. In other words, someone may want to change their timbre at certain times (such as during an RTC session). The timbre of a sound (or voice) can be thought of as consisting of different energy levels in different frequency bands.
[0018] The timbre of a sound (such as a recorded voice) can be altered. Professional audio producers, such as broadcasters or music producers, often use complex hardware or software equalizers to change the timbre of different voices or instruments in a recording. For example, a composer might record the entirety of a symphony using a single instrument across multiple tracks. Using an equalizer, the timbre of each track can be modified to match that of the instrument being recorded.
[0019] Using an equalizer essentially involves finding the parameters for adjusting various aspects of the audio spectrum. These parameters include gain (i.e., amplitude) for a specific frequency band, center frequency (i.e., adjusting the center frequency range of a selected frequency band), bandwidth, filter slope (i.e., filter steepness when low-cut or high-cut is selected), and shelving type (i.e., filter shape for the selected frequency band). For example, in terms of gain, the center frequency can be lowered or raised by a certain number of decibels (dB). Bandwidth refers to the frequency range on either side of the center frequency. Changing a specific frequency typically affects other frequencies above or below it. This affected frequency range is called the bandwidth. Different filter types are available, including low-cut, high-cut, low-shelf, high-shelf, notch, bell, and other filter types.
[0020] As can be seen from the above overview and brief description, using an equalizer can be complex, beyond the reach of the average user, and impractical in real-time applications.
[0021] Embodiments designed according to the present invention can be used to convert the timbre of speech, such as the speech of a user in a real-time communication application. As is known, in RTC, there can be a sender user and a receiver user. The sender user's audio stream (such as the sender's speech) can be sent from the sender user's sending device to the receiver user's receiving device. The sender user may wish to change the timbre of their speech to a certain style (such as a reference timbre, an ideal timbre, etc.), or the receiver user may wish to change the timbre of the sender user's voice to a specific style.
[0022] The technology described herein can be implemented on a sender user's device (i.e., a sending device), a recipient user's device (i.e., a receiving device), or on both devices simultaneously. In this context, the sender user may be speaking and their speech will be transmitted to and heard by the recipient user. The technology can also be run by a central server (e.g., a cloud-based server) that can receive audio signals from the sender user and forward them to the recipient user.
[0023] For example, a sending user can use a sending device through the user interface of an RTC application to select to convert the sending user's voice into different timbres before sending it to a receiving user. Similarly, a receiving user can use a receiving device through the user interface of an RTC application to select to convert the sending user's voice into different timbres before the receiving user listens to it (i.e., outputs it to the receiving user). A user may wish to change the timbre to a style that suits a particular situation, such as a news report or a musical style (e.g., jazz, hip-hop, etc.).
[0024] Converting a speaker's timbre (i.e., the speaker's voice) to a reference (e.g., desired, target, selected, etc.) timbre includes a preparatory (e.g., preparation, training, etc.) phase and a real-time phase. In the preparatory phase, a reference frequency response curve for the target (e.g., reference, etc.) timbre is generated. In the real-time phase, the source speech timbre can also be described by a source-domain frequency response curve. Equalizer parameters are obtained by using the difference between the source frequency response curve of the source speech timbre and the reference frequency response curve of the reference speech timbre through a mapping technique, which is described below. The equalizer parameters are then applied to the source speech.
[0025] For example, in the preparation stage, a reference sample of the target timbre can be received, and a Bark frequency response curve can be obtained from the reference sample; in the real-time stage, the Bark frequency response curve can be used to convert the speaker's source speech sample (such as a frame of the source speech sample) into the target timbre in real time.
[0026] The Bark transform is a product of psychoacoustic experiments that define each critical frequency band of human hearing as a Bark scale. The Bark scale represents spectral information processing in the human ear. In other words, the Bark domain reflects the psychoacoustic frequency response, providing useful information about how humans perceive power differences in different frequency bands.
[0027] It's also possible to use other perceptual transformations or scales. For example, the MEL scale could be used. The MEL scale reflects people's perception of pitch, while the Bark scale reflects people's subjective hearing and energy integration. However, the energy distribution in different frequency bands may be more relevant to timbre transformations (e.g., changes) than pitch.
[0028] In some cases, a constant parameter equalizer may not be suitable for long-term use. In other words, a constant parameter equalizer may not be suitable for continuous use in an RTC session. For example, five minutes into an RTC session, the speaker's timbre may change due to a change in mood or singing style; or another person with a completely different timbre may start speaking in place of the original speaker. This timbre change may require dynamic changes to the equalizer parameters so that the changed timbre style can still be converted to the target timbre. Therefore, even if the speaker's timbre changes during an RTC session, the changed timbre can still be converted to the target timbre, requiring dynamic updates to the equalizer parameters.
[0029] This invention primarily describes the timbre conversion of a single speech or sound. If there are multiple speech sounds, techniques such as speech source separation can be used to separate the sounds, and then the timbre conversion techniques described in this invention can be applied to each sound individually. Furthermore, the source speech may be noisy or reverberant. In some examples, the source speech may first be subjected to noise reduction and / or dereverberation processing before timbre conversion is performed using the techniques of this invention.
[0030] Figure 1is an example diagram of technique 100 for the preparation stage of timbre conversion according to an embodiment of the present invention. Technique 100 receives a reference sample of a target timbre and generates a reference (e.g., target) frequency response curve for the target timbre. Technique 100 can be used offline to generate the reference frequency response curve. By way of example, without loss of generality, a speaker may wish to sound like singer Justin Bieber, and therefore a reference voice sample of the singer may be used as a target timbre sample. For another example, a user may wish to sound energetic during an RTC session, and therefore a reference sample of an energetic voice may be used as a reference sample.
[0031] Technique 100 can be repeatedly performed for each desired timbre (e.g., reference timbre) style to generate a corresponding reference frequency response curve (Rf). For example, gender differences can significantly affect timbre, so for the same target timbre, a male reference sample and a female reference sample can be used to obtain two frequency response curves of the desired timbre. The lengths of the two samples (i.e., the male and female samples) can be the same or different.
[0032] At 102, the technique 100 receives a reference speech sample (i.e., a reference signal) of a desired timbre (i.e., target timbre) style. The reference speech sample may include at least one acoustic signal cycle. The reference speech sample may also be in any format. For example, the speech sample may be a waveform audio file (wave or wav file), an MP3 file, a Windows Media Audio file (wma), an Audio Interchange File Format (aiff), or the like. The length of the reference speech sample may be several minutes (e.g., 0.5, 1, 2, 5 minutes, or more or less). For example, the technique 100 may receive a longer speech sample and extract a shorter reference speech sample from it.
[0033] At 104, technique 100 converts the reference speech sample into a transform domain. Technique 100 may use a short-time Fourier transform (STFT) to convert the reference signal into the time-frequency domain. The STFT can be used to obtain the amplitude of each frequency in the reference speech sample over time. As is known, the STFT computes a fast Fourier transform (FFT) over a given window length and frequency hopping period, intercepts multiple samples from the speech sample, and calculates amplitude and phase information over time.
[0034] At 106 , the technique 100 converts the amplitude mean in the time dimension of the time-frequency domain signal into the Bark domain to obtain a reference frequency response curve (Rf) 108 , ie, a psychoacoustic frequency response curve.
[0035] We know that the time domain results of STFT can be displayed on the spectrum diagram, for example Figure 1Schematic spectrogram 120 is shown. Spectrogram 120 shows the spectral density of a signal as frequency varies over time. Time is plotted on the x-axis of spectrogram 120; frequency is plotted on the y-axis of spectrogram 120; and frequency amplitude is typically represented by color depth (i.e., grayscale in spectrogram 120).
[0036] The spectrum diagram 120 shows j frequency bins 122 (B j ,j=0,...,j-1, where j is the number of frequency bins). They can be frequency bins B j Calculate the mean of the amplitude 124 over time Where j = 0, ..., k-1. As the name suggests, the amplitude mean It can be at least one (or all) frequency bins B in all time windows (i.e., time axis, horizontal dimension) j Therefore, each Indicates frequency bin B k For example, for several words pronounced orally, the amplitude mean can represent the average performance of different (types of) word pronunciations in the reference speech sample. calculate where m t,j is the amplitude of the spectrum, t and j represent the time and frequency indices respectively, and n is the last time index of the speech sample.
[0037] According to equation (1), by dividing the FFT frequency bins The amplitude is mapped to the Bark frequency bin, and the amplitude average is converted (i.e., transformed, mapped, etc.) from the STFT domain to the i-th Bark domain amplitude
[0038]
[0039] Equation (1) represents the conversion from Fourier domain to Bark domain. For i = 1, ..., 24, the Bark domain amplitude is Constitutes the reference frequency response curve Rf.
[0040] The Bark scale can range from 1 to 24, corresponding to the first 24 critical frequency bands of hearing. In Hertz (Hz), the Bark band edges include [0, 100, 200, 300, 400, 510, 630, 770, 920, 1080, 1270, 1480, 1720, 2000, 2320, 2700, 3150, 3700, 4400, 5300, 6400, 7700, 9500, 12000, 15500 ]; in Hertz (Hz), the Bark frequency band centers include [50, 150, 250, 350, 450, 570, 700, 840, 1000, 1170, 1370, 1600, 1850, 2150, 2500, 2900, 3400, 4000, 4800, 5800, 7000, 8500, 10500, 13500]. Therefore, i ranges from 1 to 24. For another example, the Bark scale used can contain 109 frequency bins. Therefore, i can range from 1 to 109, and the entire frequency range can be from 0 to 24000 Hz.
[0041] As mentioned above, B in equation (1) i is the FFT frequency bin in the i-th Bark band; the coefficient β ij are the Bark transform parameters. Note that the Bark domain transform can remove any frequency outliers, thereby smoothing the frequency response curve. The Bark transform is an auditory filter bank that can be thought of as calculating a moving average that smoothes the frequency response curve. The coefficient β ij is the triangle shape parameter, see Figure 2 Related mentioned.
[0042] Figure 2 2 is a Bark filter bank according to an embodiment of the present invention. Please note that in order to avoid excessive clutter in the figure, the Bark filter bank 200 only shows 29 frequency bins, with a frequency range of 0 to 8000 Hz. Figure 2 The number of triangles in should be equal to the number of frequency bins. Filter bank 200 is used to illustrate how to obtain the coefficients β in equation (1). ij The coefficient β ij is the Bark transform coefficient of STFT. ij In , index i corresponds to the Bark frequency bin and index j corresponds to the FFT frequency bin. Figure 2 The x-axis in ; index i corresponds to frequency. Each coefficient β ij It must be derived in two dimensions: first by the triangle it uses, and second by The corresponding frequency range is determined.
[0043] Each Bark filter is a triangular bandpass filter with some overlap, such as filter 202. The peaks of the Bark filter bank 200, such as peak 201, represent the center frequencies of the different Bark filters. Figure 2 In the example, some triangles are drawn with thicker lines than others. This is done simply to reduce the appearance of clutter. Therefore, the thickness of the lines drawn for the triangles does not have any special meaning.
[0044] exist Figure 2 In the example, j = 6, j = 6 exists within two triangles (i.e., triangles 204 and 206). Triangle 204 roughly corresponds to the frequency band from 4200 Hz to 6300 Hz; triangle 206 roughly corresponds to the frequency band from 5300 Hz to 7000 Hz. Projecting j = 6 upward, the right side of triangle 204 intersects at point 208, which corresponds to β = 0.2; the left side of triangle 206 intersects at point 210, which corresponds to β = 1.8.
[0045] Without loss of generality, let us further illustrate: Referring to equation (1), in triangle 206, index i=28 (in (in the middle) refers to the 28th triangle or the 28th Bark frequency band; the center frequency is the horizontal axis value (i.e., x value) of the top of the triangle 206 (i.e., the peak 201), which is 6150 Hz; B in equation (1) i represents the frequency bins in the range of 5300 (i.e., frequency bin 212) to 7000 Hz (i.e., frequency bin 214), which is determined by the bottom of triangle 206. i one of the For example, the jth frequency bin is 6000 Hz. According to the y-axis projection of triangle 206 at 6000 Hz, β i=28,j is 1.8. Therefore, for Assuming the FFT size is 1024 and the sampling rate is 16kHz, there are 109 pairs And β ij The range is between 5300 and 7000 Hz, then the formula Calculated
[0046] Figure 3300 is an example diagram of a real-time stage of timbre conversion according to an embodiment of the present invention. In real-time applications, such as audio and / or video conferencing, telephone conversations, etc., the application technology 300 can convert the timbre of the source speech of at least one participant. The technology 300 receives the source speech in the form of frames, such as the source speech frame 302. For another example, the technology 300 itself can divide the received audio signal into frames. One frame can correspond to m milliseconds of audio. For example, m can be 20 milliseconds. Of course, m can also be other values. The technology 300 outputs (such as generates, obtains, produces, calculates, etc.) the transformed speech frame 306. The source speech frame 302 is the source timbre style, and the transformed speech frame 306 is the reference timbre style.
[0047] Technique 300 may be implemented by a computing device, such as Figure 4 The computing device 400 is described in relation thereto.
[0048] Technology 300 can be implemented by a sending device. Therefore, the speaker's timbre style can be converted into a reference timbre on the sending user's device and then sent to the receiving user, so the voice of the sending user received by the receiving user is already the reference timbre. Technology 300 can also be implemented by a receiving device. That is, the voice received at the receiving user's receiving device can be converted into a reference timbre selected by the receiving user. Technology 300 can be executed on the received voice to generate a transformed voice with a reference timbre. The transformed voice is then output to the receiving user. Technology 300 can also be implemented by a central server, which receives a voice sample in the source timbre from the sending device, executes technology 300 to obtain a voice with a reference timbre (such as the desired timbre), and sends (such as forwarding, relaying, etc.) the converted voice to one or more receiving devices.
[0049] The equalizer 304 may process the source speech frame 302 to generate a transformed speech frame 306. As described below, the equalizer 304 transforms the timbre using equalizer parameters that are calculated and subsequently updated when a significant change is detected, as described below.
[0050] Technique 300 is used to obtain (e.g., calculate, search, determine, etc.) the difference between a reference frequency response curve (Rf) of a reference sample and a source frequency response curve (SR) of a source sample. Technique 300 can obtain the difference (e.g., difference value) in each frequency bin. In other words, technique 300 can obtain the amplification gain between the reference frequency response curve (Rf) of the reference sample and the source frequency response curve (SR) of the source sample. For example, a logarithmic calculation can be used to obtain the gain. The difference value can be expressed in decibels (dB). As we know, decibels (dB) are the logarithmic ratio between two quantities, which helps to realistically model human auditory perception.
[0051] For each kth frequency bin in the Bark domain, technique 300 may use equation (2) to calculate the dB difference G between the source psychoacoustic frequency response curve and the reference psychoacoustic frequency response curve. b (k).
[0052] G b (k)=20*log(Rf(k) / SR(k)) (2)
[0053] To reiterate, equation (2) can be used to measure the amplification gain between the reference frequency response curve (Rf) and the source frequency response curve (SR) in each Bark frequency bin. The gain G for all Bark frequency bins is b The (k) set may constitute (eg, may be considered as) parameters of the equalizer 304. Thus, the equalizer 304 converts the timbre of the source speech into the reference timbre using the equalizer parameters.
[0054] The equalizer 304 is a set of filters. For example, the equalizer 304 may include a filter for the lower frequency f n (eg, 0Hz) to a higher frequency f n+1 (such as 800Hz) band filtering, the center frequency is (f n +f n+1 ) / 2 (eg 400 Hz). The equalizer 304 may use the equalizer parameters (ie G b (k)) Gain) is used to adjust the center frequency. This parameter determines how much the center frequency needs to be increased or decreased.
[0055] Interpolation parameters can then be derived that calculate the adjusted center frequency as an interpolation between the lower and upper frequencies of the frequency band. The interpolation parameters can also include (e.g., determining, defining, etc.) the shape of the interpolation. For example, the interpolation can be a cubic or cubic spline interpolation. Cubic spline interpolation can make the interpolation smoother than linear interpolation. The following equation (3) can be used to explain how to obtain the i-th gain The interpolation method is the cubic spline interpolation method. In equation (3), the interpolation parameter a i to d i By G close to the i-th center frequency in the equalizer b (i) Derive.
[0056]
[0057] Equalizer 304 may include (e.g., use) an initial set of equalizer parameters. For example, the initial set of equalizer parameters may be obtained by previously running technique 300. For example, memory 322 may include stored reference response curves, stored source frequency response curves, and / or corresponding equalizer parameters. Thus, memory 322 may include a reference frequency response curve 322 for a reference timbre style. Memory 322 may be a permanent memory (e.g., a database, a file, etc.) or a non-permanent memory. As another example, equalizer 304 may not include equalizer parameters. In this case, the initial equalizer parameters may be obtained by following steps 314-318.
[0058] Because equalizer 304 can add or subtract different amounts of gain for different Bark frequency bands, the total energy of the source signal may be changed. For example, technique 300 can normalize the gains so that the volume of the speech remains at the same (or approximately the same) level before and after being adjusted by equalizer 304. For another example, normalizing the gains can mean dividing each gain by the sum of all gains. Of course, other normalization methods can also be used.
[0059] When a large change in the speech signal is detected (as described below), technique 300 may perform operations 308 - 318 to obtain initial equalizer parameters.
[0060] The source speech frame 302 is received into a signal buffer 308, which can store the received speech frames and accumulate source speech samples to a certain length for further processing. For example, the time period of the source audio can be 30 seconds, 1 minute, 2 minutes, or a longer or shorter time period.
[0061] At 310, technique 300 converts the speech sample to a transform domain, such as Figure 1 As described in 104. The speech sample (i.e., the source audio of a period of time) is converted to the STFT domain. At 312, the technique 300 converts the amplitude mean in the time dimension of the time-frequency domain signal to the Bark domain to obtain the source frequency response curve (SR). The corresponding reference frequency response curve (Rf) and Figure 1 As described in step 106, a source frequency response curve (SR) can be obtained. Therefore, the source frequency response curve (SR) can be the Bark domain amplitude of the source sample A collection of .
[0062] At 314, technique 314 determines whether the timbre of the source voice has changed significantly. The change can be learned at 316. To illustrate without loss of generality, during an RTC session, the source voice can be the voice of a first speaker (e.g., a 45-year-old man). However, at some point during the RTC session, a second speaker (e.g., a 7-year-old girl) begins speaking. Thus, the source voice has changed significantly. Therefore, in this example, technique 300 can replace the source frequency response curve (originally that of the first speaker) with the source frequency response curve of the second speaker. For example, technique 300 will only replace the source frequency response curve if there is a significant change between the stored source frequency response curve and the current source frequency response curve. As described above, it can be learned at 314 that a significant change has occurred before the equalizer parameters have been obtained (e.g., initialized).
[0063] At 314, if the speech signal has not changed significantly, the technique 300 jumps to 304, where the previous equalizer parameters are still used. However, if the speech signal has changed significantly, then at 314, the technique 300 stores the current source frequency response curve in memory 322 so that the current source frequency response curve can be compared with subsequent source frequency response curves to detect any subsequent significant changes. The technique 300 can also jump to 318 to update the equalizer parameters. That is, the technique 300 obtains interpolated parameters as described in relation to equation (3).
[0064] At 314, a correlation threshold can be set to detect significant changes. A correlation coefficient can be calculated for the frequency response curve for the current time period and the stored frequency response curve (e.g., stored at 322). If the correlation coefficient is greater than the threshold, the stored frequency response curve is replaced with the current curve, and the equalizer parameters are updated. Otherwise, the equalizer and the stored frequency response curve are not updated.
[0065] The update of the equalizer parameters can be completed (e.g., executed, completed, etc.) within one frame (e.g., 10ms) of the source speech signal. Therefore, the voice style conversion implemented according to the present invention is not interrupted by the update of the equalizer parameters. In other words, there is no delay or discontinuity when updating the equalizer parameters.
[0066] Figure 4 4 is a schematic block diagram of a computing device according to an embodiment of the present invention. Computing device 400 may be a computing system including multiple computing devices, or a single computing device such as a mobile phone, tablet computer, laptop computer, notebook computer, desktop computer, etc.
[0067] The processor 402 in the computing device 400 can be a conventional central processing unit. The processor 402 can also be another type of device or multiple devices capable of manipulating or processing existing or later developed information. For example, although the examples herein may be implemented with a single processor (such as the processor 402), using multiple processors may provide advantages in speed and efficiency.
[0068] In one implementation, the memory 404 in the computing device 400 may be a read-only memory (ROM) device or a random access memory (RAM) device. Other appropriate types of storage devices may also be used as the memory 404. The memory 204 may contain code and data 406 accessed by the processor 402 using a bus 412. The memory 404 may also contain an operating system 408 and application programs 410, wherein the application programs 410 include at least one program that allows the processor 402 to execute one or more of the techniques described herein. For example, the application programs 410 may include application programs 1 through N, which include programs and techniques that may be used in implementing real-time voice timbre style conversion applications. For example, the application programs 410 may include technique 100 or its various techniques to implement a training phase. For example, the application programs 410 may include technique 300 or its various techniques to implement real-time voice timbre style conversion functionality. The computing device 400 may also include an auxiliary storage device 414, such as a memory card used with a mobile computing device.
[0069] Computing device 400 may also include one or more output devices, such as display 418. For example, display 418 may be a touch-sensitive display that combines a display with touch-sensitive elements operable for touch input. Display 418 may be coupled to processor 402 via bus 412. Other output devices that allow a user to program or use computing device 400 may also be used in addition to or in place of display 418. If the output device is or includes a display, the display may be implemented in various ways, including a liquid crystal display (LCD), a cathode ray tube (CRT) display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.
[0070] The computing device 400 may also include an image sensing device 420 (e.g., a camera), or any other image sensing device 420 now known or later developed that can sense an image (e.g., an image of a user operating the computing device 400), or communicate with the image sensing device 420. The image sensing device 420 may be positioned to face the user operating the computing device 400. For example, the position and optical axis of the image sensing device 420 may be configured such that the field of view includes an area directly adjacent to and visible to the display 418.
[0071] The computing device 400 may also include a sound sensing device 422 (such as a microphone), or any other sound sensing device 422 that is currently available or later developed and can sense sounds near the device 400, or communicate with the above-mentioned sound sensing device 422. The sound sensing device 422 can be placed in a position facing the user operating the computing device 400 and can be configured to receive sounds, and can be configured to receive sounds, such as sounds made by the user when the user operates the computing device 400, such as voice or other sounds. The computing device 400 may also include or communicate with a sound playing device 424, such as a speaker, a headset, or any other device currently available or later developed that can play sounds according to the instructions of the computing device 400.
[0072] Figure 4 The processor 402 and memory 404 of the computing device 400 are depicted as being integrated into a single processing unit, although other configurations are possible. The operations of the processor 402 may be distributed across multiple machines (each machine containing one or more processors), which may be directly coupled or coupled across a local or other network. The memory 404 may be distributed across multiple machines, such as network-based storage or storage in multiple machines that run the operations of the computing device 400. While only a single bus is described herein, the bus 412 of the computing device 400 may also be comprised of multiple buses. Furthermore, the auxiliary memory 414 may be directly coupled to other components of the computing device 400, may be accessed via a network, or may comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, the computing device 400 may be implemented in a variety of configurations.
[0073] Figure 5 5 is a flow chart of a technique for converting a speaker's speech into a target timbre according to an embodiment of the present invention. For example, technique 500 may receive an audio sample, such as a speech stream. The audio stream may be part of a video stream. For another example, technique 500 may receive frames of an audio stream and then process them. For another example, technique 500 may divide the audio sample into frames and process them according to the Figure 3 The technique 300 in processes each frame separately, as described below.
[0074] The technique 500 can be performed by a computing device such as Figure 4Technique 500 may be implemented as a software program executed by a computing device (such as computing device 400). The software program may include machine-readable instructions that may be stored in a memory (such as memory 404 or secondary memory 414) and that, when executed by a processor (such as processor 402), cause the computing device to perform technique 500. Technique 500 may be implemented using dedicated hardware or firmware. Multiple processors and / or multiple memories may also be used.
[0075] At 502, technique 500 converts a portion of a source signal of a speaker's speech into a time-frequency domain to obtain a time-frequency signal, as described above. At 504, as described above with respect to As described above, technique 500 obtains the time-varying frequency bin amplitude mean of a time-frequency signal. At 506, as described above, technique 300 converts the frequency bin amplitude mean to the Bark domain to obtain a source frequency response curve (SR). SR(i) is the amplitude mean of the i-th frequency bin.
[0076] At 508, the technique 500 obtains the corresponding gain of the frequency bin in the Bark domain for the reference frequency response curve (Rf). The method for obtaining the reference frequency response curve (Rf) is described above. Therefore, as described above, the technique 300 may include: receiving a reference sample of a reference timbre; converting the reference sample into the time-frequency domain to obtain a reference time-frequency signal; obtaining the reference frequency bin amplitude mean of the reference time-frequency signal over time. The reference frequency bin amplitude mean Convert to Bark domain to obtain reference frequency response curve (Rf). The reference frequency response curve (Rf) includes the Bark domain frequency amplitude corresponding to each Bark domain frequency bin i. Therefore, Rf(i) is the mean amplitude of the i-th frequency bin.
[0077] As described above, technique 500 can use equation (1) to convert the amplitude mean of the reference frequency bins Convert to the Bark domain to obtain a reference frequency response curve (Rf). As described above, obtaining each gain of the frequency bin in the Bark domain may include: using the ratio of the reference frequency bin amplitude mean of the kth frequency bin to the source frequency response curve (SR) of the kth frequency bin to calculate the gain G of the kth frequency bin in the Bark domain b (k). Gain G b (k) can be calculated by equation (2).
[0078] At 510, the technique 500 may use the corresponding gains of the frequency bins in the Bark domain to obtain equalizer parameters. For example, using the corresponding gains of the frequency bins in the Bark domain to obtain equalizer parameters may include: mapping the corresponding gains to the corresponding center frequencies of the equalizer to obtain the gain values of the equalizer. For example, the technique 500 may normalize the respective gains to obtain the equalizer parameters. At 512, the technique 500 uses the equalizer parameters to convert the first portion of speech into a reference timbre. For example, without loss of generality, assume that we select an equalizer with 30 frequency bands, from fc1 to fc 30 , where the center frequency of the band is fc i ; Then the gain of each frequency band of the equalizer can be the interpolation gain The interpolation gain is obtained according to equation (3).
[0079] Regarding how to detect a situation where a speech signal has undergone a significant change, technology 500 may further include the following steps: obtaining a second source frequency response curve for a second portion of the source signal; if it is detected that the difference between the source frequency response curve and the second source frequency response curve exceeds a threshold, obtaining new equalizer parameters and using the new equalizer parameters as equalizer parameters; and using the equalizer parameters to transform the second portion of the source signal (if a significant change is detected, the new equalizer parameters are used here).
[0080] To simplify the explanation, Figure 1 、 Figure 3 and Figure 5 The techniques 100, 300, and 500 are each illustrated as a series of modules, steps, or operations. However, according to the present invention, these modules, steps, or operations may occur in various orders and / or simultaneously. In addition, other steps or operations not mentioned or described herein may also be used. Furthermore, the techniques designed according to the present invention may not require all of the steps or operations shown to be implemented.
[0081] The word "example" is used herein to mean an example, instance, or illustration. Any feature or design described herein as an "example" is not necessarily superior or preferable to other features or designs. Instead, the word "example" is used to present concepts in a concrete way. The word "or" as used herein is intended to mean an inclusive "or" rather than an exclusive "or." That is, "X includes A or B" is intended to mean any natural inclusive permutation, unless otherwise specified or clear from the context. In other words, if X includes A, X includes B, or X includes A and B, then "X includes A or B" holds true in any of the aforementioned instances. Furthermore, throughout this application and the appended claims, "a" or "an" should generally be interpreted to mean "one or more," unless otherwise specified or the context clearly indicates the singular form. Furthermore, throughout this document, the phrases "a feature" or "a function" do not imply the same embodiment or function, unless specifically stated otherwise.
[0082] Figure 4 The computing device 400 shown and / or any components thereof and Figure 1 or Figure 3 Any modules or components shown (as well as the techniques, algorithms, methods, instructions, etc. stored thereon and / or executed thereby) may be implemented using hardware, software, or any combination thereof. Hardware includes, for example, intellectual property (IP) cores, application-specific integrated circuits (ASICs), programmable logic arrays, optical processors, programmable logic controllers, microcode, firmware, microcontrollers, servers, microprocessors, digital signal processors, or any other suitable circuitry. Throughout the present invention, the term "processor" should be understood to encompass any combination of one or more of the foregoing. Terms such as "signal" and "data" may be used interchangeably.
[0083] Furthermore, the techniques may be implemented using a general-purpose computer or processor with a computer program that, when executed, executes any corresponding technique, algorithm, and / or instruction described herein. Alternatively, a dedicated computer or processor equipped with dedicated hardware may be used to implement any method, algorithm, or instruction described herein.
[0084] Furthermore, all or part of the present invention may take the form of a computer program product that can be used by a computer or accessed by a computer-readable medium. A computer-usable or computer-readable medium can be any device that can contain, store, communicate, or transport a program or data structure for use by or in connection with any processor. The medium can be an electronic, magnetic, optical, electromagnetic, or semiconductor device, among others. Other suitable media may also be included.
[0085] Although the present invention has been described in conjunction with certain embodiments, it should be understood that the invention is not limited to the disclosed embodiments. On the other hand, the present invention is intended to cover various modifications and equivalent arrangements within the scope of the claims, which should be given the broadest interpretation to cover all such modifications and equivalent arrangements permitted by law.
Claims
1. A method for converting a speaker's speech into a reference timbre, comprising: converting a first portion of the speaker's speech source signal into a time-frequency domain to obtain a time-frequency signal; Obtain the frequency bin amplitude mean of the time-frequency signal changing with time; The frequency bin amplitude mean is converted to the Bark domain to obtain the source frequency response curve (SR), where SR(i) corresponds to the amplitude mean of the i-th frequency bin; Corresponding to the reference frequency response curve (Rf), the difference between the reference frequency response curve (Rf) and the source frequency response curve (SR) is calculated to obtain the gain of each frequency bin in the Bark domain; Obtaining equalizer parameters using corresponding gains of frequency bins in the Bark domain, wherein the equalizer parameters are updated according to changes in the speech source signal; Use the equalizer parameters to convert the first part of the speech into a reference timbre; as well as obtaining a second source frequency response curve of a second portion of the source signal; If the difference between the source frequency response curve and the second source frequency response curve exceeds the threshold, then Get the new equalizer parameters, and Use the new EQ parameters as EQ parameters; The second portion of the source signal is transformed using the equalizer parameters.
2. The method according to claim 1, further comprising: receiving a reference sample of a reference timbre; Converting the reference sample into the time-frequency domain to obtain a reference time-frequency signal; Obtain the reference frequency bin amplitude mean of the reference time-frequency signal that changes with time ( );as well as The reference frequency bin amplitude mean ( ) is converted to the Bark domain to obtain the reference frequency response curve (Rf).
3. The method according to claim 2, wherein the reference frequency bin amplitude mean ( ) is converted to the Bark domain to obtain the reference frequency response curve (Rf) including: Using the equation , in is the FFT frequency bin in the ith Bark band, and in is the transformation parameter of Bark transform.
4. The method of claim 2 , wherein obtaining corresponding gains of frequency bins in the Bark domain comprises: The gain of the kth frequency bin in the Bark domain is calculated using the ratio of the reference frequency bin amplitude mean of the kth frequency bin to the source frequency response curve (SR) of the kth frequency bin. .
5. The method according to claim 4, wherein G b (k) According to equation G b (k)=20·log(Rf(k) / SR(k)) is calculated.
6. The method of claim 1 , wherein obtaining equalizer parameters using corresponding gains of frequency bins in the Bark domain comprises: The corresponding gains are normalized to obtain the equalizer parameters.
7. The method of claim 6, wherein obtaining equalizer parameters using corresponding gains of frequency bins in the Bark domain further comprises: The corresponding gains are mapped to the corresponding center frequencies of the equalizer to obtain the gain values of the equalizer.
8. The method according to claim 1, further comprising: A reference timbre is received from a speaker.
9. A device for converting a speaker's voice into a reference timbre, comprising A processor configured to: converting a first portion of the speaker's speech source signal into a time-frequency domain to obtain a time-frequency signal; Obtain the frequency bin amplitude mean of the time-frequency signal changing with time; Convert the frequency bin amplitude mean to the Bark domain to obtain the source frequency response curve (SR), where SR(i) corresponds to the amplitude mean of the i-th frequency bin; Corresponding to the reference frequency response curve (Rf), the difference between the reference frequency response curve (Rf) and the source frequency response curve (SR) is calculated to obtain the gain of each frequency bin in the Bark domain; Obtaining equalizer parameters using corresponding gains of frequency bins in the Bark domain, wherein the equalizer parameters are updated according to changes in the speech source signal; Use the equalizer parameters to convert the first part of the speech into a reference timbre; as well as obtaining a second source frequency response curve of a second portion of the source signal; If the difference between the source frequency response curve and the second source frequency response curve exceeds the threshold, then Get the new equalizer parameters, and Use the new EQ parameters as EQ parameters; The second portion of the source signal is transformed using the equalizer parameters.
10. The apparatus of claim 9, wherein the processor is configured to further perform the following operations: receiving a reference sample of a reference timbre; Converting the reference sample into the time-frequency domain to obtain a reference time-frequency signal; Obtain the reference frequency bin amplitude mean of the reference time-frequency signal that changes with time ( );as well as The reference frequency bin amplitude mean ( ) is converted to the Bark domain to obtain the reference frequency response curve (Rf).
11. The apparatus according to claim 9, wherein the reference frequency bin amplitude mean ( ) to the Bark domain to obtain the reference frequency response curve (Rf) includes: Using the equation , in is the FFT frequency bin in the ith Bark band, and in is the transformation parameter of Bark transform.
12. The apparatus of claim 9, wherein obtaining corresponding gains of frequency bins in the Bark domain comprises: The gain of the kth frequency bin in the Bark domain is calculated using the ratio of the reference frequency bin amplitude mean of the kth frequency bin to the source frequency response curve (SR) of the kth frequency bin. .
13. The apparatus according to claim 9, wherein G b (k) According to equation G b (k)=20·log(Rf(k) / SR(k)) is calculated.
14. The apparatus of claim 9, wherein obtaining equalizer parameters using corresponding gains of frequency bins in the Bark domain comprises: The corresponding gains are normalized to obtain the equalizer parameters.
15. The apparatus of claim 14, wherein obtaining equalizer parameters using corresponding gains of frequency bins in the Bark domain further comprises: The corresponding gains are mapped to the corresponding center frequencies of the equalizer to obtain the gain values of the equalizer.
16. The apparatus of claim 9, wherein the processor is configured to further perform the following operations: A reference timbre is received from a speaker.
17. A non-transitory computer-readable storage medium, the storage medium containing instructions for execution by a processor, wherein the instructions may perform the following operations: converting a first portion of the speaker's speech source signal into a time-frequency domain to obtain a time-frequency signal; Obtain the frequency bin amplitude mean of the time-frequency signal changing with time; Convert the frequency bin amplitude mean to the Bark domain to obtain the source frequency response curve (SR), where SR(i) corresponds to the amplitude mean of the i-th frequency bin; Corresponding to the reference frequency response curve (Rf), the difference between the reference frequency response curve (Rf) and the source frequency response curve (SR) is calculated to obtain the gain of each frequency bin in the Bark domain; Obtaining equalizer parameters using corresponding gains of frequency bins in the Bark domain, wherein the equalizer parameters are updated according to changes in the speech source signal; Use the equalizer parameters to convert the first part of the speech into a reference timbre; as well as obtaining a second source frequency response curve of a second portion of the source signal; If the difference between the source frequency response curve and the second source frequency response curve exceeds the threshold, then Get the new equalizer parameters, and Use the new EQ parameters as EQ parameters; The second portion of the source signal is transformed using the equalizer parameters.
Citation Information
Patent Citations
Audio processing method, sound processing device, electronic equipment and readable medium
CN109686347A