Audio Transpose
The audio source separation and transposition method in karaoke systems adjusts the pitch of accompaniment signals to match the user's vocal range, addressing vocal strain and enhancing performance quality.
Patent Information
- Application Number
- JP2022575932
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-16
- Filing Date
- 2021-06-14
- Publication Date
- 2025-10-07
- Estimated Expiration
- 2041-06-14
AI Technical Summary
Karaoke systems often require karaoke singers to adapt their pitch to match the original song, leading to vocal strain and compromised performance quality, as existing audio transposition methods do not adequately account for individual vocal capabilities.
An electronic device and method for audio source separation that separates a first audio input signal into a vocal and accompaniment signal, and transposes the audio output based on a pitch ratio comparison between the vocal signals, optimizing the pitch range for the user.
Enables accurate pitch adjustment to match the user's vocal range, reducing singing effort and minimizing vocal strain while maintaining performance quality.
Smart Images

Figure 0007750250000013 
Figure 0007750250000014 
Figure 0007750250000015
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to the field of audio processing, and more particularly to an apparatus, method, and computer program for audio transposition. [Background technology]
[0002] For example, there is a great deal of audio content available in the form of compact discs (CDs), tapes, audio data files downloadable from the Internet, as well as video soundtracks stored on, for example, digital video discs.
[0003] When a music player plays a song from an existing music database, a listener may want to sing along. Of course, the listener's voice adds to and potentially interferes with the voice of the original artist present in the recording. This can hinder or distort the listener's own interpretation of the song. Therefore, karaoke systems provide playback of songs in the musical key of the original song recording for the karaoke singer to sing along to. This can lead to the karaoke singer reaching a pitch range beyond their capabilities (i.e., too high or too low). This requires high singing effort from the karaoke singer to reach the pitch range of the original song, and therefore the karaoke singer may not be able to endure long singing sessions or may damage their vocal cords. This can also force the karaoke singer to adapt their pitch to reduce their effort and protect their vocal cords, thus compromising the overall quality of the performance.
[0004] While techniques for audio transposition generally exist, it is generally desirable to improve methods and apparatus for transposing audio content. Summary of the Invention
[0005] According to a first aspect, the present disclosure relates to an electronic device comprising: a circuit configured to perform audio source separation to separate a first audio input signal into a first vocal signal and an accompaniment signal, and to transpose the audio output signal by a transposition value based on a pitch ratio, the pitch ratio being based on a comparison between a first pitch range of the first vocal signal and a second pitch range of a second vocal signal.
[0006] According to a second aspect, the present disclosure relates to a method for separating a first audio input signal into a first vocal signal and an accompaniment signal, and transposing the audio output signal by a transposition value based on a pitch ratio, the pitch ratio being based on a comparison between a first pitch range of the first vocal signal and a second pitch range of the second vocal signal.
[0007] Further aspects are set forth in the dependent claims, the following description and the drawings.
[0008] Embodiments will now be described, by way of example, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]
[0009] [Figure 1] 1 illustrates a first embodiment of a process for a karaoke system for automatically transposing an audio signal based on audio source separation and pitch range estimation. [Figure 2] 1 illustrates a general approach for audio upmix / remix by blind source separation (BSS), such as multi-source separation (MSS). [Figure 3] 2 illustrates in more detail an embodiment of the pitch analysis process performed in the pitch analysis unit in FIG. 1. [Figure 4] 2 is a flowchart illustrating the process of the pitch range determination unit of FIG. 1; [Figure 5] 10A and 10B are graphs showing pitch analysis results. [Figure 6] 1 shows a flow chart illustrating the process of the pitch range comparator of FIG. [Figure 7] 1 shows a flow chart illustrating the process of the comparison unit in FIG. [Figure 8] 10 illustrates a second embodiment of a karaoke system process for transposing an audio signal based on audio source separation and pitch range estimation; [Figure 9] The singing effort determination unit of FIG. 8 will now be described in brief. [Figure 10] 9 shows a schematic diagram of the transposition value determination unit of FIG. 8; [Figure 11] 10 illustrates a schematic diagram of a third embodiment of a process for a karaoke system for transposing an audio signal based on audio source separation and pitch range estimation. [Figure 12] 10 illustrates a schematic diagram of a fourth embodiment of a process for a karaoke system for transposing an audio signal based on audio source separation and pitch range estimation. [Figure 13] 10 illustrates a fifth embodiment of a process for a karaoke system for transposing an audio signal based on source separation and pitch range estimation; [Figure 14] An embodiment of an electronic device capable of implementing the above-described pitch range determination and transposition process will now be described generally. DETAILED DESCRIPTION OF THE INVENTION
[0010] General Description Before describing the configuration in detail with reference to FIG. 1 et seq., some general remarks will be made.
[0011] An embodiment relates to an electronic device comprising circuitry configured to perform audio source separation to separate a first audio input signal into a first vocal signal and an accompaniment, and to transpose the audio output signal by a transposition value based on a pitch ratio, the pitch ratio being based on a comparison between a first pitch range of the first vocal signal and a second pitch range of a second vocal signal.
[0012] The electronic device may be any music or video playback device, such as a karaoke booth, a smartphone, a PC, a TV, a synthesizer, or a mixing console.
[0013] The circuitry of an electronic device may include a processor, for example a CPU, memory (RAM, ROM, etc.), memory and / or storage, an interface, etc. The circuitry may comprise or be connected to input means (mouse, keyboard, camera, etc.), output means (display (e.g., liquid crystal, (organic) light-emitting diode, etc.)), speakers, etc., (wireless) interfaces, etc., which are commonly known as electronic devices (computers, smartphones, etc.). Furthermore, the circuitry may comprise or be connected to sensors (image sensors, camera sensors, video sensors, etc.) for sensing still image or video image data.
[0014] The input signal may be any type of audio signal. The input signal may be in the form of an analog signal, a digital signal, originating from a hard disk, compact disc, digital video disc, etc., or may be a data file such as a wave file, mp3 file, etc., and the present disclosure is not limited to a particular format for the input audio content. For example, the input audio content may be a stereo audio signal having a first channel input audio signal and a second channel input audio signal, and the present disclosure is not limited to input audio content having two audio channels. In other embodiments, the input audio content may include any number of channels, such as a remix of a 5.1 audio signal.
[0015] The input signal may include one or more source signals. In particular, the input signal may include several audio sources. An audio source may be any entity that generates sound waves, such as a musical instrument, a voice, a vocal, an artificially generated sound (e.g., generated from a synthesizer), etc.
[0016] The input audio content may represent or include mixed audio sources, meaning that sound information is not available individually for all audio sources of the input audio content, but that sound information for different audio sources is, for example, at least partially overlapping or mixed. The accompaniment may be a residual signal resulting from separating a vocal signal from an audio input signal. For example, the audio input signal may be a musical piece including vocals, guitar, keyboard, and drums, and the accompaniment signal may be a signal including guitar, keyboard, and drums as a residue after separating the vocals from the audio input signal.
[0017] Transposition involves changing the pitch of the tones of a piece of music at intervals, or shifting the entire piece to a different key at intervals.
[0018] A pitch ratio can be the ratio between two pitches. Transposing by pitch ratio can mean shifting the pitch of a tone in a piece of music by the ratio between two pitches, or shifting the entire piece to a different key by the number of semitones defined by the ratio between the two pitches.
[0019] Blind source separation (BSS), also known as blind signal separation, is the separation of a set of source signals from a set of mixed signals. One application of BSS is separating a musical piece into individual instrument tracks so that the original content can be upmixed or remixed.
[0020] In the following, the terms remix, upmix, and downmix can refer to the overall process of generating output audio content based on separated audio source signals resulting from mixed input audio content, while the term "mix" can refer to a mix of separated audio source signals. Thus, a "mix" of a separated audio source signal is also a "remix," "upmix," or "downmix" of the input audio content source.
[0021] In audio source separation, an input signal containing multiple sources (e.g., instruments, voices, etc.) is decomposed and separated. Audio source separation can be unsupervised (called "blind source separation" or BSS) or partially supervised. "Blind" means that blind source separation does not necessarily have information about the original sources. For example, it is not necessary to know how many audio sources the original signal contains or which sound information in the input signal belongs to which original audio source. The goal of blind source separation is to decompose the original signals so that they are separated without knowing the previous separation. The blind source separator can use any blind source separation technique known to those skilled in the art. (Blind) audio source separation can search for minimally correlated, i.e., maximally independent, audio source signals in a probabilistic or information-theoretic sense, or based on nonnegative matrix factorization structure constraints on the audio source signals. Methods for performing (blind) source separation are known to those skilled in the art and are based, for example, on principal component analysis, singular value decomposition, (independent) component analysis, nonnegative matrix factorization, artificial neural networks, etc.
[0022] Although some embodiments use blind source separation to generate the separated audio source signals, this disclosure is not limited to embodiments in which no additional information is used for separating the audio source signals, and in some embodiments, additional information is used for generating the separated audio source signals, such as information about the mixing process, information about the types of audio sources included in the input audio content, information about the spatial locations of the audio sources included in the input audio content, etc.
[0023] The circuit may be configured to perform a remix or upmix based on at least one filtered separated source and based on other separated sources obtained by blind source separation to obtain a remixed or upmixed signal. The remix or upmix may be configured to perform a remix or upmix of the separated sources, here "vocals" and "accompaniment," to generate a remixed or upmixed signal, which may be transmitted to a speaker system. The remix or upmix may further be configured to perform lyric substitution on one or more of the separated sources to generate a remixed or upmixed signal, which may be transmitted to one or more of the output channels of the speaker system.
[0024] According to some embodiments, the circuitry may be further configured to determine a first pitch range of the first vocal signal based on a first pitch analysis of the first vocal signal, and may be configured to determine a second pitch range of the second vocal signal based on a second pitch analysis of the second vocal signal.
[0025] According to some embodiments, the first vocal signal comprises the audio input signal.
[0026] According to some embodiments, the audio output signal may be musical accompaniment.
[0027] According to some embodiments, the audio output signal may be the audio input signal.
[0028] According to some embodiments, the audio output signal may be a mix of musical accompaniment and a first vocal signal.
[0029] According to some embodiments, it may be further configured to separate the accompaniment into multiple instruments.
[0030] According to some embodiments, the second audio input signal may be separated into a second vocal signal and a residual signal.
[0031] According to some embodiments, the circuitry may be further configured to determine singing effort based on the second vocal signal, the transposition value being based on the singing effort and the pitch ratio.
[0032] According to some embodiments, the singing effort may be based on a second pitch analysis of the second vocal signal and a second pitch range of the second vocal signal.
[0033] According to some embodiments, the circuitry may be further configured to determine singing effort based on a jitter value and / or a RAP value and / or a shimmer value and / or an APQ value and / or a noise-to-harmonics ratio and / or a soft voicing index.
[0034] According to some embodiments, the circuitry may be further configured to transpose the audio output signal based on the pitch ratio, such that the transposition values correspond to integer multiples of semitones.
[0035] The transposition value may be rounded to the ceiling or to the floor, and thereby to the next integral multiple of a semitone. Thus, the accompaniment may be transposed by integral multiples of a semitone.
[0036] According to some embodiments, the circuitry may include a microphone configured to capture the second vocal signal.
[0037] According to some embodiments, the circuitry may be further configured to capture the first audio input signal from a real audio recording.
[0038] A real audio recording may be, for example, any recording of music recorded with a microphone, as opposed to computer-generated sound. The real audio recording may be stored in a suitable audio file, such as WAV, MP3, AAC, WMA, AIFF, etc. That is, the audio input may be real audio, which means live, unprepared audio that is not, for example, a commercially performed song.
[0039] According to this embodiment, a method is disclosed that includes separating a first audio input signal into a first vocal signal and an accompaniment, and transposing the audio output signal by a transposition value that is based on a pitch ratio, the pitch ratio being based on a comparison between a first pitch range of the first vocal signal and a second pitch range of the second vocal signal.
[0040] According to this embodiment, a computer program is disclosed comprising instructions which, when executed on a processor, cause the processor to perform a method comprising separating a first audio input signal into a first vocal signal and an accompaniment, and transposing the audio output signal by a transposition value based on a pitch ratio, the pitch ratio being based on a comparison between a first pitch range of the first vocal signal and a second pitch range of the second vocal signal.
[0041] Hereinafter, this embodiment will be described with reference to the drawings.
[0042] Figure 1 shows a schematic diagram of a first embodiment of a karaoke system process for automatically transposing an audio signal based on source separation and pitch range estimation. An audio input signal x(n) received from a mono or stereo audio input 13 contains multiple sources (see 1, 2, ..., K in Figure 2), which are input to a source separation 12 process and separated (see separated source 2 and residual signal 3 in Figure 2), where separated source 2, i.e., the original vocal s original (n), and residual signal 3, i.e., accompaniment s Acc An exemplary embodiment of the process of Source Separation 2 is described in FIG. 2 below.
[0043] The audio output signal x*(n) is the accompaniment s Acc (n) and the audio output signal x * (n) is sent to transposition unit 17, and the original vocal s original (n) is sent to the signal adder 18 and the pitch analyzer 14 (see FIG. 3 for details). The pitch analyzer 14 calculates the original vocal s original Pitch analysis result of (n) ω f , original (n) is estimated. Pitch analysis result ω f,original (n) is input to the pitch range estimation unit 15 (described in detail in FIG. 4). The pitch range estimation unit 15 estimates the original vocal s original (n) pitch range R ω,original Estimate the pitch range R ω,original is input to the pitch comparison unit 16. The user microphone 11 acquires the separated audio input signal y(n) input to the sound source separation process 12 (see the separated sound source 2 and residual signal 3 in FIG. 2). user (n) and the unwanted residual signal 3 below. original (n) is user vocals original Pitch analysis result of (n) ω f,user The pitch analysis result ω is transmitted to the pitch analysis unit 14 (see FIG. 3 for details) which estimates the pitch (n). f,user (n) is user vocals user (n) pitch range Rω,user is input to a pitch range estimation unit 15 (described in detail in FIG. 4) which estimates the pitch range R ω,user is input to the pitch comparator 16. The pitch range estimator 15 (described in detail in FIG. 5) calculates the original vocal s original (n) pitch range R ω,original , and user vocals user (n) pitch range R ω,user Receive original vocals original (n) pitch range R ω,original The average value of the user vocal s user (n) pitch range R ω,user The average pitch ratio P ω The pitch ratio P ω is input to the transposition unit 17 (described in detail in FIG. 6). The transposition unit 17 converts the pitch ratio P ω It accepts as input a transpose amount transpose_val equal to x and outputs an audio signal x * (n)(=accompaniment s Acc (n)) and the audio output signal x * (n)(=accompaniment s Acc (n)) with the pitch ratio P ω The transposition unit 17 transposes the accompaniment s after transposition. * Acc (n) is output and input to the signal adder 18. The signal adder 18 outputs the transposed accompaniment s * Acc (n) and original vocals original (n) are input and added, and the added signal is output to the speaker system 19. ω The user's vocal s is further output to the display unit 20 and presented to the user. user (n) receives the lyrics and presents them to the user.
[0044] In the embodiment of FIG. 1, audio source separation is performed in real time on an audio input signal y(n). The audio input signal y(n) is, for example, a karaoke signal including user vocals and background sounds. The background sounds may be any noise that may be captured by the karaoke singer's microphone, such as crowd noise. The audio input signal y(n) is processed online through a voice separation algorithm to extract and potentially remove the user vocals from the background sounds. An example of real-time voice separation is described in a known paper (Uhlich, Stefan, et al. "Improving music source separation based on deep neural networks through data augmentation and network blending," 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017). In the paper, bidirectional LSTM layers are replaced by unidirectional LSTM layers.
[0045] Audio source separation is performed in real time on an audio input signal x(n). The audio input signal x(n) may be, for example, a karaoke song containing original vocals and musical accompaniment. The audio input signal x(n) may be processed online via a voice separation algorithm to extract and potentially remove the user's voice from the playback sound, or the audio input signal x(n) may be pre-processed, for example, when the audio input signal x(n) is stored in a music library. In the case of pre-processing, pitch analysis and pitch range estimation may also be performed in advance. To pre-process each song in the karaoke song database, it needs to be analyzed for pitch range.
[0046] There are karaoke booths that allow manual transposition. However, most karaoke singers (also called karaoke users) do not know whether the pitch range is suitable for their abilities, and therefore do not know how to transpose the accompaniment. Acc (n) Automatic online transposition has great advantages.
[0047] In one embodiment, the audio input x(n) is a MIDI file (see more detail in the description of FIG. 7 below). In this example, a MIDI synthesizer generates the accompaniment s for each MIDI track. Acc (n) is transposed.
[0048] In another embodiment, the audio input x(n) is an audio recording, e.g., a WAV file, an MP3 file, an AAC file, a WMA file, an AIFF file, etc. This means that the audio input x(n) is real audio, e.g., raw, unprepared audio that is not a commercially performed song. In this embodiment, no pre-prepared audio / MIDI material is required, since karaoke material does not require manual preparation, can be processed fully automatically, is online, and can provide good quality and high realism.
[0049] Karaoke systems use vocal / instrument separation algorithms (see Figure 2) to obtain clean vocal recordings from the karaoke singer's microphone or the original song (sung by the original singer) to analyze the karaoke singer's pitch range and singing effort (see Figure 8).
[0050] Although the pitch analysis and comparison sections are functionally separated in Figure 1, they are performed automatically at both stages and connected to differ from the original recording by the smallest transposition factor while minimizing singer fatigue and effort. Essentially, the system optimizes the performance experience for both the singer and the listener of the karaoke session.
[0051] Another advantage of the karaoke system described above is that the low-latency vocal / instrument separation process enables online pitch analysis and transposition. Furthermore, vocal separation allows for accurate analysis of vocal pitch range and determination of singing effort. Furthermore, the Real Audio vocal / instrument separation process does not limit karaoke to MIDI karaoke songs, thus making the music much more realistic. Furthermore, vocal / instrument separation can improve the transposition quality of Real Audio recordings.
[0052] Audio remix / upmix with audio source separation
[0053] Figure 2 shows a schematic diagram of a general approach to audio upmix / remix by blind source separation (BSS), such as multi-channel separation (MSS). First, audio source separation (also called "demixing") is performed on a source audio signal 1 (here, an audio input signal x(n) containing multiple channels I and audio from multiple audio sources Source 1, Source 2, ..., Source K (e.g., instruments, voices, etc.)), to obtain, for each channel i, a separated Source 2 (e.g., vocals S0(n)) and a residual signal 3 (e.g., accompaniment s A(n)), where K is an integer indicating the number of audio sources. Here, the residual signal is a signal obtained after separating vocals from the audio input signal. That is, the residual signal is a "rest" audio signal after removing the vocals from the input audio signal. In this embodiment, the source audio signal 1 is a stereo signal having two channels i=1 and i=2. Subsequently, the separated source 2 and residual signal 3 are remixed and rendered into a new speaker signal 4, here a signal including five channels 4a-4e, i.e., a 5.0 channel system. The audio source separation process (see 104 in FIG. 1) is described in detail, for example, in a known paper (Uhlich, Stefan, et al., "Improving music source separation based on deep neural networks through data augmentation and network blending," 2017 IEEE International Conference on Acoustics, Speech and Signal IEEE, 2017).
[0054] When the separation of the audio source signals is incomplete, for example, due to a mix of audio sources, a residual signal 3 (r(n)) is generated in addition to the separated audio source signals 2a-2d. This residual signal may represent, for example, the difference between the input audio content and the sum of all separated audio source signals. The audio signal emitted by each audio source is represented in the input audio content 1 by its respective recorded sound wave. For input audio content having two or more audio channels, such as stereo or surround sound input audio content, spatial information for the audio sources is also typically included or represented in the input audio content (e.g., as the proportion of audio source signals included in different audio channels). The separation of the input audio content 1 into the separated audio source signals 2a-2d and the residual signal 3 is performed based on blind source separation or other techniques capable of separating audio sources.
[0055] In a second step, the separated audio source signals 2a-2d and, if a residual exists, the residual signal 3 are remixed and rendered into a new speaker signal 4, here a signal including five channels 4a-4e, i.e., a 5.0 channel system. Based on the separated audio source signals and the residual signals, an output audio content is generated by mixing the separated audio source signals and the residual signals based on spatial information. The output audio content is exemplarily shown in Figure 2 and is designated by reference number 4.
[0056] In a second step, the separation and residual, if any, are remixed and rendered into a new speaker signal 4, here a signal including five channels 4a-4e, i.e., a 5.0 channel system. Based on the separated audio source signals and the residual signals, an output audio content is generated by mixing the separated audio source signals and the residual signals based on spatial information. The output audio content is exemplarily shown in FIG. 2 and is designated by reference numeral 4.
[0057] The audio input x(n) and the audio input y(n) can be separated using the method described in Figure 2. The audio input y(n) is the user vocal s user The audio input x(n) is separated into the original vocal s(n) and unused background sound. user (n) and accompaniment s acc (n) and accompaniment s acc (n) is further divided into respective tracks such as drums, piano, strings, etc. (See Figure 11). Vocal separation allows for significant improvements in the way both accompaniment and vocals are handled.
[0058] Another method for removing the accompaniment from the audio input y(n) is, for example, the crosstalk cancellation method, in which a reference of the accompaniment is subtracted in phase from the microphone signal, for example by using adaptive filtering.
[0059] Another method for separating audio input y(n) is available when the mastering recording for audio input y(n) has detailed knowledge of how audio input y(n) (i.e., the song) was mastered. In this case, the stems need to be remixed without the vocals, and the vocals need to be remixed without any accompaniment. This process uses a much larger number of stems during mastering, e.g., layered vocals, multi-microphone takes, applied effects, etc.
[0060] Pitch Analysis
[0061] Figure 3 shows in more detail an embodiment of the pitch analysis process performed in the pitch analysis unit 13 of Figure 1. As shown in Figure 1, the original vocal s original (n) and user vocals original (n) and perform pitch analysis on them, and the pitch analysis result ω f In particular, the signal framing 301 process is performed on the vocal 300, i.e., on the vocal signal s(n), to obtain the framed vocal s(n). n (i). The Fast Fourier Transform (FFT) Spectral Analysis 302 process obtains the framed vocal S n (i) is performed on the FFT spectrum S ω (n) is obtained. FFT spectrum S ω A pitch measurement analysis 303 is performed on (n) to obtain the pitch measurement R P (ω f ) is obtained.
[0062] In the signal framing 301, the framed vocal s n A windowed frame such as (i) can be obtained by Equation 1.
[0063]
number
[0064] where s(n+i) represents the discretized audio signal shifted by n samples (i represents the sample number and therefore time), and h(i) is a framing function around time n (respectively sample n), such as the Hamming function well known to those skilled in the art.
[0065] In the FFT spectrum analysis 302, each frame of vocal is transformed into its respective short-term power spectrum. The short-term power spectrum S(ω), also known as the power spectral density, obtained by the discrete Fourier transform, can be obtained by Equation 2:
[0066]
number
[0067] where S n (i) is a framed vocal S as defined above n (i) is the signal in the windowed frame, ω is the frequency in the frequency domain, and |S ω (n)| are the components of the short-term power spectrum S(ω), and N is the number of samples in the windowed frame, for example, in each framed vocal.
[0068] The pitch measurement analysis 303 can be performed, for example, as described in the known paper by Der-Jenq Liu and Chin-Teng Lin, "Fundamental frequency estimation based on the joint time frequency analysis of harmonic spectral structure," in IEEE Transactions on Speech and Audio Processing, vol. 9, no. 6, pp. 609-621, September 2001.
[0069] Pitch measurement value R P (ω f ) is the fundamental frequency candidate ω f For each frame window S n The power spectral density S ω (n) can be obtained by equation 3.
[0070]
number
[0071] where R E (ω f ) is the fundamental frequency candidate ω f is the energy measurement of R I (ω f ) is the fundamental frequency candidate ω f is an impulse measurement value.
[0072] Fundamental frequency candidate ω f Energy measurement value R E (ω f ) is obtained by equation 4.
[0073]
number
[0074] Here, K(ω f ) is the fundamental frequency candidate ω f is the number of harmonics of h in (nω f ) is the fundamental frequency candidate ω f Harmonics of lω f is the internal energy associated with , and E is the total energy, where E is given by equation 5.
[0075]
number
[0076] The internal energy is given by equation 6.
[0077]
number
[0078] The internal energy is the length W in is the area under the curve of the spectrum bounded by the inner window of , and the total energy is the sum of the areas under the curve of the spectrum.
[0079] Fundamental frequency candidate ω f Impulse measurement value R I (ω f ) can be obtained using equation 7.
[0080]
number
[0081] where ω f is a candidate fundamental frequency, and K(ω f ) is the fundamental frequency candidate ω f is the number of harmonics of h in (lω f ) is the harmonic nω f is the internal energy of the candidate fundamental frequency associated with h out (lω f ) is the harmonic lω f is the external energy related to
[0082] The external energy is given by equation 8.
[0083]
number
[0084] The external energy is out is the area under the curve of the spectrum enclosed by the outer window of
[0085] Frame Window S n Pitch analysis result ω^ f (n) is obtained from equation 9.
[0086]
number
[0087] Here, ω^ f (n) is the fundamental frequency of the window S(n), and R P (ω f) is the fundamental frequency candidate ω obtained by the pitch measurement value analysis 303 as described above. f is the pitch measurement value.
[0088] Fundamental frequency ω^ at sample n f (n) is a pitch measurement showing the vocal pitch at sample n in vocal signal s(n).
[0089] Furthermore, the pitch measurement result ω^ f A low-pass filter (LP) 304 is applied to (n) to obtain the pitch analysis result ω f (n)Get 305.
[0090] The low-pass filter 305 may be an M-th order causal discrete-time low-pass finite impulse response (FIR) filter, given by:
[0091]
number
[0092] α i is i for 0≦i≦M th is the value of the impulse response at time instant e p In (n), each value in the output sequence is a weighted sum of the most recent input values.
[0093] Filter parameters M and α i can be chosen according to the design choice of one skilled in the art, e.g., α = 1 for normalization. The parameter M can be chosen, for example, on a time scale up to 1 sec.
[0094] The pitch analysis process described above with respect to Figure 3 original (n) is performed on the original vocal pitch analysis result ω f,original (n) is obtained. The pitch analysis process original (n) is executed to obtain the user pitch analysis result ωf,user (n) is obtained.
[0095] In the embodiment of FIG. 3, the fundamental frequency ω f It is proposed to perform a pitch measure analysis, such as pitch measure analysis 303, to estimate the fundamental frequency ω f can be estimated based on a fast adaptive representation (FAR) spectral algorithm.
[0096] Other methods for pitch analysis and estimation of monophonic signals that can be used instead of or in addition to the method described in Figure 3 are described in the following scientific papers: The multiplicative autocorrelation method is described in "New methods of pitch extraction," by Sondhi, M. M, published in EEE Trans. Audio Electroacoust. AU-16, 262-266, in 1968. The mean magnitude difference function method is described in "Average magnitude difference function pitch extractor" by Ross, M. J, Shaffer, H. L., Cohen, A., Freudberg, R., and Manley, H. J, published in IEEE Trans. Acoust. Speech Signal Process. ASSP-22, 353-362, in 1974. The comb filtering method is described in "The optimum comb method of pitch period analysis of continuous digitized speech" by Moorer, J. A., published in IEEE Trans. Acoust. Speech Signal Process. ASSP-22, 330-338, in 1974. A method based on linear predictive analysis is described in "Linear Prediction of Speech" by Moorer, J. A., published in Springer-Verlag, New York, in 1974. A method based on cepstrum analysis is described in "Cepstrum pitch determination" by Noll, A. M., published in J. Acoust. Soc. Am. 41, 293-309, in 1966.The period histogram method is described in "Period histogram and product spectra: New methods for fundamental frequency measurement," by Schroeder, MR, published in J. Acoust. Soc. Am. 43, 829-834, in 1968.
[0097] Furthermore, other more advanced methods for pitch analysis and estimation that can be used instead of or in addition to the method described in Figure 3 are described in the scientific paper "Fundamental frequency estimation of musical signals using a two-way mismatch procedure", by R.C. Maher and J.W. Beauchamp, published in the Journal of the Acoustical Society of America 95(4), in April 1994.
[0098] For robust pitch determination, it is necessary to use pitch tracking (avoiding pitch doubling errors and voiced / unvoiced detection), which is often done by using dynamic programming for pitch F0 candidates, as described in one of the methods given above. Pitch tracking methods are described in "An integrated pitch tracking algorithm for speech systems," B. Secrest and G. Doddington, published in ICASSP '83. IEEE International Conference on Acoustics, Speech, and Signal Processing, Boston, Massachusetts, USA, 1983, pp. 1352-1355, doi: 10.1109 / ICASSP.1983.1172016.
[0099] Furthermore, pitch analysis and (key) transposition are better when vocals and accompaniment are separate.
[0100] Pitch range judgment
[0101] 4 is a flowchart illustrating the process of the pitch range determination unit 15 in FIG. 1. In step 41, the pitch analysis result ω f (n) is received as input to pitch range determination unit 15. In step 42, sample number n is tested to be zero. If the query in step 42 is answered with a yes, the process proceeds to step 43. In step 43, the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] lower limit min_ω f (n) is min_ω f (0)=ω f Initialize with (0). Pitch range R ω (n)=[min_ω f (n),max_ω f (n)] upper limit max_ω f (n) is max_ω f (0)=ω f After step 43, the process continues to step 51. In step 51, the pitch range R ω =[min_ω f (n),max_ω f (n)] is output by the pitch range determination unit 15 and stored in storage, for example, in storage memory 1202. If the query in step 42 is answered with a No, the process proceeds to step 44. In step 44, the old pitch range R ω,old =[min_ω f (n-1),max_ω f (n-1)] is loaded from storage. In step 45, the pitch analysis result ω f (n) is the old pitch range R ω (n)=[min_ω f (n-1),max_ωf (n-1)] lower limit min_ω f If the query in step 45 is answered with a Yes, the process proceeds to step 46. In step 46, min_ω is f (n)=ω f (n) is the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] lower limit min_ω f (n) and proceed to step 50. In step 50, max_ω f (n)=max_ω f (n-1) is the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] upper limit min_ω f (n) and proceed to step 51. In step 51, the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] is output by the pitch range determination unit 15 and stored in storage, for example, in storage memory 1202. If the query in step 45 is answered with a No, the process proceeds to step 47. In step 47, the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] lower limit max_ω f (n) and min_ω f (n)=min_ω f (n-1) is set, and the process proceeds to step 48. In step 48, the pitch analysis result ω f (n) is the old pitch range R ω =[min_ω f (n-1),max_ω f (n-1)] upper limit max_ω f If the query in step 48 is answered with a Yes, the process proceeds to step 49. In step 49, the pitch range R ω =[min_ω f(n),max_ω f (n)] upper limit max_R ω (n) is max_ω f (n)=ω f (n), and the process proceeds to step 51. In step 51, the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] is output by the pitch range determination unit 15 and stored in storage, for example, in storage memory 1202. If the query at step 48 is answered with a No, the process proceeds to step 50. At step 50, the pitch range R ω =[min_ω f (n),max_ω f (n)] upper limit max_ω f (n) is max_ω f (n)=max_ω f (n-1) and proceed to step 51. In step 51, the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] is output by the pitch range determination unit 15 and stored in a storage, for example, the storage memory 1202.
[0102] The above pitch range judgment process is performed based on the original vocals. original Pitch analysis result of (n) ω f,original (n) and user vocals user Pitch analysis result of (n) ω f,user (n) and (n) can be carried out based on the above.
[0103] The pitch determination process of the pitch determiner 15 as described above in FIG. 4 can be performed online, which means that the pitch analysis process 14 and pitch range determination process 15 are performed for each sample (or frame) of the audio input y(n) (e.g., a user's karaoke performance).
[0104] In another embodiment, the pitch range determination process of pitch determiner 15 as described above may be performed on a pre-stored audio input x(n) (e.g., a song whose pitch range is to be determined stored in a karaoke system), where the pitch range R ω (n)=[min_ω f (n),max_ω f (n)] upper limit max_ω f (n) is max_ω f (n)=(n) where max is the maximum function and N is the total number of samples of the stored audio input x(n).
[0105]
number
[0106] Pitch Range R ω (n)=[min_ω f (n),max_ω f (n)] lower limit min_ω f (n) is determined by setting Equation 12, where min is the minimum function.
[0107]
number
[0108] In yet another embodiment, the pitch range determination process of pitch determiner 15 as described above may be performed on pre-stored audio input y(n) (i.e., stored user karaoke performances for a number of existing songs, for which pitch range and singing effort (see below) profiles can be compiled), where pitch range R ω (n)=[min_ω f (n),max_ω f (n)] may be determined as explained in the previous paragraph.
[0109] A graph of the pitch analysis results is shown schematically in Figure 5. On the x-axis of the graph 50, the number of samples n of the audio input ty(n) or tx(n) is shown, with the total number of samples being N.
[0110] On the y-axis of Graph 50, the pitch range analysis result ω f (n) is shown. Graph line 53 shows the pitch range analysis result ω f (n) shows the pitch range R across all N samples. ω (n)=[min_ω f (n),max_ω f (n)] lower limit min_ω f (n) is the minimum value min_ω of all N samples that the graph line 53 reaches f (N). The pitch range R over all N samples ω (n)=[min_ω f (n),max_ω f (n)] upper limit max_ω f (n) is the maximum value max_ω of all N values that the graph line 53 reaches f (N).
[0111] Pitch range comparison
[0112] FIG. 6 is a flow chart illustrating the process of the pitch range comparison unit 16 of FIG. 1. In step 61, the original vocal s original The pitch range R of (n) (also called the first vocal signal) ω,original (n)=[min_ω f (n),max_ω f (n)] (also called the first pitch range) is received and input to step 63. In step 62, the user vocal s user The pitch range R of (n) (also called the second vocal signal) ω,user (n)=[min_ω f (n) , max_ω f(n)] (also called the second pitch range) is received and input to step 64. In step 63, the original vocal pitch range average value avg_ω f,original is avg ωf,original (n)=[max_ω f,original (n)-min_ω f,original (n)] / 2+min_ω f,original In step 64, the user vocal pitch range average value avg_ω is determined as (n). f,user (n) is avg_ω f,user (n)=[max_ω f,user (n)-min_ω f,user (n)] / 2+min_ω f,user In step 65, the pitch ratio P ω (n) to P ω (n)=[(avg ωf,user (n)-avg ωf,original (n)) / avg_ω f,original (n)+1]. In step 66, the pitch ratio P ω (n) is output by the pitch range comparison process of the pitch range comparison unit 16.
[0113] The pitch range comparison process of the pitch range comparator 16 as described above is performed on the user vocal s user (n) for every n samples. In other words, while the user is singing karaoke, the pitch ratio P ω (n) can be adapted to the final pitch ratio P of all samples n=1...N after the user has finished their karaoke performance. ω (N) may be stored in a database, for example, storage 1202, and linked to the user.
[0114] Pitch ratio P ω (n) is the original vocal pitch range average value avg_ω f,original (n), and since it is centered around 1, the original vocal pitch frequency ω f,original It can be seen as a kind of "transposition factor" to be applied to (n).
[0115] As described above, the pitch analysis result ω f (n) and the pitch range R from the pitch range determination unit 15 ω As with (n), the pitch ratio ω (n) can be determined online for each sample n from audio input y(n) (e.g., a user's live karaoke performance) and from audio input x(n) (e.g., a selected song on which the karaoke performance is to be performed).
[0116] User pitch range R ω,user If (N) is known in advance (i.e., before karaoke is performed for a song and audio input y(n) is obtained) (e.g., from another song performed by a user and stored in storage 1202), the pitch ratio P ω (N) is the user's known range R ω,user and the user's pre-known range R ω,original (N) can be judged based on the above.
[0117] In the field of music and musical transposition, it is often described how many semitones or whole steps a piece of music is transposed. An octave consists of 12 semitones, and an octave is a pitch ratio P ω Since (n)=2, transposition up a semitone corresponds to the pitch ratio P ω (n)=2 1 / 12 = 1.087. Transposing down a semitone corresponds to the pitch ratio P ω (n)=(1 / 2) 1 / 12 = 0.920. This corresponds to the pitch ratio P ω (n) and semitone transposition specifications can be easily converted. Therefore, in another embodiment, the pitch ratio P ω (n) is the pitch ratio P ω It may be rounded to the next semitone relative to the ceiling or floor (i.e., above or below) so that (n) always corresponds to a transposition of an integer multiple of a semitone.
[0118] Transposition
[0119] As mentioned above, the goal is to allow the user to accompany their own voice during their karaoke performance. Acc (n) so that the song accompaniment s can be more easily matched Acc (n) is to transpose the accompaniment s Acc The "transposition factor" by which (n) should be transposed is determined as described above in Figure 6. Transposition of the audio input can be performed, for example, by a standard pitch scale modification technique, where all frequencies are multiplied by a predetermined transposition value, transpose_val(n). Standard pitch scale modification techniques include the steps of time scale modification and resampling.
[0120] Figure 7 shows a schematic flow chart illustrating the process of the transposition unit 17 of Figure 1. In step 71, a transposition value transpose_val is received. In this embodiment, the transposition value transpose_val(n) is calculated by multiplying the pitch ratio P ω,user (n), i.e., transpose val(n) =R ω,user (n). In step 72, the accompaniment s Acc (n) is received as input. In step 73, accompaniment s Acc The time scale modification of (n) uses transpose_val(n) as the time coefficient along with the transpose value. Acc The time scale modification of (n) is performed using a phase vocoder. Acc Without changing the frequency of (n), transpose the accompaniment s by a factor of transpose_val. Acc (n) is expanded or shortened, so that the timescale-corrected accompaniment s is obtained as the output of step 73 and as the input to step 74. Acc,mod In step 74, the time scale corrected accompaniment s Acc,mod (n) is resampled with a new sampling period ΔT*transpose_val(n), where ΔT is the time AccThis is the sampling period used when sampling (n). This is the timescale corrected accompaniment s during resampling with the new sampling period ΔT*transpose_val(n). Acc,mod (n) is accompaniment s Acc (n), whereby all frequencies are multiplied by the transpose value transpose_val(n), resulting in the transposed accompaniment s * Acc This means that the transposed accompaniment s (n) is obtained in step 75. * Acc (n) is output by the transposition unit 17.
[0121] In this embodiment, the audio output signal x * (n) is accompaniment Acc (n). In general, other audio output signals x * 7 can be applied to (n). For example, in another embodiment, the audio output signal is x * (n), which may be equal to the audio input signal x(n). The same transposition as described in FIG. 7 is then applied to the audio output signal x * In this example, the output signal of the comparator is the transposed signal s * It is sometimes called (n).
[0122] Time-scale modified phase vocoders and resampling are described in more detail, for example, in the scientific paper "New phase-vocoder techniques for pitch-shifting, harmonizing and other exotic effects," published in Proc. 1999 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, New Paltz, New York, October 17-20, 1999, or in the papers mentioned therein. Still further an improved phase-vocoder is explained in more detail, for example, in the paper "Improved Phase Vocoder Time-Scale Modification of Audio," by Jean Laroche and Mark Dolson, published in IEEE transactions on speech and audio processing, vol. 7, no. 3, May 1999.
[0123] If the transpose value transpose_val(n) is less than 1, steps 73 and 74 in FIG. 7 may be reversed.
[0124] As mentioned above, the pitch ratio P ω (n) can be determined online for each sample n, and the transposed accompaniment s * Acc (n) can be determined online for each n depending on the current transposition value transpose_val(n) (which can also be thought of as the transposition key), and then applied to the entire song in real time.
[0125] The pitch ratio P ωIf (N) (and the transpose value transpose_val(n)) is known in advance, the transposed accompaniment s * Acc (n) may also be determined in advance.
[0126] As described above, the accompaniment s output by the sound source separator 12 (see FIG. 2) Acc (n) can include all instruments (tracks), e.g., drums, piano, strings, etc. In this example, the transposition process of the comparator is performed to obtain the "complete" accompaniment s, as explained in Figure 7. Acc (n) (also known as polyphonic pitch transposition). Polyphonic pitch transposition can be of lower quality than single-track pitch transposition (see Figure 11). This is because it is difficult to deal with very different attack / release, melodic / percussive, and multi-note-on / note-off times on tracks with multiple instruments. This can result in artifacts such as pre-echoes in percussive parts and comb / flange effects in melodic parts.
[0127] As mentioned above, the pitch ratio P ω (n) can also be written in semitones or whole steps, and exactly the same is true for the transpose value transpose_val(n).
[0128] In yet another embodiment, the audio input signal x(n) may be MIDI (Musical Instrument Digital Interface) enabled, and thus the accompaniment s Acc (n) Or a single track of accompaniment may also be available as a MIDI file. In this case, the MIDI file's accompaniment Acc The transposition of (n) can be done with normal MIDI commands like transpose filters, i.e. in this case the transposition is performed by simply transposing the key of the MIDI track by the required transposition value transpose_val(n) before instrument synthesis.
[0129] The above-described comparison unit can therefore process any type of recording (synthetic MIDI, third-party covers, or commercially released recordings) that can have high separation quality and improved transposition quality through pitch analysis and transposition value determination.
[0130] Judging singing effort
[0131] 8 shows a schematic diagram of a second embodiment of a karaoke system process for transposing an audio signal based on source separation and pitch range estimation. An audio input signal x(n) received from a mono or stereo audio input 13 contains multiple sources (see 1, 2, ..., K in FIG. 2) and is input to a source separation 12 process (see separated source 2 and residual signal 3 in FIG. 2), where separated source 2, i.e., original vocal s original (n), and residual signal 3, i.e., accompaniment s Acc An exemplary embodiment of the source separation 2 process is described in Figure 2 below. The audio output signal x*(n) is separated into accompaniment s Acc (n), the audio output signal is x*(n) sent to the transposing unit 17, and the original vocal s original (n) is the signal adder 18 and the original vocal s original Pitch analysis result of (n) ω f,original The pitch analysis result ω is transmitted to the pitch analysis unit 14 (see FIG. 3 for details) which estimates the pitch (n). f,original (n) is the original vocals original (n) pitch range R ω,original The pitch range R is input to a pitch range estimation unit 15 (described in detail in FIG. 4) that estimates the pitch range R ω,original is input to the pitch comparison unit 16. The user microphone 11 acquires the separated audio input signal y(n) input to the sound source separation process 12 (see the separated sound source 2 and residual signal 3 in FIG. 2). user (n) and the unwanted residual signal 3 below. original (n) is a signal summation unit 22, a signal summation unit 18, and a user vocal soriginal Pitch analysis result of (n) ω f,user The pitch analysis result ω is transmitted to the pitch analysis unit 14 (see FIG. 3 for details) which estimates the pitch (n). f,user (n) is user vocals user (n) pitch range R ω,user The pitch range R is input to a pitch range estimation unit 15 (described in detail in FIG. 4) that estimates the pitch range R ω,user is input to the pitch comparator 16. The pitch comparator 16 (described in detail in FIG. 5) compares the original vocal s original (n) pitch range R ω,original and user vocals user (n) pitch range R ω,user Receive the original vocals original (n) pitch range R ω,original The average value of and user vocal s user (n) pitch range R ω,user The average pitch ratio P ω The pitch ratio P ω is input to the transposition value determination unit 23. The singing effort determination unit 22 determines the user vocal s original (n), user vocals original Pitch analysis result of (n) ω f,user (n), and user vocals user (n) pitch range R ω,user and judges the singing effort (see FIG. 9). The singing effort judgment unit 22 outputs the singing effort flag E input to the transposition value judgment unit 23. The transposition value judgment unit 23 judges the singing effort based on the pitch ratio P ω and the singing effort flag E, the transpose value determination unit 23 outputs the transpose value transpose_val to the transposing unit 17. The comparing unit receives the transpose value transpose_val and calculates the audio output signal as x*(n) (=accompaniment s Acc (n)) and the audio output signal x*(n) (=accompaniment s Acc The transposition unit 17 transposes the transposed accompaniment s*(n) by the transpose value transpose_val. Acc The signal adder 18 outputs the transposed accompaniment s*. Acc(n) and original vocals original (n) are input, and the sum is output to the speaker system 19. The transposition value transpose_val is further output to the display unit 20 and presented to the user. The display unit 20 also displays the user vocal s user (n) receives the lyrics and presents them to the user.
[0132] Singing effort and vocal cord pathology
[0133] The karaoke system can further estimate the singing effort of the karaoke singer. Singing effort indicates whether the karaoke user has to exert a lot of effort to reach the pitch range of the original song, i.e., whether the karaoke user has to exert a lot of effort to sing high or low in the original song. If an amateur karaoke user sings beyond their inherent ability for a longer period of time, the user may not be able to endure the long singing session, which may damage their vocal cords and result in poor performance quality.
[0134] User vocals user (n) and / or user pitch analysis result ω f,user There are various characteristic parameters that can be inferred from the analysis of (n), which are indicative of high singing effort. These different characteristic parameters are, for example:
[0135] Jitter value (in percent) of the user pitch analysis result ω in the analyzed voice sample. f,user Relative evaluation of period-to-period (very short-term) fluctuations in (n). Voice break regions are excluded.
[0136] RAP value (in percent). A relative assessment of the period-to-period variation of pitch within the analyzed audio sample, with a smoothing factor of three periods. Voice break regions are excluded.
[0137] Shimmer value (in percent). A relative assessment of period-to-period (very short-term) fluctuations in peak-to-peak amplitude within the analyzed audio sample. Voice break regions are excluded.
[0138] APQ value (in percent). Relative assessment of period-to-period (very short-term) fluctuations in peak-to-peak amplitude within an analyzed audio sample over 11 periods of smoothing. Voice break regions are excluded.
[0139] Noise-to-Harmonic Ratio (NHR) value. The average ratio of the low-frequency spectral energy in the frequency range 1500-4500 Hz to the harmonic spectral energy in the frequency range 70-4500 Hz. This is a general assessment of the noise present in the analyzed signal.
[0140] The soft phonation index (SPI) value, which is the average ratio of low-frequency harmonic energy in the range of 70-1600 Hz to high-frequency harmonic energy in the range of 1600-4500 Hz. This parameter reflects vocal approximation. A high SPI value is said to correlate with incomplete vocal fold adduction and is a better indicator of breathiness than electroglottography (EGG). Both NHR and SPI are calculated using pitch-synchronous frequency-domain methods.
[0141] User vocals user (n) and / or user pitch analysis result ω f,userBased on (n), a more detailed analysis of the above-mentioned parameters and methods for measuring and detecting them can be found in the scientific paper "Vocal Folds Disorder Detection using Pattern Recognition Methods", J. Wang and C. Jo, published in 200729th Annual International Conference of the IEEE Engineering in Medicine and Biology Society, Lyon, 2007, pp. 3253-3256, doi: 10.1109 / IEMBS.2007.4353023.
[0142] Most of the above parameters are related to the vocal cords. Some of these, such as jitter (vibrato), are related to expressiveness during singing, but continuous chaotic vocal cord behavior throughout a karaoke singing session may be an indicator of developing short-term vocal cord problems such as swelling. NHR values can also be used to detect aphasia. A karaoke system can monitor these above-mentioned indicators and their variations over a user's karaoke session to determine singing effort and possible vocal cord damage (e.g., through a gradual deterioration in singing quality).
[0143] 9 is a schematic diagram of the singing effort determination unit 22 of FIG. 8. In step 91, the singing effort determination unit 22 receives a user vocal s user (n) is input. In step 92, the user pitch analysis result ω f,user (n) is input to the singing effort determination unit 22. In step 93, the singing effort determination unit 22 receives the user vocal s user (n) pitch range R ω,user (n)=[min_ω f,user (n) , max_ω f,user In step 94, the jitter value jitter_val is calculated based on the user pitch analysis result ω f,user (n) and user vocals user(n). This is described in more detail in the paper by J. Wang and C. Jo (cited above in the paper cited therein). In step 95, a first singing effort value pitch_high(n) is initialized as pitch_high(n)=0, where setting the first singing effort value pitch_high(n)=1 indicates that the karaoke singer is exerting too much effort or is unable to reach a high pitch. Also in step 95, a second singing effort value pitch_low(n) is initialized as pitch_low(n)=0, where setting the second singing effort value pitch_low(n)=1 indicates that the karaoke singer is exerting too much effort or is unable to reach a low pitch. In step 96, a jitter value jitter_val(n) is tested to see if it is greater than a 5% threshold. In alternative embodiments, the jitter threshold may have a different value. If the query in step 96 is answered "yes," the process proceeds to step 97. In step 97, the user pitch analysis result ω f,user (n) and low pitch range R ω,user The absolute value of the difference between (n) is the user pitch analysis result ω f,user (n) and high pitch range R ω,user (n) is greater than the absolute value of the difference, i.e., │ω f,user (n)-min_ω f,user (n)│>│ω f,user (n)-max_ω f,user (n)| is tested. If the query in step 97 is answered with Yes, the process proceeds to step 98. In step 98, the first singing effort value pitch_high(n) is set to 1, i.e., pitch_high(n)=1, and the process proceeds to step 100. If the query in step 97 is answered with No, the process proceeds to step 99. In step 99, the second singing effort value pitch_low(n) is set to 1, i.e., pitch_low(n)=1, and the process proceeds to step 100. If the query in step 96 is answered with No, the process proceeds to step 100. In step 100, the singing effort determination unit 22 outputs the singing effort value E(n)={pitch_low(n), pitch_high(n)}.
[0144] In the above embodiment, the singing effort value E(n) is a "binarized" value of the jitter value jitter_val(n), i.e., a flag is set when the threshold is exceeded and a flag is not set when the threshold is exceeded. In another embodiment, the singing effort value E(n) may be a quantitative value, e.g., a value directly proportional to the jitter value jitter_val(n).
[0145] In yet another embodiment, any of the other different characteristic parameters described above may be used instead of jitter or to determine the first and second singing effort values, as described in FIG. 9.
[0146] In yet another embodiment, the singing effort value E(n) may be a quantitative value, for example, a value directly proportional to any linear or non-linear combination of the different characteristic parameters mentioned above.
[0147] In another embodiment, the karaoke system can suggest stopping or pausing singing to prevent more serious vocal problems. Further details of methods for recognizing pathological speech, which can also be used to detect high singing effort, are described, for example, in "A system for automatic recognition of pathological speech," by Dibazar, Alireza & Narayanan, Shrikanth, published in Proceedings of the Asilomar Conference on Signals, Systems and Computers, November 2002. In this paper, standard Mel-Frequency Cepstral Coefficients (MFCCs) and pitch features are used for classification of several speech production-related pathologies.
[0148] The singing effort determination unit 22 calculates the singing effort value E and the pitch ratio P ω If the above is determined, the transpose value transpose_val can be determined.
[0149] 10 is a schematic diagram of the transposition value determining unit 23 of FIG. 8. In step 101, the pitch ratio P ω is input to the transposition value determining unit 23. In step 102, the singing effort value E={pitch_low(n), pitch_high(n)} is received as input to the transposition value determining unit 23. In step 103, the pitch ratio P ω is set equal to the transpose value transpose_val(n), where transpose_val(n)=P ω In step 104, it is tested whether the first singing effort value is set to pitch_high=1. If the query in step 104 is answered with a Yes, the process proceeds to step 105. In step 105, the transposition value transpose_val is decreased by 0.05, i.e., transpose_val(n)=transpose_val-0.05, and the process proceeds to step 108. If the query in step 104 is answered with a No, the process proceeds to step 106. In step 106, it is tested whether the second singing effort value is set to pitch_low=1. If the query in step 106 is answered with a Yes, the process proceeds to step 107. In step 107, the transposition value transpose_val(n) is increased by 0.05, i.e., transpose_val(n)=transpose_val(n)+0.05, and the process proceeds to step 108. In step 108, the transposition value transpose_val is output by the transposition value determination unit 23.
[0150] Figure 11 shows a schematic diagram of a third embodiment of a karaoke system process for transposing an audio signal based on source separation and pitch range estimation. The embodiment of Figure 11 is similar to the embodiment of Figure 1. However, in Figure 11, the accompaniment s Acc (n) is a first instrument s, such as a drum, piano, or strings, which is separated by sound source separation 12. A1 (n), second instrument A2 (n) and third instruments A3(n) can be separated into various instruments (tracks). A1 (n), s A2 (n), s A3 Each of (n) is transposed by transposing unit 17 according to the same transposition as described above in FIG. * The transposing unit 17 may be set as (n) for the first instrument s A1 The first instrument s after transposition for the input of (n) * A1 (n), or second instrument A2 (n) input for the second instrument s after transposition * A2 (n), and third instruments A3 (n) input for the transposed third instrument s * A3 (n) is output. The first instrument after transposition is s * A1 (n), the second instrument after transposition A2 * (n), transposed third instrument * A3 (n) is integrated by adders 1101 and 1102 to produce the complete accompaniment s * Acc (n) is received.
[0151] In yet another embodiment, the accompaniment s Acc (n) can be separated into melodic / harmonic and percussion tracks, and the same single-track (single instrument) transpositions as above can be applied. Acc If (n) is separated into two or more tracks (instruments), the transposition process of the transposition unit 17 is applied to each of the separated tracks individually, and the individual transposed tracks are then combined into a stereo recording to produce the complete transposed accompaniment s * Acc (n) is received.
[0152] 12 illustrates a fourth embodiment of a karaoke system process for transposing an audio signal based on source separation and pitch range estimation. The embodiment of FIG. 12 is similar to the embodiment of FIG. 1. However, in FIG. 12, an audio output signal x transposed by a transposition value transpose_val(n) is used. * (n) is equal to the audio input signal x(n), which is the original vocal s original (n) (and accompaniment s acc (n)) is also transposed by the value transpose_val(n) as described above. The output of the comparator, i.e., the transposed signal s * (n) is input to adder 18 and proceeds as described in FIG.
[0153] Figure 13 shows a schematic diagram of a fifth embodiment of the process of a karaoke system for transposing an audio signal based on source separation and pitch range estimation. The embodiment of Figure 13 is most similar to the embodiment of Figure 1. However, in Figure 13, the audio output signal x*(n) transposed by a transposition value transpose_val(n) is the same as the accompaniment s acc Original vocals mixed with (n) original For example, the output signal x*(n) is composed of the original vocal s original (n) has gain G (meaning it is amplified or attenuated) and accompaniment s acc The output of the comparator, i.e. the transposed signal s * (n) is input to adder 18 and proceeds as described in FIG.
[0154] FIG. 14 illustrates an embodiment of an electronic device capable of implementing the above-described pitch range determination and transposition processes. The electronic device 1200 includes a CPU 1201 as a processor. The electronic device 1200 further includes a microphone array 1210, a speaker array 1211, and a convolutional neural network unit 1220 connected to the processor 1201. The processor 1201 may implement, for example, a pitch analysis unit, a pitch range determination unit, a pitch comparison unit, a singing effort determination unit, a transposition determination unit, or a comparison unit that implements the processes described with reference to FIGS. 1, 8, 3, 4, 5, 6, 7, 9, and 10 in more detail. The CNN 1220 may be, for example, an artificial neural network in hardware, e.g., a neural network on a GPU, or any other hardware dedicated to implementing an artificial neural network. The CNN 1220 may implement, for example, source separation 104. The speaker array 1211, such as the speaker system 111 described with reference to FIGS. 1 and 8, consists of one or more speakers distributed throughout a given space and configured to render any type of audio, such as 3D audio. The electronic device 1200 further includes a user interface 1212 connected to the processor 1201. The user interface 1212 functions as a human-machine interface, enabling interaction between an administrator and the electronic system. For example, an administrator can use the user interface 1212 to configure the system. The electronic device 1200 further includes an Ethernet interface 1221, a Bluetooth interface 1204, and a WLAN interface 1205. These units 1204 and 1205 function as input / output interfaces for data communication with external devices. For example, additional speakers, microphones, and video cameras with Ethernet, WLAN, or Bluetooth connections can be connected to the processor 1201 via the interfaces 1221, 1204, and 1205. The electronic device 1200 further comprises a data storage 1202 and a data memory 1203 (here a RAM).Data memory 1203 is arranged to temporarily store or cache data or computer instructions for processing by processor 1201. Data storage 1202 is configured as long-term storage for recording sensor data obtained from microphone array 1210 and provided to or retrieved from CNN 1220, for example. Data storage 1202 can also store audio data representing audio messages that a public announcement system can transmit to people moving within a given space.
[0155] It should be noted that the above description is merely an example configuration, and alternative configurations may be implemented using additional or other sensors, storage, interfaces, etc.
[0156] It should be understood that the above-described embodiments describe methods with example orderings of method steps. However, the particular ordering of method steps is provided for illustrative purposes only and should not be construed as binding.
[0157] It should also be noted that the division of the electronics in Figure 1 into units is done for illustrative purposes only, and the present disclosure is not limited to any particular division of functionality in particular units. For example, at least some of the circuitry may be implemented by respectively programmed processors, field programmable gate arrays (FPGAs), dedicated circuitry, etc.
[0158] All units and entities described in this specification and recited in the accompanying claims may be implemented, for example, as integrated circuit logic on a chip, unless otherwise stated, and the functions provided by such units and entities may be implemented by software, unless otherwise stated.
[0159] To the extent that embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be understood that computer programs providing such software control, and transmission, storage, or other media on which such computer programs are provided, are contemplated as aspects of the present disclosure.
[0160] The present disclosure may have the following configurations.
[0161] (1) Audio source separation separates a first audio input signal (x(n)) into a first vocal signal (s original (n)) and accompaniment (s Acc (n);s A1 (n), s A2 (n), s A3 (n)) and the pitch ratio (P ω a circuit configured to transpose the audio output signal (x*(n)) by a transpose value (transpose_val(n)) based on the audio output signal (x*(n)); The pitch ratio (P ω (n)) is the first vocal signal (s original (n)) first pitch range (R ω,original (n)) and the second vocal signal (s user (n)) second pitch range (R ω,user Based on comparison with (n) electronic equipment. (2) The circuit comprises: The first vocal signal (s original The first pitch analysis result of (n) (ω f,original (n)) based on the first vocal signal (s original (n)) of the first pitch range (R ω,original (n)) is determined, The second vocal signal (s user The second pitch analysis result of (n) (ω f,user (n)) based on the second vocal signal (s user (n)) of the second pitch range (R ω,user(n)) (1) The electronic device described in (1). (3) The circuit comprises: The first vocal signal (s original (n)) based on the first pitch analysis result (ω f,original (n)) is determined, The second vocal signal (s user (n)) based on the second pitch analysis result (ω f,user (n)) An electronic device according to (1) or (2). (4) The accompaniment (s Acc (n);s A1 (n), s A2 (n), s A3 (n)) is the first vocal signal (s original (n)) An electronic device according to any one of (1) to (3). (5) The audio output signal (x*(n)) is Acc (n);s A1 (n), s A2 (n), s A3 (n) An electronic device according to any one of (1) to (4). (6) The audio output signal (x*(n)) is the first audio input signal (x(n)). An electronic device according to any one of (1) to (5). (7) The audio output signal (x*(n)) is Acc (n);s A1 (n), s A2 (n), s A3 (n)) and the first vocal signal (s original (n)) An electronic device according to any one of (1) to (6). (8) The circuit Acc (n);s A1 (n), s A2 (n), s A3 (n)) to multiple instruments (s A1 (n), s A2 (n), s A3 (n) An electronic device according to any one of (1) to (8). (9) The circuitry is further configured to separate the second audio input signal (y(n)) by audio source separation. An electronic device according to any one of (1) to (8). (10) The second audio input signal (y(n)) is user (n)) and the residual signal (9) The electronic device described in (9). (11) The circuitry receives the second vocal signal (s user (n)) and further configured to determine singing effort (E(n)) based on the The transposition value (transpose_val(n)) is a function of the singing effort (E(n)) and the pitch ratio (P ω Based on (n) An electronic device according to any one of (1) to (10). (12) The singing effort (E(n)) is calculated by multiplying the second vocal signal (s user The second pitch analysis result (ω f,user (n)) and the second vocal signal (s user (n)) of the second pitch range (R ω,user (n)) and based on (11) The electronic device described in (11). (13) The circuitry is further configured to determine the singing effort (E(n)) based on a jitter value (jitter_val), a RAP value, a shimmer value, an APQ value, a noise-to-harmonic ratio, and / or a soft voicing index. The electronic device according to (11) or (12). (14) The circuit adjusts the pitch ratio (P) so that the transpose value (transpose_val(n)) corresponds to an integer multiple of a semitone. ω configured to transpose the audio output signal (x*(n)) based on An electronic device according to any one of (1) to (13). (15) The circuitry receives the second vocal signal (s user (n) An electronic device according to any one of (1) to (14). (16) The circuit is configured to capture the first audio input signal (x(n)) from a real audio recording. An electronic device according to any one of (1) to (15). (17) The first audio input signal (x(n)) is converted into the first vocal signal (s original (n)) and accompaniment (s Acc (n);s A1 (n), s A2 (n), s A3 (n)) and Pitch ratio (P ω transpose the audio output signal (x*(n)) by a transpose value (transpose_val(n)) based on the The pitch ratio (P ω (n)) is the first vocal signal (s original (n)) first pitch range (R ω,original (n)) and the second vocal signal (s user (n)) second pitch range (R ω,user Based on comparison with (n) method. (18) A computer program comprising instructions that, when executed on a processor, cause the processor to perform the method of (17). Computer program.
Claims
1. a circuit configured to perform audio source separation to separate a first audio input signal into a first vocal signal and an accompaniment signal, and to transpose the audio output signal by a transposition value based on the pitch ratio; the pitch ratio is based on a comparison of a first pitch range of the first vocal signal and a second pitch range of the second vocal signal; the circuitry is further configured to determine singing effort based on the second vocal signal; The transposition value is based on the singing effort and the pitch ratio. electronic equipment.
2. The circuit comprises: determining the first pitch range of the first vocal signal based on a first pitch analysis of the first vocal signal; and determining the second pitch range of the second vocal signal based on a second pitch analysis of the second vocal signal. The electronic device according to claim 1 .
3. The circuit comprises: determining a first pitch analysis result based on the first vocal signal; and determining a second pitch analysis result based on the second vocal signal. Consists of The electronic device according to claim 1 .
4. The accompaniment includes all parts of the first audio input signal except the first vocal signal. The electronic device according to claim 1 .
5. The audio output signal is the accompaniment The electronic device according to claim 1 .
6. The audio output signal is the first audio input signal. The electronic device according to claim 1 .
7. The audio output signal is a mix of the accompaniment and the first vocal signal. The electronic device according to claim 1 .
8. The circuitry is further configured to separate the accompaniment into multiple instruments. The electronic device according to claim 1 .
9. The circuitry is further configured to separate the second audio input signal by audio source separation. The electronic device according to claim 1 .
10. The second audio input signal is separated into the second vocal signal and a residual signal.
10. The electronic device according to claim 9.
11. The singing effort is based on the second pitch analysis of the second vocal signal and the second pitch range of the second vocal signal. The electronic device according to claim 3 .
12. The circuitry may be configured to measure jitter values and / or RAP values and / or shimmer values and / or Based on the APQ value and / or the noise to harmonic ratio and / or the soft voicing index, Further configured to determine singing effort. The electronic device according to claim 1 .
13. The circuitry is configured to transpose the audio output signal based on a pitch ratio such that the transposition values correspond to integer multiples of semitones. The electronic device according to claim 1 .
14. a microphone configured to capture the second vocal signal. The electronic device according to claim 1 .
15. The circuitry is configured to capture the first audio input signal from a real audio recording. The electronic device according to claim 1 .
16. Separating a first audio input signal into a first vocal signal and an accompaniment; transposing the audio output signal by a transposition value based on the pitch ratio; the pitch ratio is based on a comparison of a first pitch range of the first vocal signal and a second pitch range of the second vocal signal; determining singing effort based on the second vocal signal; The transposition value is based on the singing effort and the pitch ratio. method.
17. 17. A computer program comprising instructions which, when executed on a processor, cause the processor to perform the method of claim 16. Computer program.
Citation Information
Patent Citations
Karaoke device
JP1995072881A
Karaoke device
JP1995199978A
Karaoke sing-along machine
JP1997044174A
Method for pitch change of prerecorded background music and karaoke system
JP1997097091A
Karaoke machine with automatically adjusted accompaniment key
JP2004317934A