Voice modification method with visual and audio feedback
The method addresses the lack of real-time audio and visual feedback in vocal training by providing high temporal resolution and low latency feedback, enabling singers to correct intonation and rhythmic errors during singing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-04-02
AI Technical Summary
Existing methods for vocal performance training lack real-time audio and visual feedback, preventing singers from correcting intonation and rhythmic errors during singing, and fail to visualize elements like grace notes and glissando.
A method that provides real-time vocal training through high temporal resolution and low latency audio and visual feedback, allowing singers to see and hear their performance deviations instantly, with a latency of no more than 60 ms for audio and 100 ms for visual feedback, by using complex-modulated filters and instantaneous FFT for harmonic analysis and synthesis.
Enables singers to correct intonation and rhythmic errors in real-time, visualize performance elements like grace notes and glissando, and develop professional vocal techniques by seeing and hearing their performance deviations on a monitor screen.
Smart Images

Figure IMGF000006_0001 
Figure IMGF000009_0001 
Figure IMGF000009_0002
Abstract
Description
[0001] A METHOD OF VOICE MODIFICATION WITH VISUAL AND AUDIO FEEDBACK
[0002] Field of technology
[0003] 5 The invention relates to computing technology, multimedia systems and can be used for professional development of vocal performance techniques, correct intonation and rhythmic presentation of musical compositions in karaoke devices, as well as for developing performance abilities in singing training.
[0004] 10 Prior art
[0005] Methods for assessing the quality of vocal performance of a musical work are known, based on a computer system for teaching vocal performance of a musical work based on the visualization of the fundamental tone frequency contour on the screen (RU, No. 2356105), (US, No. 7271329).
[0006] Known methods include converting musical notation of a musical work into a first graphic image in time-pitch coordinate axes and visualizing it on a screen, a user's vocal performance of said musical work, an audio recording of the musical work, and digital conversion of the recording. Using the converted digital recording, during each time interval corresponding to the sound of a single note of the musical work, multiple calculations are performed to determine the fundamental frequency of the user's vocal performance of the musical work. Using these values, a second graphic image of the user's vocal performance in time-pitch coordinate axes is constructed.The first and second 25 graphic images are compared and the places of discrepancy between the second graphic image and the first graphic image are identified, according to which the quality of the vocal performance of the musical work by the user is assessed, at least for the time corresponding to the sound of each individual note of the musical work.
[0007] 30 However, in known methods, the melody is represented by notes (discrete values with gaps), while the voice, as is known, can continuously move from note to note, and the singer performs the notes themselves with individual intonation. These methods lack audio feedback when the user's voice is in real-time.
[0008] The signal is synthesized with corrected intonation during signal processing (“on the fly”), which does not allow the user to correct their singing by ear.
[0009] A method is known, implemented on the basis of a karaoke device, which allows a singer-performer to understand, based on the sound reproduction of his corrected singing, how he needs to change his singing style (US, No. 8027631).
[0010] In this method, the target performance model's voice data is compared with the input singer's voice data along the time axis. The singer's pitch is then adjusted to match the pitch of the corresponding frame of the target performance model's voice data. The singer's voice data is then temporally scaled so that the length of the singer's voice data segment matches the corresponding length of the target performance model's voice data segment. The corrected singer's voice data is then regenerated, and a corresponding corrected audio signal is generated from the loudspeaker.
[0011] This method allows for the creation of a detailed graphic image of the target performance of a musical work throughout the duration of its vocal performance and the formation and recording of a corrected audio signal of the singer-performer, which allows the singer-performer to perceive his singing by ear and make certain changes to his manner of performing the musical work.
[0012] However, this method does not allow for dynamic analysis of a performed musical work or for the performer to observe the successes or failures achieved during the process of working on the musical work. This is due to the fact that the correction (increase or decrease) of the frame of a certain length of the singer-performer's vocal data is carried out in a number of sections in accordance with the separator information stored in memory, indicating predetermined sections of the vocal data of the target performance model in the direction of the time axis. Therefore, there is no acoustic feedback effect, and the singer-performer can only become aware of his singing with a delay. Since this method does not evaluate the fundamental frequency contour (FFC) and does not visualize it, it is impossible to correct rhythmic errors, loss of intonation due to lack of air, etc. during singing.
[0013] Thus, the existing technical solution essentially only displays the melody, but lacks visualization of elements that characterize performance style. Specifically, it does not provide the ability to practice vocal strokes such as grace notes, glissando, and vibrato.
[0014] 5 The closest to the proposed method of voice modification with visual and audio feedback is the method of voice modification implemented on the basis of karaoke (RU, No. 2591640).
[0015] The known method enables two operating modes: correction of the singer-performer's input voice based on notes and correction of the singer-performer's voice based on a reference performance, which enables the karaoke singer to accurately perform a given melody, as well as correction of the karaoke singer's voice based on a reference performance of a song and melody, allowing the singer-performer to imitate the singing skill of a professional singer. A generalized functional diagram of the voice modification device implementing the known method 15 according to the first or second embodiment is shown in Figure 1.
[0016] The central processor (CPU) performs overall synchronization of the device's operation, while the singer's voice (microphone signal), its modification, and output to the speaker are performed by the audio processor (AP). A device implementing this known method can be implemented on modern computing platforms, such as personal computers and mobile computing systems such as smartphones. These computing platforms typically have multiple processing cores, one of which functions as the audio processor.
[0017] The device's parameter table contains several pre-prepared sets of parameters for storing a song—a musical composition (melody and lyrics). The central processor selects one of the desired parameter sets from the parameter table and configures the audio processor with this selected set of parameters. The output audio signal, generated by the audio processor in accordance with the selected set of parameters and representing a vocal output signal close to the target singer, is fed through the audio output device to the loudspeaker. As in most known approaches, harmonic modeling is chosen for manipulating the fundamental frequency, which involves decomposing the signal into periodic components of several
[0018] frequencies and their parameter extraction [J. Bonada and X. Serra, "Synthesis of the singing voice by performance sampling and spectral models," IEEE Signal Processing Magazine, vol. 24, issue 2, pp. 67–79, March 2007; J. Laroche, Y. Stylianou, and E. Moulines, “HNS: Speech modification based on a harmonic+noise model,” in IEEE ICASSP 1993 – IEEE International Conference on Acoustic, Speech, and Signal Processing, April 27–30, Minneapolis, USA, Proceedings, 1993. – pp. 550–553.]. Speech separation into deterministic and stochastic components is required to reduce audible artifacts.
[0019] The selection of a set of parameters characterizing a musical composition (melody and lyrics) is configured on the control panel within the control tool and displayed on screen 10. Karaoke accompaniment is generated based on the performance track data provided sequentially in time, and parameter sets are selected sequentially based on the control track data provided sequentially in time, synchronous with the performance data: the lyrics are displayed on the monitor screen. Performance and control track data are generated by the central processor.
[0020] 15 A well-known method of modifying the voice of a singer singing a karaoke song is carried out according to the following steps:
[0021] Step 1. Initialize the voice modification device.
[0022] Step 2. Selecting voice modification parameters f T
[0023] 0 ( n ) from the parameter table using the control tool.
[0024] 20 Step 3. In parallel with the karaoke accompaniment, the singer's voice is input through the audio signal input device from the microphone input into the audio processor.
[0025] Step 4. A parametric analysis (mathematical model of the signal: harmonics plus noise) of this signal frame is performed in the audio processor to obtain the parameter vector: instantaneous amplitudes AS(n), fundamental frequency (FPF)f0 S ( n ) ,25 instantaneous values of the phases ^S(n ) and the noise component of the signal r S(n ).
[0026] Step 5. Formation of the output circuit of the frequency response f0̂( n ) in accordance with the target melody in the dynamic parameters generation tool.
[0027] Step 6. Transformation of the signal frame parameters in the dynamic parameters generator to obtain the vector of output parameters [Aˆ( n ) ,f0̂( n ) ,30 ^(n ) , r^ ( n ) ] based on the parameters of the singer-performer of instantaneous amplitudes AS(n ) , з^phase values ^S(n ) and noise component r S(n ) . 5
[0028] Step 7. Parametric synthesis is performed in accordance with these parameters in the audio processor, according to which a modified frame of the singer-performer's output voice signal is formed.
[0029] Step 8. Next, in the audio output device, the output frame of the singer's voice signal 5 is mixed with the musical accompaniment transmitted to the audio output device by the central processor from the parameter table, and is output to the loudspeaker.
[0030] Step 9. If the musical composition is not finished, the process is repeated by entering a new frame of the audio signal of the singer-performer's input voice from microphone input 10 (go to step 3).
[0031] It can be noted that the work is carried out in real time and the central processor synchronizes the parallel work of the audio processor, the audio signal input device and the audio signal output device according to the principle of frame signal processing (Vanhoof, J., Rompaey, K., Bolsens, I., 15 Goossens, G., Man, H.: High-Level Synthesis for Real-Time Digital Signal Processing. Springer US, Boston, MA (1993)).
[0032] It should be noted that in the known method, the process of forming the output contour of the fundamental tone frequency is not reflected on the monitor screen, but an automatic correction of the input voice of the singer-performer is performed according to the notes or 20 correction of the voice of the singer-performer according to the reference performance.
[0033] Analysis of this method of voice modification shows that:
[0034] - This method doesn't allow the singer to learn professional vocal technique, correct intonation, and rhythmic presentation of musical compositions, nor does it allow the singer-performer to immediately see deviations from correct intonation and their own rhythmic errors. They can't correct these errors directly while singing, which prevents the singer-performer from making corrections to their singing by ear;
[0035] - insufficient temporal resolution in assessing the instantaneous frequency of the fundamental tone causes smoothing of the phonation frequency contour, which does not allow one to see in real time 30 phonation modes, fragments of wheezing, loss of intonation due to lack of air, etc. 6
[0036] Brief description of the invention
[0037] The problem solved by the invention is to provide real-time training in the professional development of vocal performance techniques and rhythmic presentation of musical compositions.
[0038] 5 The technical result obtained by implementing the voice modification method is an increase in the temporal resolution in assessing the instantaneous frequency of the fundamental tone, a decrease in the latency of processing and modification of the singer-performer's voice.
[0039] In computer and network technologies, latency is a delay or expectation that increases the actual response time compared to the expected one. The proposed technical solution achieves an audio signal latency of no more than 60 ms, which, taking into account the time it takes for a person to hear the transmitted sound, allows the singer / performer to avoid discomfort and singing disturbances caused by echo, independently correct their performance, and achieve correct intonation. High temporal resolution and the display of the singer / performer's fundamental frequency contour with a delay of no more than 100 ms, without smoothing, provide the singer / performer with more information about the sound of their voice.
[0040] To solve the stated problem with achieving the specified technical result, as in the known method-analogue of voice modification (RU, No. 2591640), in parallel and synchronously with the karaoke accompaniment, the singer-performer's voice is input through the audio signal input device 20 from the microphone input into the audio processor to perform parametric analysis (mathematical model of the signal: harmonics plus noise) of this signal frame and obtain a vector of parameters: instantaneous amplitudes
[0041] S S A S (n), the fundamental frequency (FPF) f0(n), instantaneous phase values ^ (n) and the noise component of the signal r S(n). Next, the output circuit FPF25 f (n) is obtained in accordance with the target melody f T (n), on the basis of which and the parameters^ 0 0
[0042] singer-performer – instantaneous amplitudes AS(n), phase values ^S(n) and noise component r S(n) – a vector of output parameters is formed [ ^A( n) , ^f 0 ( n) ,
[0043]
[0044] ^ ) r^ ( n ) ]. In accordance with these parameters, parametric synthesis is performed in the audio processor, according to which a frame of the output signal of the singer-performer's voice is formed. Then, in the audio signal output device, the frame of the output signal of the singer-performer's voice is mixed with the musical accompaniment, 7
[0045] transmitted to the audio signal output device by the central processor from the parameter table of the support device, and is output to the loudspeaker.
[0046] Unlike its closest analogue, the claimed method of voice modification with visual and audio feedback introduces a temporal scaling of the frame 5 of the microphone signal, matched with the instantaneous FFT of this frame, a complex-modulated bank of analysis and synthesis filters, respectively, for isolating the harmonic components of the signal and interpolating the modified harmonic parameters to form the output signal, inverse temporal scaling of the output signal in accordance with the output FFT circuit f0̂( n ) and10 summation with the stochastic component to obtain a modified signal of the singer-performer's voice and transmission to the loudspeaker, as well as the output in real time of the melody notes f T on the monitor screen
[0047] 0 ( n ) , instantaneous PFC contour f S
[0048] 0( n ) of the singer-performer's voice and the input circuit f0̂( n ) FET.
[0049] There may be additional embodiments of the method, in which it is advisable that:
[0050] - the processing time from the moment the singer-performer's voice signal was input into the microphone of the audio signal input device until it was output to the loudspeaker was no more than 60 ms, and until the notes of the target melody, the contour of the instantaneous frequency of the fundamental tone of the singer-performer's voice, and the output contour of the fundamental tone frequency 20 were output to the monitor screen was no more than 100 ms;
[0051] - synchronously with the display of the notes of the target melody, the contour of the instantaneous frequency of the fundamental tone of the singer-performer's voice and the output contour of the fundamental tone frequency, a line of the text of the performed vocal piece was displayed on the screen;
[0052] 25 - to determine the periodicity of the singer-performer's voice signal, the deterministic part of the model of which is the sum of periodic components with non-stationary parameters - amplitude, frequency and phase, resampling of the input frame of the singer-performer's voice signal is performed for each candidate period of the periodicity, the energy of each resampled input 30 frame is normalized to 1, an estimate of the instantaneous parameters of the sinusoidal model is made, the period candidate formation function (PCF) of the fundamental frequency is calculated for the corresponding set of parameters - amplitude, frequency and phase, multiplying 8
[0053] the calculated value of the function for forming candidates of the fundamental tone period on the weighting function for limiting low-frequency candidates of the fundamental tone, the best continuous contour of the fundamental tone frequency is determined using dynamic programming that maximizes the sum of the FFCP on a local sequence of 5 frames, as a result of which the best candidate of the fundamental tone is selected, which is the initial estimate of the fundamental tone frequency, then a refined estimate of the fundamental tone is calculated using the instantaneous parameters of the sinusoidal model obtained for the best candidate.
[0054] The advantage of the proposed method of voice modification with visual and audio feedback is that it ensures the most accurate matching of the singer-performer's voice to a given melody by incorporating the training process into the cycle of setting up the voice modification method using the principles of visual / audio (biological) feedback for training singing skills. The singer-performer hears the target note generated by the audio processor as a result of the modification of his voice, i.e. his voice, and also sees on the screen the contours of the melody, the fundamental tone of his voice and the fundamental tone of the modified voice of the singer-performer (his altered voice), i.e. he sees and hears how accurately he "hit" the melody.Since the process of processing and synthesizing the modified voice occurs in real time, the singer-performer can adjust their performance during the performance, taking into account the time it takes for a person to hear sound levels. High temporal resolution and the display of the singer-performer's fundamental frequency contour without smoothing provide the singer-performer with more information about the sound of their voice. Phonation modes, fragments of wheezing, and loss of intonation due to lack of air are visible in real time. The singer-performer sees the fundamental frequency contour of their voice on the monitor screen for the required time, which provides ample opportunities for intonation analysis and error correction directly during the singing process. The ability to visualize in real time elements characterizing the performance style allows for the development of singing touches such as grace notes, glissando, and vibrato.
[0055] The said advantages, as well as the features of the present invention 30, are explained using a variant of its implementation with reference to the figures. 9
[0056] Brief list of drawings
[0057] Fig. 1 depicts a generalized functional diagram of a voice modification device implementing the closest analogue method (prior art);
[0058] Fig. 2 depicts a generalized functional diagram of a voice modification device 5 that implements the claimed voice modification method with visual and audio feedback;
[0059] Fig. 3 shows the notes of the melody f T observed by the singer-performer on the monitor screen
[0060] 0 ( n ) , the contour of the instantaneous frequency of the fundamental tone of one's own voice f0 S ( n ) and the output circuit of the fundamental frequency f0̂( n );
[0061] 10 Fig. 4 – illustration of calculation of new moments of time for signal scaling;
[0062] Fig. 5 illustrates the calculation of signal values at arbitrary points in time;
[0063] Fig.6 illustrates the interpolation of a signal by sinc functions with an offset from 0 to 15 1;
[0064] Fig. 7 illustrates spectrograms of the input signal for (a) Fourier analysis and (b) harmonic analysis (FOT dependent);
[0065] Fig. 8 illustrates the scaling of the analysis filter bank for each candidate PFC, where (a) is the candidate period formation function (CPFF), (b)20 is the amplitude spectrum and filter bank for the candidate ^
[0066]
[0067] (c) – amplitude spectrum of the filter bank for the candidate
[0068]
[0069] (d) – original signal;
[0070] Fig.9 illustrates a multi-speed scheme for calculating the fundamental tone FFCP (M is the number of candidates for the fundamental tone period);
[0071] Fig.10 – illustration of the formation of candidates for the period of the partial pressure flux, where (a) is the original25 signal, (b) is the partial pressure flux ^(n, l ), (c) is the partial pressure flux obtained on the basis of the sinusoidal model ^inst(n, l ), (d) is the proposed partial pressure flux ^ms(n, l ) (V ^ 1 );
[0072] Fig. 11 – illustration of the evaluation of the temporal resolution of pitch detection algorithms.
[0073] 30 Detailed disclosure of the invention
[0074] The declared new device (Fig. 2) differs from the previously known one (Fig. 1) in that the known technical solution includes a connection between the audio processor output and the input 10
[0075] central processor for transmitting the f0 frequency control circuit S( n ) the current processed frame of the microphone signal s(n) of the singer-performer to the central processor, as well as the connection of the output of the support means (parameter table) and the output of the output circuit generation unit of the dynamic parameters generation means with the inputs 5 of the central processor, which allow the melody circuit f T to be transmitted to the central processor, respectively
[0076] 0 ( n ) and the output contour of the partial frequency response f0̂( n ) of the processed voice of the singer-performer. The output of the central processor is connected to the input of the control means and the monitor, which allows the central processor to form an image of the corresponding contours of the partial frequency response frame based on the frame of the processed signal 10 of the microphone s(n) of the singer-performer synchronously with the text of the performed work and display it on the monitor screen. This allows the singer-performer to see the contours of the melody, the fundamental tone of his voice and the fundamental tone of the modified voice of the singer-performer (his altered voice) on the monitor screen, and the output of the modified voice of the singer-performer to the loudspeaker makes it possible to implement the principles of visual / audio (biological) feedback for training singing skills: the singer-performer hears in the earphone the target note formed by the audio processor as a result of the modification of his voice, i.e. his voice, and also sees and hears how accurately he “hits” the melody.
[0077] To address the stated objective—providing real-time training in professional vocal performance technique and rhythmic composition—the present invention utilizes the principle of interactive feedback, both visual and aural. The singer's voice is instantly extracted (during real-time signal processing) and graphically superimposed on the melody of the composition (25) in sync with the playing accompaniment (visual feedback). The singer's voice is also synthesized with the corrected intonation and output to the loudspeaker (audio feedback).The delay time or latency (the delay of the microphone signal in the audio processor - the time of processing and output to the loudspeaker of the synthesized signal), with which the singing intonation is displayed on the screen and the synthesized voice sounds in the loudspeaker, must be small (no more than 60 ms) so that the information displayed on the monitor screen (control panel) and the synthesized signal in the loudspeaker are related to the current moment 11.
[0078] (the current state of the singer – the fragment of the musical composition he is performing). Thus, taking into account the time of human auditory perception of sound levels, the singer-performer sees and hears his deviations from the correct intonation, rhythmic errors, and can correct his performance during the performance of the piece (Fig. 3).
[0079] Research shows that latency should not exceed 60 ms for sound and 100 ms for visual displays. For example, the dynamic ranges of individual musical and speech signals, measured using devices whose readings correspond to the auditory perception of loudness (integration time of 60 ms), average 60-70 dB for a symphony orchestra, 35 dB for pop music, 20 dB for a jazz orchestra, 47 dB for a choir, 35 dB for vocal soloists, and 25 dB for a speaker's speech.
[0080] To achieve the stated technical result, it is necessary to improve harmonic analysis / synthesis methods, as well as perform real-time processing: instantaneous harmonic parameterization, modification, and synthesis of the singer-performer's voice parameters. The device (Fig. 2) for implementing the claimed method operates on frames of variable length, which depends on the value of the instantaneous frequency response with an offset of approximately 5 ms. The analysis / synthesis frame length is proportional to an integer number of periods of the frequency response. f0 S ( n ) The parameters of the basic harmonic model 20 are updated approximately every 5 ms. The sampling frequency is 44.1 kHz.
[0081] The parameters of the singer-performer input signal are determined in the following order:
[0082] 1) Temporal scaling is applied to the input signal, which involves adaptive resampling of the signal with a variable sampling frequency of 25, proportional to the extracted instantaneous PF. The signal is sampled in such a way that each PF period contains an equal number of samples.
[0083]
[0084] . Each input sample of the speech signal s ( n ) is assigned a value of n
[0085] phase correspondence to the fundamental tone period
[0086]
[0087] ^ f 0 ( i ) , where f 0
[0088] i ^ 0( i ) – normalized circular frequency of the fundamental tone at time i. New time instants m, in 30 which require recalculation of the input signal are defined as m^^^ 1
[0089]
[0090] ( p / Nf 0 ) , where p is the sample index in the scaled time domain. Inverse function 12
[0091] ^ ^ 1( ^ ) is calculated by linear interpolation of its known values for integers n, as shown in Fig.4.
[0092] Calculations of the signal values at a given point in time are performed according to Kotelnikov's theorem (see Fig. 5, where the dotted line is a continuous signal; the square 5 marker is a signal sampled with an equal time interval; the round marker is the signal values at points in time that are not multiples of the sampling interval), according to which
[0093] ^
[0094]
[0095] where sinc (( t^ nT ) / T ), T is the sampling interval, s ( nT ) are the signal values at discrete10 moments of time nT . For simplicity, we will further assume T ^ 1 , and s ( nT ) will be denoted as s ( n ) .
[0096] In practice, due to summation over infinite limits, a finite number of signal points are selected N pt preceding the moment and N pt subsequent points:
[0097]
[0098] 15 where w(•) is a window function with the center of symmetry at point 0.
[0099] For each output sample, the time stamps are recalculated so that the current instant always falls within the range from 0 to 1, as shown in Fig. 6. This allows the use of a table with pre-calculated sinc function values. Using table values allows for a significant reduction in computational costs.
[0100] Temporal scaling stabilizes the PF and reduces smoothing in the frequency domain caused by PF modulations (Fig. 7 shows the spectrograms of the input signal for Fourier analysis and for harmonic analysis (PF dependent));
[0101] 2) Harmonic parameters are extracted from the scaled signal using a complex-modulated filter bank. Since the scaled signal has a stable sampling interval, the filter bank analysis window always contains a fixed number of PF periods. Using estimates of the amplitude, instantaneous frequency, and phase of each harmonic, the spectral envelope is estimated;
[0102] 3) Next, using the instantaneous values of the frequency response, each signal in the filter bank subband is classified as periodic or stochastic. The solutions for each harmonic are combined into voiced / unvoiced spectral regions;
[0103] 4) the voice part of the signal is synthesized and subtracted from the input signal to obtain the signal residue (stochastic component r S(n)), which does not change when the pitch changes (PFT).
[0104] Thus, the vector of parameters of the singer-performer is determined: instantaneous amplitudes A S(n ) , fundamental frequency (FPF) f S
[0105] 0( n ), instantaneous phase values ^S(n ) and the noise component of the signal r S(n ).
[0106] 10 Modification of harmonic parameters and generation of the output signal are performed in the following order:
[0107] 1) the harmonic parameters determined at the analysis stage are changed in accordance with the target PFC f T
[0108] 0 ( n ) (the contour of the melody notes), while maintaining the harmonic spectrum envelope shape of the input signal. New harmonic phases are generated continuously, approaching the original relative phase shift of the PF harmonic. Due to processing in the warped time domain, where harmonic frequencies are fixed, it is possible to implement an effective synthesis scheme with anti-aliasing filtering based on a complex-modulated synthesis filter bank, in which each channel corresponds to a harmonic component.
[0109] 20 Decimated subchannel sequences are generated using modified harmonic parameters and interpolated by a filter bank, resulting in an output signal with a stable frequency response f0̂( n );
[0110] 2) to obtain the final output signal, the output signal is inversely time-scaled in accordance with the output contour of the partial differential equation f0̂( n ) and25 summed with the stochastic component r S(n ) to obtain a modified signal of the singer's voice, and then transmitted to the loudspeaker.
[0111] When generating the output tone contour f0̂( n ) based on notes from the parameter table, the melody notes of the selected musical piece are read. The output tone contour f0̂( n ) is generated based on the melody notes, introducing minimal distortion into the processed signal. Octave 14 is selected first.
[0112] melody that most closely resembles the user's voice. To do this, the melody's frequency contour is multiplied and divided by factors of 2 and 4, and then compared with the frequency response of the singer's input voice signal f0 S ( n ) . After this, the contour of the input signal of the singer's voice f0 is equalized S ( n ) and the melody in time 5 by using time scaling based on dynamic programming. This procedure reduces the level of audible artifacts introduced during the transitions of the melody from note to note. Then, the FOT contour of the input signal of the singer's voice f0 S ( n ) is attracted to the melody. The initial shape of the FOT contour of the input signal of the singer's voice f0 S ( n ) is retained at the boundaries of 10 voiced segments in order to reduce the "computer accent" effect.
[0113] Thus, the algorithm of the voice modification method with visual and audio feedback, corresponding to the device (Fig. 2), is performed in the following sequence of steps:
[0114] Step 1. Initialize the voice modification device.
[0115] 15 Step 2. Selecting voice modification parameters from the support tool parameters table.
[0116] Step 3. In parallel and synchronously with the accompaniment of the target melody, the singer's voice is input from the microphone input into the audio processor through the audio signal input device.
[0117] 20 Step 4. Determining the contour of the partial pressure flux f0 S ( n ) singer-performer in the audio processor. In ближайшем аналогеf0 S ( n ) determined at the parametric analysis step, but according to the stated method, a dependent NFC analysis is performed, therefore, the NFC is first found, time scaling is performed, and only then do they move on to parametric analysis.
[0118] 25 Step 5. Temporal scaling of the input frame of the singer's voice in the audio processor.
[0119] Step 6. Performing a dependent parametric analysis in the audio processor based on a complex-modulated filter bank of the given scaled signal frame to obtain a parameter vector: instantaneous amplitudes A S (n ) ,30 instantaneous values of phases ^S(n ) and the noise component of the signal r S(n ) . 15
[0120] Step 7. Formation of the output circuit of the FOT f0̂( n ) in accordance with the target melody f T
[0121] 0 ( n ) in the dynamic parameters generation tool.
[0122] Step 8. Transformation of the signal frame parameters in the dynamic parameters generator to obtain the vector of output parameters [Aˆ( n ) ,f0̂( n ) ,5 ^(n ) , r^ ( n ) ] based on the parameters of the singer-performer: instantaneous amplitudes AS(n ) , з ^phase values^
[0123]
[0124] Step 9. Performing parametric synthesis of the modified harmonic component in accordance with the parameters [Aˆ( n ) ,f0̂( n ) , ^ ^(n ) ] in the audio processor.
[0125] Step 10. Inverse time scaling of the modified harmonic 10 component based on the complex-modulated synthesis filter bank.
[0126] Step 11. Formation of a modified frame of the singer-performer's output voice signal as the sum of the harmonic component and the noise component r S(n ) .
[0127] Step 12. Next, in the audio signal output device, the output frame of the singer's voice signal 15 is mixed with the musical accompaniment transmitted to the audio signal output device by the central processor from the parameter table, and is output to the loudspeaker, and the central processor also displays the notes of the melody f T on the monitor screen
[0128] 0 ( n ) , the contour of the instantaneous frequency of the fundamental tone of the singer-performer's voice f S
[0129] 0( n ) and the output contour of the fundamental frequency f0̂( n ) .
[0130] 20 Step 13. If the musical composition is not finished, the process is repeated with the input of a new frame of the audio signal of the singer-performer's input voice from the microphone input (go to step 3).
[0131] The processing time from the moment the audio signal is input to the display of the melody notes, the instantaneous frequency contour of the singer's fundamental voice, the output frequency contour of the fundamental tone, and the modified voice of the singer into the loudspeaker is no more than 60 ms. For the user's convenience, as in karaoke, a floating line of the lyrics of the performed vocal piece can be displayed on the monitor screen (Fig. 3).
[0132] The melodies are synchronized with correspondent databases (note contours of 30 melodies). An optimization procedure is applied to the target melody to soften the "computer accent" effect. Original tone contours and voice contours / 16
[0133] Unvoiced decisions are analyzed to find the best moments for transitions between notes. The system can also perform a polyphonic effect, mixing multiple outputs with different target frequency contours.
[0134] Real-time processing is a prerequisite for achieving the technical result 5. This requires double buffering of the singer-performer's input signal: while the current signal frame is being processed according to the generalized functional diagram (Fig. 2), the next frame is buffered. Therefore, the frame length must be at least 60 ms, which corresponds to a buffer size in samples of N = 60 / (sampling frequency period). Thus, for a sampling frequency of 44.1 kHz, the buffer size will be 2646 samples of the singer-performer's input signal. Implementation of the claimed method on the iPhone 5s mobile platform demonstrated the possibility of generating a modified singer-performer's voice in real time with a latency of no more than 60 ms: the input / output of the singer-performer's signal here takes approximately 20 ms, with approximately another 40 ms spent on processing. It should be noted that the iOS operating system of the iPhone 5s mobile platform (multi-core, similar to that shown in Fig.2) utilizes parallel computation: two computational processes can be performed in parallel: 1) synthesis of the harmonic component and determination of the stochastic component, 2) modification of the harmonic component, synthesis of the modified component, and inverse 20-time scaling. In a real device implementing the claimed method, the latency for sound was no more than 60 ms, and for visual display, no more than 100 ms.
[0135] Since the singer-singer signal analysis scheme depends on the instantaneous frequency response (FDR), the analysis efficiency is determined by the frequency-time resolution of the FDR estimation. The FDR estimation algorithm is based on a sinusoidal model, which represents the deterministic portion of the signal as a sum of periodic components with non-stationary parameters:
[0136] K
[0137] s ( n ) ^^ A k ( n ) cos(^ k ( n )) ^ r ( n )
[0138] k ^ 1 (1) n
[0139] где
[0140]
[0141] ^ ^ ^ ^ , K – number of periodic components, r(n)
[0142] i ^ 1 – шумовая component. The parameters of this model (instantaneous amplitude Ak(n ) and frequency ^k(n ) in 30 rad / sample) are used as initial data to estimate the frequency of the fundamental 17
[0143] tones. To obtain the model parameters, the signal s(n) is decomposed into complex subband components by a complex-modulated filter bank, a brief description of which is given below.
[0144] The uniform frequency grid corresponding to the central frequencies of the analysis filter bank5 is defined as k^ step , k ^1,2,..., K , K ^^ / ^ step where ^ step is the frequency step in rad / sample. The impulse response of the k-th analysis filter is determined by the expression:
[0145]
[0146] ^
[0147] where ^ bw is half the bandwidth of the filter, w(n) is an even window function.
[0148] 10 The output of each channel of the filter bank is an analytical signal S k (n ) is a band-limited signal that can be represented as a convolution of the input signal s(n ) with the impulse response:
[0149] ^
[0150]
[0151] The instantaneous parameters of the subband components can be obtained as follows:
[0152]
[0153] ^ k( n) ^ arctan
[0154]
[0155] To avoid discontinuity points, phase unwrapping is used in expression (5). The instantaneous parameters Ak(n), ^k(n) are used as initial data for calculating the period candidate generation function of the PF. Given the assumption that the fundamental frequency variation is proportional to its current value, the parameters of the 20-filter bank of analysis filters should be scaled for each period candidate as follows:
[0156] ^
[0157]
[0158] ^^^ ^ ^ ^ ^^ 18
[0159] where ^ 0 is the candidate frequency in rad / sample, ^ is the permissible relative pitch variation. The duration of the analysis frame ( N ) should be chosen to include an integer number of fundamental pitch periods:
[0160] N ^2^L / ^ 0 , (7)where L is the number of periods in the analysis frame.
[0161] 5 Changing the filter bank parameters for each period candidate is generally computationally expensive. An alternative to this approach is to apply a filter bank with fixed parameters to a signal with a variable sampling frequency. In this case, the sampling frequency F s can be determined as a multiple of the period candidate frequency (f 0 ):
[0162] Fs ^Rf 0 ,
[0163]
[0164] ) 10 where f 0 is the frequency in Hz, R is an integer. Considering that f0 ^^0F S / (2 ^ ) , expression (7) takes the form of a fixed-duration analysis frame for all candidates for the period:
[0165] N ^ RL . (9)The parameter R determines the number of harmonics that are left in the pre-sampled version of the signal:
[0166]
[0167] 15 This approach is illustrated in Fig.8.
[0168] Since only the first few harmonics are of practical significance for determining the frequency response, very short frames can be used for analysis. Expression (6), which describes the scaling of the filter bank parameters, takes the form
[0169] ^
[0170]
[0171] ^^ ^ ^ ^^ 20 The parameters of the sinusoidal model required to calculate the candidate period formation function (CPFF) are extracted from the signal using a multi-speed circuit (Fig. 9), which shows: the resampling stage in channels in accordance with expression (11), an array of complex-modulated filter banks, and the stage of determining the period formation function PF. 19
[0172] As a rule, various metrics based on the autocorrelation function are used as the pitch period candidate formation function (PPCF), for example, the normalized cross-correlation function:
[0173] ^ ^
[0174]
[0175] where l is the delay in counts, ^ ^
[0176]
[0177] ^
[0178] The function averages the data within an analysis frame5 and therefore produces smoothed values. To improve temporal resolution, a normalized cross-correlation function based on a sinusoidal signal model can be used:
[0179]
[0180] This function assumes that the bandwidth of the analysis filters is narrower than the minimum allowable value of the fundamental frequency, as a result of which every 10th harmonic of the signal always falls into a separate channel.
[0181] The multi-rate analysis scheme (Fig. 9) is subject to the phenomenon of harmonic mixing in the channels, which occurs for high-frequency candidates of the fundamental period when the processed voice is low. As a result, rare single emissions appear in the high-frequency region. In order to
[0182]
[0183] 15 To reduce the influence of the harmonic mixing effect, the following FFKP is used, which uses instantaneous parameters obtained for 2V ^ 1 adjacent samples:
[0184]
[0185] ^ ^ In each individual channel of the circuit in Fig. 9, calculation (14) is performed, m, which is responsible for a specific delay value l^1 / ^ , where the index m ^1,... M , for which in 0
[0186]
[0187] Channel 20 is supplied with a signal frame resampled with a coefficient of
[0188]
[0189] ^ (n, l ) in accordance with expression (8). Balancing the values for different ms
[0190] candidates is performed by normalizing each resampled signal frame to unit energy. The use of amplitudes not squared in expression (14), unlike expression (13), allows for the contribution of amplitudes of 20
[0191] various harmonics more balanced. As a rule, the effect of harmonic mixing occurs over short time periods and can be significant K
[0192] reduced by multiplying several terms of the form
[0193]
[0194] ^ cos( ^ k l )
[0195] k ^ 1. Figure 10 shows the period candidates generated by the functions ^(n, l ), ^inst(n, l ) and the proposed5 function ^ms(n, l ) for a short speech fragment, where (a) is the original signal, (b) is the period candidate function ^(n, l ), (c) is the period candidate function obtained based on the sinusoidal model, ^inst(n, l ), (d) is the proposed period candidate function ^ms(n, l ) (V ^ 1 ).
[0196] It is obvious that the function ^ms(n, l ) has a higher frequency and time resolution compared to ^(n, l ) and ^inst(n, l ).
[0197] 10 Thus, the algorithm for estimating the fundamental tone frequency consists of the following steps:
[0198] 1) resample the input signal frame x(n) for each candidate pitch period with the sampling frequency according to expression (8);
[0199] 2) normalize the energy of each resampled frame to 1;
[0200] 15 3) estimate the instantaneous parameters of the sinusoidal model according to expressions (1)–(5). This step is repeated for 2V ^ 1 overlapping frames of each frame. The Hamming window can be used as a window function;
[0201] 4) calculate the function of forming candidates of the period of the fundamental tone frequency according to expression (14), using the corresponding set of parameters;
[0202] 20 5) multiply the obtained value of the function for forming candidates of the fundamental tone period by a weighting function for limiting low-frequency candidates периода:
[0203]
[0204] ^
[0205] 0.8 ;
[0206] 6) the best continuous contour of the fundamental tone frequency is determined using the dynamic programming method, maximizing the sum of the FFCPs on the local 25 frame sequence; as a result of this step, the best candidate^0,best( n ) is selected, which is a rough (initial) estimate of the fundamental tone frequency;
[0207] 7) calculate the refined pitch estimate ^0,fine( n ) using the instantaneous parameters of the sinusoidal model obtained for the best candidate: 21
[0208]
[0209] The computational complexity of the algorithm (the number of multiplications required to estimate a single value of the fundamental frequency) is low, since the implementation of the filter bank is based on the fast Fourier transform. The computational complexity of the algorithm is roughly estimated as
[0210] O ( IKN^ 2( V ^ 1) KN log( N )) , (16)5 where I is the order of the low-pass filter used for decimation / interpolation in the resampling process, N is the number of samples in the analyzed signal frame, K is the number of harmonics of the frequency response, and V is the number of adjacent signal frames.
[0211] For the practical implementation of the algorithm, the following values of parameters were used: K^ 8, L^ 4, R^2 K^1^ 17, N^ 64, M^ 100, I^ 121, V^ 1 (L10 corresponds to formula (7), R – to formula (9), M – to formula (14) and Fig. 9, the number of channels in the filter bank). The argument of the function O(•) shows how the computational complexity changes depending on the parameters I, K, N, V. The search range for the fundamental tone frequency is 50–450 Hz. This range is divided linearly in a logarithmic scale into 100 intervals, each of which corresponds to one candidate for the period of the fundamental tone frequency of 15. The duration of the peridiscretized frames varies from 80 ms (the lowest-frequency candidate) to 9 ms (the highest-frequency candidate).
[0212] The previously published Halcyon pitch estimation algorithm (E. Azarov, M. Vashkevich, A. Petrovsky. Intantaneous pitch estimation algorithm based on multirate sampling / / ICASSP-2016, Shanghai, China, 2016) was compared with five well-known and widely used algorithms: RAPT, YIN, SWIPE, IRAPT, and PEFAC. The comparison was performed in terms of 1) the gross pitch error (GPE) and 2) the mean fine pitch error (MFPE). GPE is calculated as the percentage of voiced frames with a pitch estimation error exceeding ^ 20% of the true pitch value; frames containing gross errors were not taken into account when calculating MFPE. To evaluate the temporal resolution of the algorithm and its robustness to rapid changes in the fundamental frequency, model signals with varying fundamental frequency in the range from 100 to 350 Hz were synthesized.The experimental results obtained were divided into 6 groups depending on the rate of change of the fundamental tone frequency, measured as a percentage of tone change per millisecond (0–0.3, 0.3–0.6, 30 0.6–0.9, 0.9–1.2, 1.2–1.5, >1.5). The average error values are shown in Fig. 11. 22.
[0213] The IRAPT and Halcyon algorithms demonstrate greater robustness to changes in the fundamental frequency—their gross error rate remains negligible, down to 1.5% / ms. The MFPE plot shows that the Halcyon algorithm outperforms all other algorithms in terms of frequency / time resolution. It and 5 were used in implementing the proposed method.
[0214] The most successfully declared method of voice modification with visual and audio feedback is industrially applicable in interactive training systems with the ability to ensure the most accurate matching of the singer-performer's voice to a given melody using the principles of visual and audio 10 (biological) feedback for training in the art and skill of singing.
Claims
23 CLAUSES OF THE INVENTION 1. A method for modifying the voice of a singer-performer with visual and audio feedback, which consists in the fact that provide sets of parameters, each of which characterizes the notes of the target 5 melody, from among the specified sets of parameters, the required set of parameters is specified for a specific target melody in the form of notes of the target melody, the singer-performer's voice signal, which has a frequency spectrum corresponding to it, is introduced synchronously with the musical accompaniment of the target melody's notes, 10 process the singer-performer's voice signal, ensuring the creation of a modified frame of the singer-performer's output voice signal, the target melody f T 0 ( n ) , the contour of the instantaneous frequency of the fundamental tone of the singer-performer's voice f S 0 ( n ) and the output circuit of the fundamental frequency f ( n ), the microphone is indicated ^ 0sh yu a modified frame with musical accompaniment of the target melody and output them to the loudspeaker, and also display the notes of the target melody on the monitor screen f T 0 ( n ) , the contour of the instantaneous frequency of the fundamental tone of the singer-performer's voice f S 0( n ) and the output circuit of the fundamental frequency up and et^f 0 ( n ), ^ The time from the moment the singer's voice signal is input to the output of the modified frame to the loudspeaker is no more than 60 ms, and the time for the output of the target melody notes to the monitor screen is up to 20 ms. 0 ( n ) , the contour of the instantaneous frequency of the fundamental tone of the singer-performer's voice f0 S ( n ) , the output frequency contour of the fundamental tone f 0 ( n ) is no more than 100 ms.
2. A method for modifying the voice of a singer-performer with visual and audio feedback, which consists in the fact that 25 provide sets of parameters, each of which characterizes the notes of the target melody, from among the specified sets of parameters, the required set of parameters is specified for a specific target melody in the form of notes of the target melody, the signal 30 of the singer-performer's voice, having a frequency spectrum corresponding to it, is introduced synchronously with the musical accompaniment of the notes of the target melody, 24 determine the frequency contour of the singer-performer's fundamental tone, perform a temporal scaling of the input frame of the singer-performer's voice signal, carry out a parametric analysis dependent on the fundamental tone frequency using a complex-modulated filter bank for analyzing the input frame of the singer-performer's voice signal to obtain a vector of parameters - instantaneous amplitudes AS(n), instantaneous phase values ^S(n) and the noise component r S(n) of the signal, and form an output contour of the fundamental tone frequency f (n) in accordance with the target melody f T 0 ( n ) , 0 ^ 0 1 transform the parameters of the input frame of the singer-performer's voice signal to obtain a vector of output parameters [A(n), f(n), ^(n), r^(n)], based on the amplitudes A(n) and the corresponding ^0 parameters of instantaneous ^ phase values ^ сигнала the voices of the singer-performer, carry out parametric synthesis of the modified harmonic15 component in accordance with the parameters [A ^( n), ^f 0 ( n), ^(n) ] produce the reverse time ^, scaling of the modified harmonic component using a complex-modulated synthesis filter bank, form a modified frame of the output signal of the singer-20 performer's voice as the sum of the modified harmonic component and the noise component r S(n ), and mix the modified frame of the singer-performer's output voice signal with the musical accompaniment of the target melody and output them to the loudspeaker, and also output the notes of the target melody f T to the monitor screen 0 ( n ) , 25 contour of the instantaneous frequency of the fundamental tone of the singer's voice f0 S ( n ) and the output contour of the fundamental frequency f ( n ) .
3. С оо п п2 ол чю ^ 0 p s b o . , t i a , which is characterized by the fact that the processing time from the moment of input of the singer-performer's voice signal to the output to the loudspeaker is no more than 60 ms, and until the output of the notes of the target melody to the monitor screen, the contour of the instantaneous frequency of the main 25 the tone of the singer's voice, the output frequency contour of the fundamental tone is no more than 100 ms.
4. The method according to paragraph 2, characterized in that, synchronously with the output of the notes of the target melody, the contour of the instantaneous frequency of the fundamental tone of the singer-performer’s voice and the output contour of the fundamental tone frequency, a line of the text of the performed vocal work is output to the screen.
5. The method according to paragraph 2, characterized in that in order to determine the fundamental frequency of the singer-performer's voice signal, the deterministic part of the model of which is the sum of periodic components with non-stationary parameters - 10 amplitude, frequency and phase, the input frame of the singer-performer's voice signal is resampled for each candidate of the fundamental frequency period, the energy of each resampled input frame is normalized to 1, the instantaneous parameters of the sinusoidal model are estimated, the function of forming candidates of the fundamental frequency period is calculated for the corresponding set of 15 parameters - amplitude, frequency and phase, the calculated value of the function of forming candidates of the fundamental frequency period is multiplied by a weighting function for limiting low-frequency candidates of the fundamental frequency, the best continuous contour of the fundamental frequency is determined using dynamic programming,maximizing the sum of the 20-period candidate generation function on the local frame sequence, select the best candidate for the fundamental frequency, which is the initial estimate of the fundamental frequency, and then calculate a refined estimate of the fundamental frequency using the instantaneous parameters of the sinusoidal model obtained for the best candidate.
Citation Information
Patent Citations
Vocal music vocal training device
CN110033670A
Song recording method, sound correction method and electronic device
EP3905246A1
Method of modifying voice and device therefor (versions)
RU2591640C1
Computer-aided learning system employing a pitch tracking line
US20050262989A1