Information processing device, information processing method, and program
The voice quality conversion system effectively addresses the challenge of converting singing voices to specific speakers in real-time by utilizing a voice quality conversion unit with speaker feature estimation, achieving high-quality and robust voice transformation despite noise and data variations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2022-02-09
- Publication Date
- 2026-07-22
AI Technical Summary
Existing voice quality conversion technologies struggle to accurately convert singing voices to specific speakers, especially in real-time, due to variations in pitch and voice quality, and are hindered by noise from sound source separation, making high-quality conversion difficult.
A voice quality conversion system using a voice quality conversion unit with an encoder, feature mixing unit, and decoder, which extracts and mixes speaker features from both vocal signals, and a speaker feature estimation unit that estimates features over varying durations, enabling high-quality real-time conversion by utilizing parallel data relationships and robustness to sound source separation noise.
Enables high-quality, real-time conversion of singing voices to match specific speakers, even with noisy sound sources, by accurately estimating and blending speaker features, ensuring smooth and robust voice quality transformation.
Smart Images

Figure 0007893251000079 
Figure 0007893251000080 
Figure 0007893251000081
Abstract
Description
Technical Field
[0006] , , , ,
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Proposals have been made regarding voice quality conversion technology for converting the voice quality of one's own speech (including singing) into the voice quality of another company. Voice quality refers to the attributes of human speech generated by a speaker and perceived by a listener over a plurality of voice units (e.g., phonemes). More specifically, it refers to elements that are perceived as different by the listener even for utterances with the same pitch and timbre. Patent Document 1 below describes voice quality conversion technology for converting general speech voices into the voice quality of another speaker while maintaining the speech content.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In this field, it is desired to perform appropriate voice quality conversion processing.
[0005] One of the objectives of the present disclosure is to provide an information processing apparatus, an information processing method, and a program for performing appropriate voice quality conversion processing.
Means for Solving the Problems
[0006] The present disclosure, for example, a voice quality conversion unit that inputs a first vocal signal separated from a mixed sound signal and a second vocal signal that has been picked up, and performs voice quality conversion on the second vocal signal using the first vocal signal, <00 The voice quality conversion unit includes an encoder, a feature mixing unit, and a decoder. The encoder extracts feature quantities from the first vocal signal and the second vocal signal. The feature mixing unit mixes the features extracted by the encoder. The decoder generates a vocal signal by transforming the voice quality of the second vocal signal based on the features supplied from the feature mixing unit and the speaker features supplied from the speaker feature estimation unit. The speaker feature estimation unit includes a first speaker feature estimation unit, a second speaker feature estimation unit, and a feature merging unit. The first speaker feature estimation unit performs the estimation for a predetermined time or longer. First Vocal signal and the second vocal signal Based on this, we estimate the features related to the speaker, The second speaker feature estimation unit is shorter than a predetermined time. First Vocal signal and the second vocal signal Based on this, we estimate the features related to the speaker, The feature merging unit merges the speaker features estimated by the first speaker feature estimation unit with the speaker features estimated by the second speaker feature estimation unit. The speaker feature estimation unit estimates the combined features of the feature concatenation unit as speaker features. It is an information processing device.
[0007] This disclosure includes, for example, A voice quality conversion unit, which receives a first vocal signal separated from the mixed sound signal and a second vocal signal that has been captured, uses the first vocal signal to perform voice quality conversion on the second vocal signal. The speaker feature estimation unit estimates speaker features related to the speaker. The encoder in the voice conversion unit extracts the characteristic quantities of the first vocal signal and the second vocal signal. The feature mixing unit of the voice conversion unit mixes the feature quantities extracted by the encoder, The decoder in the voice quality conversion unit generates a vocal signal by converting the voice quality of the second vocal signal based on the feature quantities supplied from the feature quantity mixing unit and the speaker feature quantities supplied from the speaker feature quantity estimation unit. The first speaker feature estimation unit of the speaker feature estimation unit has been performing the estimation for a predetermined time or longer. First Vocal signal and the second vocal signal Based on this, we estimate the features related to the speaker, The second speaker feature estimation unit of the speaker feature estimation unit is shorter than a predetermined time. First Vocal signal and the second vocal signal Based on this, we estimate the features related to the speaker, The feature concatenation unit of the speaker feature estimation unit concatenates the speaker features estimated by the first speaker feature estimation unit and the speaker features estimated by the second speaker feature estimation unit. The speaker feature estimation unit estimates the combined features of the feature concatenation unit as speaker features. It is an information processing method.
[0008] This disclosure includes, for example, A voice quality conversion unit, which receives a first vocal signal separated from the mixed sound signal and a second vocal signal that has been captured, uses the first vocal signal to perform voice quality conversion on the second vocal signal. The speaker feature estimation unit estimates speaker features related to the speaker. The encoder in the voice conversion unit extracts the characteristic quantities of the first vocal signal and the second vocal signal. The feature mixing unit of the voice conversion unit mixes the feature quantities extracted by the encoder, The decoder in the voice quality conversion unit generates a vocal signal by converting the voice quality of the second vocal signal based on the feature quantities supplied from the feature quantity mixing unit and the speaker feature quantities supplied from the speaker feature quantity estimation unit. The first speaker feature estimation unit of the speaker feature estimation unit has been performing the estimation for a predetermined time or longer. First Vocal signal and the second vocal signal Based on this, we estimate the features related to the speaker, The second speaker feature estimation unit of the speaker feature estimation unit is shorter than a predetermined time. FirstVocal signal and the second vocal signal Based on this, estimate the feature amount related to the speaker, The feature amount combining part included in the speaker feature amount estimation part combines the feature amount related to the speaker estimated by the first speaker feature amount estimation part and the feature amount related to the speaker estimated by the second speaker feature amount estimation part, The speaker feature amount estimation part estimates the combined feature amount by the feature amount combining part as the speaker feature amount It is a program that causes a computer to execute an information processing method.
Brief Description of Drawings
[0009] [Figure 1] FIG. 1 is a diagram for explaining the outline of an embodiment. [Figure 2] FIG. 2 is a block diagram showing a configuration example of a smartphone according to an embodiment. [Figure 3] FIG. 3 is a block diagram showing a configuration example of a voice quality conversion part according to an embodiment. [Figure 4] FIG. 4 is a diagram for explaining an example of learning performed by the voice quality conversion part according to an embodiment. [Figure 5] FIG. 5 is a diagram referred to when explaining the operation of a smartphone according to an embodiment. [Figure 6] FIG. 6 is a diagram for explaining an example of a process performed in association with the voice quality conversion process performed in an embodiment. [Figure 7] FIG. 7 is a diagram for explaining another example of a process performed in association with the voice quality conversion process performed in an embodiment. [Figure 8] FIG. 8 is a diagram for explaining a modification example. [Figure 9] FIG. 9 is a diagram for explaining a modification example.
Mode for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. The description will be made in the following order. <Background of the Present Disclosure> <One Embodiment> <Variation> The embodiments described below are preferred examples of the present disclosure, and the content of the present disclosure is not limited to these embodiments.
[0011] <Background of this disclosure> First, to facilitate understanding of this disclosure, the background of this disclosure will be explained. In recent years, in karaoke, instead of using pre-created MIDI (Musical Instrument Digital Interface) sound sources or recorded sound sources as accompaniment, it has become increasingly common to separate the original sound source containing vocals into a vocal signal and an accompaniment signal, and then use the separated accompaniment signal.
[0012] The evolution of such sound source separation technology offers advantages such as reduced costs for creating backing tracks and the ability to enjoy karaoke with the original song's sound intact. On the other hand, while effects such as reverberation, chorus (which alters the pitch of the singing voice), and voice changers (which change the voice quality to an unspecified one) are commonly used in karaoke, it is still difficult to change the singing voice to that of a specific person. Therefore, it is difficult to smoothly convert one's voice quality to that of a specific singer, for example, by "making one's voice slightly closer to the original artist's voice."
[0013] As described in Patent Document 1 above, voice conversion technologies have been proposed that convert ordinary speech to the voice quality of another speaker while preserving the content of the speech. However, singing voices generally have more variations in pitch, voice quality, and various musical expressions (such as vibrato) compared to ordinary speech, making voice conversion difficult. Therefore, currently, voice conversion is limited to unspecified voice qualities such as robotic or anime-like voices, or gender conversion, or voice conversion of specific speakers for whom a sufficient amount of clean audio can be obtained beforehand. It is difficult to convert to speakers for whom a sufficient amount of clean audio cannot be obtained beforehand. Obtaining a sufficient amount of clean audio generally takes a lot of time and cost, and for example, converting to the voice of a famous singer is practically very difficult.
[0014] Furthermore, in karaoke applications, real-time voice quality conversion is necessary, and since future information cannot be used, high-quality conversion is even more difficult. In addition, since the sound source separated by sound source separation may contain noise generated during the sound source separation process, the converted audio that references such separated audio tends to contain a lot of noise, making high-quality conversion even more difficult. Taking the above points into consideration, one embodiment of this disclosure will be described in detail below.
[0015] <One Embodiment> [Summary of one embodiment] First, an overview of one embodiment will be described with reference to Figure 1. A sound source separation process (PA) is performed on the mixed sound source shown in Figure 1. The mixed sound source can be provided by a recording medium such as a CD (Compact Disc) or by distribution via a network. The mixed sound source includes, for example, the artist's vocal signal (an example of a first vocal signal, which will also be appropriately referred to as the vocal signal VSA). The mixed sound source also includes signals other than the vocal signal VSA (such as instrument sounds, which will also be appropriately referred to as accompaniment signals).
[0016] Meanwhile, the singing voice of the karaoke user is captured by a microphone or similar device. The user's singing voice (an example of a second vocal signal) is also referred to as the vocal signal VSB as appropriate.
[0017] A voice quality conversion process (PB) is performed on the vocal signals VSA and VSB. In the voice quality conversion process (PB), one of the vocal signals, VSA or VSB, is made to sound more like the other. At this time, the amount of change to make one of the vocal signals sound more like the other can be set according to a predetermined control signal. For example, the vocal quality conversion process is performed to make the karaoke user's vocal signal VSB sound more like the artist's vocal signal VSA. Then, an addition process (PC) is performed to add the vocal signal VSB that has undergone the voice quality conversion process to the accompaniment signal, and a playback process (PD) is performed on the signal that has undergone the addition process (PC). As a result, the user's singing voice, which has been converted to sound more like the artist's vocal signal, is played back.
[0018] [Example of an information processing device configuration] (Example of overall structure) Figure 2 is a block diagram showing an example configuration of an information processing device according to one embodiment. An example of the information processing device according to this embodiment is a smartphone (smartphone 100). Using smartphone 100, a user can easily perform karaoke with voice quality conversion. In this embodiment, karaoke, i.e., singing, is used as an example for explanation, but this disclosure is not limited to singing and can also be applied to voice quality conversion processing for speech such as conversation. Furthermore, the information processing device according to this disclosure is not limited to smartphones, but can also be applied to portable electronic devices such as smartwatches, personal computers, and stationary karaoke devices.
[0019] The smartphone 100 includes, for example, a control unit 101, a sound source separation unit 102, a voice quality conversion unit 103, a microphone 104, and a speaker 105.
[0020] The control unit 101 comprehensively controls the entire smartphone 100. The control unit 101 is configured as, for example, a CPU (Central Processing Unit) and has ROM (Read Only Memory) where programs are stored and RAM (Random Access Memory) used as work memory (note that these memories are not shown in the diagram).
[0021] The control unit 101 has a speaker feature estimation unit 101A as a functional block. The speaker feature estimation unit 101A estimates feature quantities corresponding to features that do not change over time as the singing progresses, specifically, feature quantities related to the speaker (hereinafter referred to as speaker features as appropriate).
[0022] Furthermore, the control unit 101 has a feature mixing unit 101B as a functional block. The feature mixing unit 101B mixes, for example, two or more speaker features with appropriate weights.
[0023] The sound source separation unit 102 separates the input mixed sound signal into a vocal signal and an accompaniment signal (sound source separation processing). The sound source separated vocal signal is supplied to the voice quality conversion unit 103. The sound source separated accompaniment signal is supplied to the speaker 105.
[0024] The voice quality conversion unit 103 performs voice quality conversion processing on the vocal signal corresponding to the user's singing voice picked up by the microphone 104, so that it approaches the voice quality of the vocal signal separated by the sound source separation unit 102. Details of the processing performed by the voice quality conversion unit 103 will be described later. In this embodiment, voice quality includes not only speaker features but also features such as pitch and volume.
[0025] The microphone 104 captures, for example, the singing or speech (singing in this example) of the user of the smartphone 100. The vocal signal corresponding to the captured singing is supplied to the voice conversion unit 103.
[0026] The accompaniment signal supplied from the sound source separation unit 102 and the vocal signal output from the voice quality conversion unit 103 are added together by an additive unit (not shown). The added signal is then played back from the speaker 105.
[0027] Note that the smartphone 100 may have configurations other than those shown in Figure 2 (for example, a display configured as a touch panel or buttons).
[0028] (Example of the voice quality conversion unit configuration) Figure 3 is a block diagram showing an example configuration of the voice conversion unit 103. The voice conversion unit 103 includes an encoder 103A, a feature mixing unit 103B, and a decoder 103C. The encoder 103A extracts features from the vocal signal using a learning model obtained through predetermined learning. The features extracted by the encoder 103A are, for example, features that change over time as the singing progresses, and specifically include at least one of pitch information, volume information, and speech (lyrics) information.
[0029] The feature mixing unit 103B mixes the features extracted by the encoder 103A. The features mixed by the feature mixing unit 103B are supplied to the decoder 103C.
[0030] The decoder 103C generates a vocal signal based on the feature quantities supplied from the feature quantity mixing unit 103B and the speaker feature quantities.
[0031] (Regarding the learning process conducted in the voice conversion section) Next, an example of the learning method performed in the voice conversion unit 103 will be described with reference to Figure 4. Note that in Figure 4, the feature mixing unit 103B and the feature mixing unit 101B in the voice conversion unit 103 are omitted from the illustration.
[0032] During training, the voice conversion unit 103 is trained using vocal signals from multiple singers (which may include normal speech). The vocal signals may be parallel data in which multiple singers sing the same content, or they may not be parallel data. In this example, non-parallel data, which is more realistic and difficult to train, is used. As shown in Figure 4, the vocal signals from multiple singers are stored in a suitable database 110.
[0033] A predetermined vocal signal is input as input singing data x to the speaker feature estimation unit 101A and encoder 103A described above. The speaker feature estimation unit 101A estimates speaker features from the input singing data x. The encoder 103A extracts from the input singing data x, for example, pitch information, volume information, and spoken content (lyrics) as examples of features. These features are defined, for example, by an embedding vector represented by a multidimensional vector. Each feature defined by the embedding vector is then processed, Speaker embedding TIFF0007893251000001.tif7163 Pitch embedding TIFF0007893251000002.tif7163 Volume embedding TIFF0007893251000003.tif7163 Content embedding TIFF0007893251000004.tif7163 It may be referred to as such as appropriate.
[0034] Decoder 103C takes these features as input and performs the process of constructing speech. During training, Decoder 103C is trained so that its output reconstructs the input singing voice data x. For example, Decoder 103C is trained to minimize the loss function between the input singing voice data x calculated by the loss function calculation unit 115 shown in Figure 4 and the output of Decoder 103C.
[0035] By training the speaker feature estimation unit 101A and encoder 10AC so that each embedding reflects only the corresponding feature and does not contain information about other features, it is possible to transform only the corresponding features by replacing some embeddings with others during inference. For example, Speaker embedding TIFF0007893251000005.tif7163 By replacing only certain elements with those of another person, it is possible to transform the voice quality (voice quality in the narrow sense, excluding pitch, etc.) while preserving pitch, volume, and speech content. Thus, methods for obtaining an embedding vector that separates features include obtaining embedding from feature quantities that reflect only specific features, and training an encoder that extracts only specific features from data (a given vocal signal).
[0036] As the former, the fundamental tone f0 is extracted using a fundamental tone extractor. Pitch embedding TIFF0007893251000006.tif8163 To obtain Average power p to volume embedding TIFF0007893251000007.tif8163 To obtain Speaker embedding from speaker label n TIFF0007893251000008.tif8163 To obtain Features obtained from speech recognition TIFF0007893251000009.tif7163 Content embedding from (Automatic Speech Recognition) TIFF0007893251000010.tif7163 There are methods such as obtaining it.
[0037] The latter method (a method of training an encoder that extracts only specific features from data) can be considered using adversarial learning or methods involving information loss through quantization. For example, Pitch embedding TIFF0007893251000011.tif7163 Volume embedding TIFF0007893251000012.tif7163 Speaker embedding TIFF0007893251000013.tif7163 Each of these is obtained through adversarial learning. In addition, content embedding, for which obtaining accurate labels is difficult. TIFF0007893251000014.tif7163 This can be obtained by learning from data.
[0038] As a concrete example, content embedding TIFF0007893251000015.tif7163 This section describes an example of learning performed on encoder 103A, which extracts data. First, a specific example using adversarial learning will be explained.
[0039] From input vocal data x, Content embedding TIFF0007893251000016.tif7163 Encoder for extracting TIFF0007893251000017.tif7163 teeth, Content embedding TIFF0007893251000018.tif7163 From other features TIFF0007893251000019.tif9163 Critics who estimate TIFF0007893251000020.tif7163 Loss function using TIFF0007893251000021.tif7163 Loss function for reconstructing the input TIFF0007893251000022.tif7163 You can learn by adding this.
[0040] Specifically, learning is performed using the following formula. TIFF0007893251000023.tif20167 TIFF0007893251000024.tif11163 However, in the above formula TIFF0007893251000025.tif7163 This shows the loss function for training encoder 103A and decoder 103C. Also, TIFF0007893251000026.tif7163 is a critic TIFF0007893251000027.tif7163 This is the loss function for, TIFF0007893251000028.tif8163 These are the weight parameters. TIFF0007893251000029.tif7163 TIFF0007893251000030.tif7163 TIFF0007893251000031.tif7163 TIFF0007893251000032.tif7163 TIFF0007893251000033.tif7163 Each of these is a parameter of encoder 103A and decoder 103C, TIFF0007893251000034.tif8163 is a critic TIFF0007893251000035.tif7163 These are the parameters.
[0041] Next, we will explain a specific example of a method that utilizes information loss due to quantization. Content embedding from input vocal data x TIFF0007893251000036.tif7163 Encoder for extracting TIFF0007893251000037.tif7163 The output is vector-quantized and the information is compressed, thereby reducing the other information given to the decoder. TIFF0007893251000038.tif8163 Content embedding only information not included in the main content. TIFF0007893251000039.tif7163 It can be guided to hold it in that position.
[0042] Learning can be performed by minimizing the following loss function. TIFF0007893251000040.tif17163 However, sg() is a gradient stopping operator that prevents the neural network's gradient information from being passed to subsequent layers, and V() is a vector quantization operation. Loss function for reconstruction TIFF0007893251000041.tif7163 Regarding this, various forms are possible depending on the type of decoder or encoder. For example, in the case of a variational autoencoder (VAE) or a vector quantization VAE, variational lower bound (ELBO) TIFF0007893251000042.tif30163 This can be used, and in the case of a Generative Adversarial Network, the input-output error and adversarial loss are used. TIFF0007893251000043.tif7163 It can be expressed as a weighted sum (using the formula below). TIFF0007893251000044.tif10163
[0043] The learning process described above is performed without altering the speaker information estimated by the speaker feature estimation unit. Once learned, the speaker information may change. Furthermore, future information may be used during the learning process.
[0044] In the above, speaker embedding, which determines voice quality, uses speaker label n. TIFF0007893251000045.tif8163 The method for determining this was explained. However, this method requires that the target singer be included in the training data beforehand, and it is not possible to perform voice conversion on an arbitrary singer (unknown speaker). Therefore, a method for determining speaker embedding from an audio signal will be explained. For example, the following two methods can be considered.
[0045] The first method is a speaker embedding estimation method that estimates the speaker information of a given speaker (for example, a speaker whose singing voice data has similar characteristics to the singing voice data of the target singer) based on the vocal signal of that speaker. Speaker embedding learned using speaker label n TIFF0007893251000046.tif8163 The singing sound of speaker n TIFF0007893251000047.tif7163 The speaker feature estimation unit F() is trained, which estimates speaker features from the speaker embedding. F can be constructed as a neural network or similar and is trained to minimize the distance from the speaker embedding. The distance is measured using the Lp norm. TIFF0007893251000048.tif11163 It can be used.
[0046] The second method involves training a singer identification model that estimates the speaker information of a given speaker based on a predetermined vocal signal. singing sound TIFF0007893251000049.tif7163 Speaker embedding TIFF0007893251000050.tif8163 The speaker feature estimation unit G(), which extracts the speaker's characteristics, is trained prior to the training of the voice quality conversion unit 103. G can be trained by minimizing the following objective function L using singing sound data of multiple singers with singer labels. TIFF0007893251000051.tif9165 However, K(x,y) is the cosine distance between x and y. TIFF0007893251000052.tif7163 These are different singing voices by singer n. TIFF0007893251000053.tif7163 This is a vocal recording by singer (m≠n). Using G learned in this way, speaker embedding TIFF0007893251000054.tif8163 The following values are obtained and used for training the voice quality conversion unit 103. TIFF0007893251000055.tif17163
[0047] In any of the above methods, it is preferable that the input speech to the speaker feature estimation unit G() be sufficiently long in order to obtain accurate speaker embedding. This is because it is not possible to sufficiently extract the singer's features from short speech. On the other hand, excessively long input has the disadvantage of requiring a huge amount of memory. Therefore, a recurrent neural network can be used in G(), or the average of speaker embeddings obtained using multiple short segments can be used.
[0048] [Example of operation] Voice conversion is performed by the voice conversion unit 103, which has learned the voice quality as described above. The voice conversion process performed in the smartphone 100 will be explained with reference to Figure 5.
[0049] In Figure 5, the vocal signal VSB is the singing voice data of the karaoke user. The vocal signal VSA is the singing voice data of the singer whose voice the karaoke user wants to emulate, and is a vocal signal separated from the sound source.
[0050] The vocal signal VSA and the vocal signal VSB are input to the voice quality conversion unit 103. The encoder 103A extracts characteristic quantities such as pitch and volume from the vocal signal VSA and the vocal signal VSB.
[0051] The feature mixing unit 103B receives, for example, a control signal that specifies the feature to be replaced. For example, if a control signal is input that converts the pitch information extracted from the vocal signal VSB to the pitch information extracted from the vocal signal VSA, the feature mixing unit 101B replaces the pitch information extracted from the vocal signal VSB with the pitch information extracted from the vocal signal VSA. The features mixed by the feature mixing unit 101B are then input to the decoder 103C.
[0052] The vocal signals VSA and VSB are input to the speaker feature estimation unit 101A. The speaker feature estimation unit 101A estimates speaker information from each vocal signal. The estimated speaker information is supplied to the feature mixing unit 101B.
[0053] The feature mixing unit 101B receives a control signal indicating whether or not to replace the speaker features, and if so, with what weight. In response to the control signal, the feature mixing unit 101B appropriately replaces the speaker features. For example, if the speaker features obtained from the vocal signal VSB are replaced with those obtained from the vocal signal VSA, the voice quality (voice quality in the narrow sense) defined by the speaker features is replaced from the voice quality of the karaoke user to the voice quality of the singer corresponding to the vocal signal VSA. The speaker features mixed by the feature mixing unit 101B are then supplied to the decoder 103C.
[0054] The decoder 103C generates singing voice data based on the feature quantities supplied from the feature quantity mixing unit 101B and the speaker feature quantities supplied from the feature quantity mixing unit 101B. The generated singing voice data is played back from the speaker 105. As a result, a singing voice is played back in which some of the voice qualities of the karaoke user are replaced with some of the voice qualities of a professional singer or other vocalist.
[0055] [Processing performed in conjunction with voice quality conversion processing] Next, we will explain the processes that are performed in conjunction with the voice conversion process. First, we will explain the process that achieves smooth voice conversion. There is a demand to change one's own singing voice to that of the original singer for use in karaoke and other applications. This is because, during inference (when the voice conversion process is executed), the singing voice of singer A (yourself) is changed to that of another singer (the original singer), for example, by embedding the speaker of singer A. TIFF0007893251000056.tif8163 singer B's speaker embedding TIFF0007893251000057.tif8163 This can be achieved by replacing it with...
[0056] However, for uses such as karaoke, there is a demand not to completely change one's singing voice to that of singer B, but rather to make it slightly similar to singer B's voice. To achieve this, speaker embedding of singer A is used. TIFF0007893251000058.tif8163 singer B's speaker embedding TIFF0007893251000059.tif8163 An interpolation function that smoothly changes the value. TIFF0007893251000060.tif8163 The following is used: α is a scalar variable that determines the amount of change, and can also be determined by the user. Linear interpolation or spherical linear interpolation can be used as the interpolation function.
[0057] In addition, TIFF0007893251000061.tif8163 but also TIFF0007893251000062.tif7163 TIFF0007893251000063.tif7163 TIFF0007893251000064.tif7163 Similarly, this can be interpolated using linear interpolation or spherical linear interpolation. For example, the pitch of a karaoke user. TIFF0007893251000065.tif9163 The original singer's pitch TIFF0007893251000066.tif9163 If you want to get closer to it, TIFF0007893251000067.tif10163 Linear interpolation can be performed as shown.
[0058] Next, let's discuss real-time processing. Many common singing voice conversion algorithms use batch processing that utilizes past and future information. However, for applications like karaoke, real-time conversion is required. In this case, the inability to use future information makes it difficult to achieve high-quality conversion.
[0059] Therefore, in this embodiment, we focus on the parallel data relationship between the singing in the original sound source and the user's singing, which in most cases constitutes identical utterances (lyrics), and utilize this characteristic to enable high-quality conversion even in real time. Below, we will describe a specific example of the process that realizes such conversion.
[0060] First, the encoder 103A and decoder 103C of the voice conversion unit 103 are made into functions that do not utilize future information. This can be achieved by constructing the encoder 103A and decoder 103C, which are composed of recurrent neural networks (RNNs) or convolutional neural networks (CNNs), using unidirectional RNNs or causal convolutions that do not utilize future information.
[0061] While this enables real-time processing, accurate estimation of speaker embeddings requires sufficiently long inputs. Therefore, for the initial period after singing begins, sufficient input length is unavailable, making high-quality conversion difficult. To address this, voice conversion in karaoke utilizes the relationship between parallel data during inference, using only short-duration inputs for speaker embedding estimation. Here, "short-duration" refers to the duration of singing speech containing one or a few phonemes, for example, several hundred milliseconds to several seconds. Generally, voice conversion between the same phonemes in different speakers is relatively easy, allowing for high-quality conversion. Therefore, making speaker embedding phoneme-dependent enables high-quality conversion even with short-duration information. However, since the training assumes the absence of parallel data, the model must be trained under the constraint that speaker embeddings are time-invariant. In other words, simply obtaining speaker embeddings from short-duration information—in other words, learning phoneme-dependent speaker embeddings—is not possible.
[0062] Therefore, the encoder 103A and decoder 103C are first trained with time-invariant speaker embedding, and after freezing the parameters of these models, the speaker feature estimator uses these models to estimate the speaker embedding of the incident. TIFF0007893251000068.tif8163 This is learned. Therefore, speaker embedding is treated as an event feature when performing this processing. TIFF0007893251000069.tif7163 The objective function for learning is TIFF0007893251000070.tif12163 It can be expressed as follows. Note that the parameters for encoder 103A and decoder 103C are fixed here. TIFF0007893251000071.tif7163 The receptive field is limited to the above short time interval and can be obtained by minimizing the above objective function.
[0063] The speaker feature estimation unit F, which has been trained in this manner, TIFF0007893251000072.tif7163 This is an estimation machine that determines speaker embeddings based on the specified utterance content (phonemes), enabling high-quality real-time conversion based only on short-term information.
[0064] On the other hand, when singing is sustained for a certain period of time and speaker embedding is required from a sufficiently long input audio, the speaker feature estimation unit F, which has undergone training as explained in Figure 4, etc., may exhibit higher temporal stability.
[0065] Therefore, as shown in Figure 6, for example, the speaker feature estimation unit 101A is configured to have a speaker feature estimation unit that uses long-duration information of a predetermined time or longer (hereinafter appropriately referred to as the global feature estimation unit 121A), a speaker feature estimation unit that uses short-duration information shorter than the predetermined time (hereinafter appropriately referred to as the local (phoneme) feature estimation unit 121B), and a feature concatenation unit 121C. The speaker features can then be obtained using both the global feature estimation unit 121A and the local feature estimation unit 121B. The speaker features obtained from both estimation units are concatenated by the feature concatenation unit 121C and used to determine the final speaker embedding. Weighted linear combinations and spherical linear combinations can be used for concatenation, and the concatenation weight parameters can be determined from the duration and input signal, etc. For example, speaker embedding TIFF0007893251000073.tif7163 This can be calculated as follows. TIFF0007893251000074.tif8163 However, T is the input length from the start of the conversion. α can also be determined as follows, depending only on T. TIFF0007893251000075.tif21163 Alternatively, it can be calculated using a neural network from the input x, as in α(x), or it can be calculated using either T or x information.
[0066] Next, we will explain the process for handling singing mistakes. The real-time processing described above assumes that the singing content in the original song and the user's singing content match during inference (the assumption of parallel data). However, users may make singing mistakes, and this assumption does not always hold true. When speaker embedding is determined using only the short-time input method described above between significantly different phonemes, the quality of the conversion may deteriorate significantly.
[0067] Therefore, when performing this process, a similarity calculation unit 103D is provided in the voice quality conversion unit 103, as shown in Figure 7. The similarity calculation unit 103D performs content embedding of the target singer and the original singer. TIFF0007893251000076.tif7163 The similarity is calculated. The calculation result from the similarity calculation unit 103D is supplied to the speaker feature estimation unit 101A.
[0068] The speaker feature estimation unit 101A modifies the coupling coefficients of global and local features (the weighting of each speaker feature estimated by each speaker feature estimation unit) and the weights of other feature mixtures in speaker feature estimation according to the similarity. Specifically, when the similarity is low, the weight of the coupling to speaker features based on short-time information is reduced because the utterance content is different, thereby lowering the dependence. In other words, the processing results of the global feature estimation unit 121A are mainly used. In addition, in other feature mixtures, excessive transformation is suppressed by increasing the weight of the original speaker's features, thereby suppressing significant degradation of sound quality.
[0069] Next, we will discuss robustness to separated sound sources. Generally, clean, noise-free data is preferred for training in singing voice conversion. However, in this disclosure, the target speaker's singing voice is source-separated audio, and this separation includes noise. Therefore, the noise degrades the estimation accuracy of each embedding, and the sound quality of the converted speech tends to be noisy. To prevent this, we will describe a method for constructing a system robust to source separation noise.
[0070] Robustness against sound source separation noise can be achieved by constraining the encoder, decoder, and speaker feature estimation unit during training so that the embedding vectors extracted from the sound source-separated audio and the original clean audio are identical. Specifically, if the clean audio signal is x, the accompaniment signal is b, and the sound source separator is h(), then the regularization term TIFF0007893251000077.tif11163 Add this to the learning objective function. Here, E is the encoder or feature extractor. The loss function for reconstruction. TIFF0007893251000078.tif7163 By using only clean audio for the calculations related to this, it is possible to keep the output of decoder 103C clean while training encoder 103A so that the feature extraction results from separated audio match those for clean audio.
[0071] While it is preferable that all of the processes associated with the voice conversion process described above be performed, some of these processes may be performed, or they may not necessarily be performed at all.
[0072] <Variation> Although one embodiment of the present disclosure has been described above, the present disclosure is not limited to the embodiment described above, and various modifications are possible without departing from the spirit of the present disclosure.
[0073] Not all of the processing described in one embodiment needs to be performed on the smartphone 100. Some of the processing may be performed by a device other than the smartphone 100, such as a server. For example, as shown in Figure 8, the sound source separation processing and speaker feature estimation processing may be performed by the server, and the voice quality conversion processing and playback processing may be performed on the smartphone. Alternatively, as shown in Figure 9, the sound source separation processing may be performed by the server, and the voice quality conversion processing, playback processing, and speaker feature estimation processing may be performed on the smartphone. Processing results are transmitted and received between the server and the smartphone via a network.
[0074] Furthermore, this disclosure can be implemented in any form, such as an apparatus, method, program, or system. For example, a program that performs the functions described in the embodiments described above can be made downloadable, and an apparatus that does not have the functions described in the embodiments can download and install the program, thereby enabling the apparatus to perform the control described in the embodiments. This disclosure can also be implemented by a server that distributes such a program. In addition, the matters described in each embodiment and modification can be combined as appropriate. Furthermore, the effects exemplified herein should not be interpreted as limiting the content of this disclosure.
[0075] This disclosure may also be structured as follows: (1) It has a voice conversion unit that separates the vocal signal and accompaniment signal from the mixed sound signal and performs voice conversion using the results of the sound source separation. Information processing device. (2) The first vocal signal is separated from the mixed sound signal by the aforementioned sound source separation. The second vocal signal that has been picked up is input to the aforementioned voice quality conversion unit. The voice quality conversion unit brings either the first vocal signal or the second vocal signal closer to the other vocal signal. (1) The information processing device described above. (3) The amount of change required to bring one of the signals closer to the other's vocal signal can be set. (2) The information processing device described above. (4) Furthermore, it has a speaker feature estimation unit that estimates features related to the speaker, The voice quality conversion unit has an encoder and a decoder. (2) The information processing device described above. (5) The aforementioned speaker features are features that do not change over time. The encoder extracts feature quantities corresponding to time-varying features from the input vocal signal. The decoder generates a vocal signal based on the features estimated by the speaker feature estimation unit and the features extracted by the encoder. (4) The information processing device described above. (6) The feature quantity corresponding to the aforementioned time-independent feature is speaker information. The feature quantity corresponding to the aforementioned time-varying feature includes at least one of pitch information, volume information, and speech information. (5) The information processing device described above. (7) The aforementioned feature quantities are defined by the embedding vector. (6) The information processing device described above. (8) The encoder uses a learning model obtained by learning to acquire an embedding vector from feature quantities that reflect only specific features, or by learning to extract only specific features from a vocal signal, to extract the embedding vector of the feature quantity corresponding to the time-changing feature. (7) The information processing device described above. (9) The speaker feature estimation unit estimates speaker features using a learning model obtained through learning that estimates speaker information of a predetermined speaker based on the speaker's vocal signal. An information processing device as described in any of (6) to (8). (10) The speaker feature estimation unit estimates speaker features using a learning model obtained through learning that estimates speaker information of a speaker based on a predetermined vocal signal. An information processing device as described in any of (6) to (8). (11) The speaker feature estimation unit includes a first speaker feature estimation unit and a second speaker feature estimation unit. The system includes a feature merging unit that combines the speaker features estimated by the first speaker feature estimation unit and the speaker features estimated by the second speaker feature estimation unit. An information processing device as described in any of (4) through (10). (12) The first speaker feature estimation unit estimates speaker features based on a vocal signal of a predetermined duration or longer, and the second speaker feature estimation unit estimates speaker features based on a vocal signal shorter than the predetermined duration. (11) The information processing device described above. (13) The coupling coefficient in the feature coupling unit is changed according to the similarity between the first vocal signal and the second vocal signal. (11) The information processing device described above. (14) The aforementioned coupling coefficient is a weighting of the speaker features estimated by the first speaker feature estimation unit and the speaker features estimated by the second speaker feature estimation unit. (13) The information processing device described above. (15) The voice conversion unit separates the vocal signal and accompaniment signal from the mixed sound signal, and performs voice conversion using the results of this sound source separation. Information processing methods. (16) The voice conversion unit separates the vocal signal and accompaniment signal from the mixed sound signal, and performs voice conversion using the results of this sound source separation. A program that instructs a computer to execute information processing methods. [Explanation of Symbols]
[0076] 100... Smartphone 102...Sound source separation section 101A...Speaker Feature Estimation Unit 101B...Speaker feature blending section 103...Voice quality conversion unit 103A... Encoder 103C...Decoder 103D...Similarity calculation section 121A...Global Feature Estimation Unit 121B...Local feature estimation unit
Claims
1. A first vocal signal separated from the sound source from a mixed sound signal and a second vocal signal that has been recorded are input, and a voice quality conversion unit performs voice quality conversion of the second vocal signal using the first vocal signal. It has a speaker feature estimation unit that estimates speaker features related to the speaker, The voice quality conversion unit comprises an encoder, a feature quantity mixing unit, and a decoder. The encoder extracts feature quantities from the first vocal signal and the second vocal signal. The feature mixing unit mixes the feature quantities extracted by the encoder, The decoder generates a vocal signal by converting the voice quality of the second vocal signal based on the feature quantities supplied from the feature quantity mixing unit and the speaker feature quantities supplied from the speaker feature quantity estimation unit. The speaker feature estimation unit includes a first speaker feature estimation unit, a second speaker feature estimation unit, and a feature merging unit. The first speaker feature estimation unit estimates speaker features based on the first vocal signal and the second vocal signal for a predetermined duration or longer. The second speaker feature estimation unit estimates speaker features based on the first vocal signal and the second vocal signal, which are shorter than the predetermined time. The feature merging unit merges the speaker features estimated by the first speaker feature estimation unit with the speaker features estimated by the second speaker feature estimation unit. The speaker feature estimation unit estimates the feature quantities combined by the feature merging unit as the speaker features. Information processing device.
2. The voice quality conversion unit performs voice quality conversion to bring the second vocal signal closer to the first vocal signal, and the amount of change to bring the second vocal signal closer to the first vocal signal is settable. The information processing apparatus according to claim 1.
3. The speaker features relating to the aforementioned speaker are features that correspond to features that do not change over time. The encoder extracts feature quantities corresponding to time-varying features from the input vocal signal. The information processing apparatus according to claim 1.
4. The feature quantity corresponding to the aforementioned time-independent feature is speaker information. The feature quantity corresponding to the aforementioned time-varying feature includes at least one of pitch information, volume information, and speech information. The information processing apparatus according to claim 3.
5. The aforementioned feature quantities are defined by the embedding vector. The information processing apparatus according to claim 4.
6. The encoder uses a learning model obtained by learning to acquire an embedding vector from feature quantities that reflect only specific features, or by learning to extract only specific features from a vocal signal, to extract the embedding vector of the feature quantity corresponding to the time-changing feature. The information processing apparatus according to claim 5.
7. The coupling coefficient in the feature coupling unit is changed according to the similarity between the first vocal signal and the second vocal signal. The information processing apparatus according to claim 1.
8. The aforementioned coupling coefficient is a weighting of the speaker features estimated by the first speaker feature estimation unit and the speaker features estimated by the second speaker feature estimation unit. The information processing apparatus according to claim 7.
9. A voice quality conversion unit, which receives a first vocal signal separated from the sound source of the mixed sound signal and a second vocal signal that has been captured, uses the first vocal signal to perform voice quality conversion on the second vocal signal. The speaker feature estimation unit estimates speaker features related to the speaker. The encoder in the voice conversion unit extracts the characteristic quantities of the first vocal signal and the second vocal signal. The feature quantity mixing unit of the voice quality conversion unit mixes the feature quantities extracted by the encoder, The decoder in the voice quality conversion unit generates a vocal signal in which the voice quality of the second vocal signal has been converted, based on the feature quantities supplied from the feature quantity mixing unit and the speaker feature quantities supplied from the speaker feature quantity estimation unit. The first speaker feature estimation unit of the speaker feature estimation unit estimates speaker features based on the first vocal signal and the second vocal signal for a predetermined duration or longer. The second speaker feature estimation unit of the speaker feature estimation unit estimates speaker features based on the first vocal signal and the second vocal signal which are shorter than the predetermined time, The feature merging unit of the speaker feature estimation unit combines the speaker feature quantities estimated by the first speaker feature estimation unit and the speaker feature quantities estimated by the second speaker feature estimation unit. The speaker feature estimation unit estimates the feature quantities combined by the feature merging unit as the speaker features. Information processing methods.
10. A voice quality conversion unit, which receives a first vocal signal separated from the sound source of the mixed sound signal and a second vocal signal that has been captured, uses the first vocal signal to perform voice quality conversion on the second vocal signal. The speaker feature estimation unit estimates speaker features related to the speaker. The encoder in the voice conversion unit extracts the characteristic quantities of the first vocal signal and the second vocal signal. The feature quantity mixing unit of the voice quality conversion unit mixes the feature quantities extracted by the encoder, The decoder in the voice quality conversion unit generates a vocal signal in which the voice quality of the second vocal signal has been converted, based on the feature quantities supplied from the feature quantity mixing unit and the speaker feature quantities supplied from the speaker feature quantity estimation unit. The first speaker feature estimation unit of the speaker feature estimation unit estimates speaker features based on the first vocal signal and the second vocal signal for a predetermined duration or longer. The second speaker feature estimation unit of the speaker feature estimation unit estimates speaker features based on the first vocal signal and the second vocal signal which are shorter than the predetermined time, The feature merging unit of the speaker feature estimation unit combines the speaker feature quantities estimated by the first speaker feature estimation unit and the speaker feature quantities estimated by the second speaker feature estimation unit. The speaker feature estimation unit estimates the feature quantities combined by the feature merging unit as the speaker features. A program that instructs a computer to execute information processing methods.