Information processing device, information processing method, and information processing program
A three-module system for singing voice synthesis trains HuBERT features from lyrics and MIDI data without labeled data, addressing data scarcity challenges and improving the diversity and applicability of singing voice synthesis.
Patent Information
- Application Number
- PCT/JP2025/004058
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-01
- Filing Date
- 2025-02-07
- Publication Date
- 2025-10-09
AI Technical Summary
Conventional singing voice synthesis technologies require labeled data for training, limiting the breadth of applications and diversity of generated singing voices, especially in cases where data for specific languages, styles, or singers is scarce.
A three-module system is proposed, where Module M1 estimates HuBERT features from lyrics, Module M2 estimates features from MIDI data and HuBERT features, and Module M3 generates singing voices, all trained separately without requiring labeled data, using unlabeled datasets and existing methods to estimate essential sound parameters.
Enables high-quality singing voice synthesis adaptable to user intentions, overcoming data scarcity issues and enhancing the diversity and applicability of singing voice synthesis across languages, styles, and singers.
Smart Images

Figure JP2025004058_09102025_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and information processing program
[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing program.
[0002] In recent years, singing voice synthesis technology has been attracting attention. For example, a technique for synthesizing singing voices using a combination of acoustic and linguistic models is known, using HuBERT (Hidden-Unit BERT).
[0003] Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, Abdelrahman Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” [online], [Retrieved November 22, 2024], Internet <https: / / arxiv.org / abs / 2106.07447>, Ryuichi Yamamoto, Reo Yoneyama, Tomoki Toda, “NNSVS: A Neural Network-Based Singing Voice Synthesis Toolkit,” [online], [Retrieved November 22, 2024], Internet <https: / / arxiv.org / abs / 2210.15987>
[0004] However, conventional techniques require labeled data for training, which limits the breadth of applications and diversity of the singing voice synthesis technology.
[0005] Therefore, the present disclosure proposes an information processing device, an information processing method, and an information processing program that promote improvements in the breadth of applications and diversity of generation in singing voice synthesis technology.
[0006] In order to solve the above problems, an information processing device according to one embodiment of the present disclosure includes a first estimation unit that estimates a first feature from lyrics, a second estimation unit that estimates a second feature from performance data that has been digitized from music performance information and the first feature, and a generation unit that generates a singing voice based on the first feature and the second feature.
[0007] 1 is a diagram showing the relationship between three modules according to an embodiment. FIG. 1 is an explanatory diagram for explaining details of learning of a first model according to an embodiment. FIG. 2 is an explanatory diagram for explaining details of learning of a second model according to an embodiment. FIG. 3 is an explanatory diagram for explaining details of learning of a third model according to an embodiment. FIG. 4 is an explanatory diagram for explaining a modified example of module M3 according to an embodiment. FIG. 5 is an explanatory diagram for explaining an estimation phase according to an embodiment. FIG. 6 is a diagram (1) showing an example of a UI screen of the estimation phase according to an embodiment. FIG. 7 is a diagram (2) showing an example of a UI screen of the estimation phase according to an embodiment. FIG. 8 is a sequence diagram (1) showing an overview of an information processing device according to an embodiment. FIG. 9 is a sequence diagram (2) showing an overview of an information processing device according to an embodiment. FIG. 10 is a diagram showing an example of the configuration of a user terminal according to an embodiment. FIG. 11 is a diagram showing an example of the configuration of an information providing device according to an embodiment. FIG. 12 is a diagram showing an example of the configuration of an information processing device according to an embodiment. FIG. 13 is a diagram showing an example of a model storage unit according to an embodiment. FIG. 14 is a diagram showing an example of a singing voice synthesis result storage unit according to an embodiment. FIG. 15 is a flowchart showing the flow of learning processing (module M1) in a control unit. FIG. 16 is a flowchart showing the flow of learning processing (module M2) in a control unit. FIG. 17 is a flowchart showing the flow of learning processing (module M3) in a control unit. FIG. 18 is a flowchart showing the flow of estimation processing (modules M1 to M3) in a control unit. FIG. 19 is a hardware configuration diagram showing an example of a computer that realizes the functions of an information processing device.
[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0009] The present disclosure will be described in the following order of items: 1. Embodiment 1-1. Overview of information processing system according to embodiment 1-2. Details of learning phase according to embodiment 1-2-1. Details of learning of first model 1-2-2. Details of learning of second model 1-2-3. Details of learning of third model 1-3. Details of estimation phase according to embodiment 1-4. Overview of information processing device according to embodiment 1-4-1. Learning phase of information processing device 1-4-2. Estimation phase of information processing device 1-5. Configuration of user terminal according to embodiment 1-6. Configuration of information providing device according to embodiment 1-7. Configuration of information processing device according to embodiment 2. Other embodiments 3. Hardware configuration
[0010] (1. Embodiments) (1-1. Overview of Information Processing System According to Embodiments) In recent years, singing voice synthesis technology using neural networks has been attracting attention. For example, a technology that uses HuBERT to synthesize singing voices by combining an acoustic model and a linguistic model and performing training is known.
[0011] Singing voice synthesis technology is a technology that generates natural and diverse singing voices from information such as lyrics, MIDI (Musical Instruments Digital Interface) data (an example of performance data that digitizes musical performance information), and sound timing. Singing voice synthesis technology is a technology that can be applied not only to the music field but also to speech fields such as speech and conversation.
[0012] Compared to speech, singing voices are difficult to model because they have a higher pitch and the sound changes depending on the singing style, even when the lyrics are the same. However, in recent years, advances in neural networks have made it possible to synthesize high-quality singing voices.
[0013] One of the reasons why high-quality singing voice synthesis is difficult is that labeled data is required for training. Singing voice synthesis generally requires vocal information about the singing voice to be generated, but collecting such data in large quantities is difficult.
[0014] Deep learning models require labels for training. Singing voice synthesis models generate singing voices from lyrics and MIDI data, so this information is required as labels for training. Since learning is not possible without information such as lyrics and MIDI data, the singing voices that can be trained are limited.
[0015] One example of a situation where singing voice synthesis becomes difficult is when there is a lack of singing voice data for a particular language. If singing voices for a particular language are needed as labels, and there are not enough of such singing voices, it becomes difficult to synthesize singing voices for that language.
[0016] Another example is when there is little data on singing voices in a currently popular style (genre).Singing voices in a currently popular style are needed as labels, and if there are not enough such voices, it becomes difficult to synthesize singing voices in that style.
[0017] Another example is when there is little data on the singing voice of a particular singer. If the singing voice of a particular singer is needed as a label, and there is not enough of that singing voice, it becomes difficult to synthesize the singing voice of that singer.
[0018] As described above, in conventional technologies, the range of applications and the diversity of the results generated by singing voice synthesis technology are limited because labeled data is required for learning. Therefore, this disclosure proposes an information processing device, an information processing method, and an information processing program that promote the improvement of the range of applications and the diversity of the results generated by singing voice synthesis technology.
[0019] In the following embodiment, we propose a method that enables training of a singing voice synthesis model without requiring labeled data. Specifically, the system is divided into three modules: a module that estimates first features (hereinafter referred to as "Module M1"), a module that estimates second features (hereinafter referred to as "Module M2"), and a module that generates singing voices (hereinafter referred to as "Module M3"), and each module is trained separately. Singing voice synthesis is then performed using these three modules.
[0020] In the following embodiments, "audio" refers to all music including vocals, "vocal" refers to vocals included in audio, "speech" refers to non-vocal voices such as speech or conversation, and "lyrics" refers to text information.
[0021] Although details will be described later, in the following embodiment, HuBERT (Hidden Unit BERT) features are used for singing voice synthesis. The HuBERT features are estimated from lyrics, and the singing voice is generated using the HuBERT features.
[0022] Here, the three modules will be briefly explained. FIG. 1 is a diagram showing the relationship between the three modules according to the embodiment. Module M1 estimates a first feature from lyrics (T1). Module M2 estimates a second feature from MIDI data (D1) and the first feature. Module M3 generates a singing voice (Z1) using these feature amounts.
[0023] The information processing system according to the embodiment is divided into a learning phase and an estimation phase. Details of each phase will be described separately below.
[0024] (1-2. Details of the Learning Phase According to the Embodiment) Details of the learning phase according to the embodiment will be described. As described above, learning is performed separately for each module. That is, learning is performed separately for module M1, module M2, and module M3. Hereinafter, the model trained in module M1 will be referred to as the "first model," the model trained in module M2 as the "second model," and the model trained in module M3 as the "third model."
[0025] The first model is a model that estimates a first feature, and estimates the first feature from lyrics. The first feature is a HuBERT feature. The second model is a model that estimates a second feature, and estimates the second feature from MIDI data and the first feature. The third model is a model that generates singing voices. The third model is a diffusion model.
[0026] The first model may be, for example, a model called FastSpeech, and the second model may be, for example, a model called Tactron-AR.
[0027] The learning of each model will be described in detail below with reference to FIGS.
[0028] (1-2-1. Details of Learning the First Model) FIG. 2 is an explanatory diagram for explaining details of learning the first model according to the embodiment. FIG. 2 is an explanatory diagram for explaining details of learning in module M1. In FIG. 2, model N1 (corresponding to the first model) is trained.
[0029] The feature (F1) is a HuBERT feature. HuBERT is a model that estimates features from audio. HuBERT features are information that can be estimated from audio and are output when audio is input.
[0030] The information that comes out when audio is fed into a neural network is called the true HuBERT feature. More precisely, the HuBERT feature that comes out when audio is fed into a neural network is called the true HuBERT feature. The intention is that these are more accurate HuBERT features than the estimates from model N1.
[0031] Using these HuBERT features as targets (teacher), model N1 is trained using lyrics (T1). Model N1 is a model trained to estimate the true HuBERT features from lyrics (T1).
[0032] In other words, true HuBERT features can be estimated from audio, and this is used as a target to train a neural network so that it can estimate true HuBERT features from lyrics (T1). The true HuBERT features are the information that comes out when audio is input into a neural network, and this is used as a target to train another neural network.
[0033] The learning in module M1 is performed by training this model N1 using dataset L1, which includes a dataset of singing voice and speech.
[0034] Additionally, learning is performed by including information indicating the timing of sounds (phonemes), such as when to pronounce them. Specifically, this information is about which sounds to pronounce and when, and which lyrics to pronounce at which sounds. This kind of timing is also important in singing voice synthesis. Learning is performed by including information about which parts of the lyrics (T1) to pronounce and when.
[0035] By learning the timing of the sounds, users can freely decide which lyrics to pronounce at which parts and generate singing voices, which makes it possible to synthesize singing voices that are more in line with the user's intentions.
[0036] Information indicating such timing is referred to as timing (P1). Model N11 is a model capable of estimating the timing (P1) of a sound from audio. A pre-trained model is used for model N11.
[0037] Model N1 is trained using lyrics (T1) and timing (P1) with the true HuBERT features as the target. Model N1 is a model trained to estimate the true HuBERT features from lyrics (T1) and timing (P1).
[0038] The input data for model N1 will be explained. Lyrics (T1) and timing (P1) obtainable from dataset L1 are used to train model N1. By training with a labeled dataset, it becomes possible to handle languages for which there is no vocal dataset.
[0039] On the other hand, during estimation, lyrics (T1) specified by the user and timing (P1) specified by the user are input to the model N1.
[0040] Module M1 is trained by focusing on linguistic features. By training only on these features, singing voice synthesis can be performed without the linguistic features being affected by other data.
[0041] Even if the lyrics are the same, there are many variations, such as different MIDI data (D1) or different vocals, and training data is required for each variation. By narrowing down the features, the training data required for learning is reduced, enabling efficient learning.
[0042] (1-2-2. Details of Learning the Second Model) Next, details of learning the second model will be described. FIG. 3 is an explanatory diagram for explaining details of learning the second model according to the embodiment. FIG. 3 is also an explanatory diagram for explaining details of learning in module M2. In FIG. 3, model N2 (corresponding to the second model) is trained.
[0043] The training in module M2 is performed by training model N2 using singing voice dataset L2, which is an unlabeled dataset.
[0044] Model N21 is a model that can estimate MIDI data (D1) from audio. Model N22 is a model that can estimate a feature (K1) of a singer's singing voice from audio. Model N23 is a model that can estimate a feature (F1) from audio. Models N21 to N23 are all models that use audio as input. Pre-trained models are used for all of models N21 to N23.
[0045] Below, an example will be described in which three pieces of information, namely, MIDI data (D1), feature (K1), and feature (F1), are input to model N2. Of these, feature (K1) is optional and need not be required (it need not be used in learning). If feature (K1) is not included in learning, model N22 becomes unnecessary. Furthermore, feature (K1) becomes unnecessary during estimation. In FIG. 3, feature (K1) is indicated by shading as optional information. Correspondingly, model N22 is also indicated by shading. The shading has been added for ease of understanding and does not indicate any particular limitation.
[0046] The input data for model N2 will be explained below. The input data is MIDI data (D1), feature (K1), and feature (F1). First, regarding the MIDI data (D1) and feature (K1), during training, the MIDI data (D1) estimated from the audio and the feature (K1) estimated from the audio are used for training model N2.
[0047] On the other hand, for the MIDI data (D1) and the feature (K1), at the time of estimation, the MIDI data (D1) and the feature (K1) specified by the user are input to the model N2.
[0048] In this case, the user specifies the audio (any audio of the singer will do). The singing voice is then generated using the singer's voice. If reference audio is available, it can also be specified. By making it possible to specify the feature (K1) by specifying the audio, it is expected that versatility will increase.
[0049] Instead of specifying the audio, the feature (K1) may be directly specified. For example, feature values for the singing voices of a specific number of singers may be prepared in advance, and the feature value specified by the user may be used as the feature (K1).
[0050] The features (F1) are the same as those in module M1. During training, true HuBERT features that can be estimated from audio are used. Therefore, during training, model N2 is trained using true HuBERT features. In contrast, during estimation, HuBERT features (estimated values) estimated by module M1 are used. HuBERT features estimated using the trained model N1 are used.
[0051] In this way, input data for learning the model N2 is generated. Next, the target information for learning the model N2 will be described.
[0052] One of the target information is F0 (A1). F0 (A1) is information that represents the fundamental frequency of the sound. Compared to MIDI data (D1), this information is capable of expressing sound changes that are closer to reality. The other is volume (B1), which represents the loudness of the sound.
[0053] The other is data (C1) indicating whether a sound is voiced or unvoiced. The data (C1) indicating whether a sound is voiced or unvoiced is information that uses the values "0" and "1" to indicate whether a sound is voiced or unvoiced. For ease of explanation, the data indicating whether a sound is voiced or unvoiced will hereinafter be referred to as "voiced / unvoiced_flags" where appropriate.
[0054] The following description will be given using an example in which three pieces of information, F0 (A1), volume (B1), and voiced / unvoiced_flags (C1), are used as target information. Of these, volume (B1) and voiced / unvoiced_flags (C1) are optional and do not necessarily have to be used (they do not have to be used in learning). If volume (B1) or voiced / unvoiced_flags (C1) are not included in learning, they will not be included in the output results during estimation, and will not be taken into account in module M3, described below. In Figure 3, volume (B1) and voiced / unvoiced_flags (C1) are shaded to indicate that they are optional information. The shading has been added for ease of understanding and does not indicate any particular limitations.
[0055] This information (F0 (A1), volume (B1), and voiced / unvoiced_flags (C1)) can be estimated with relatively high accuracy from audio using existing methods. Furthermore, since the input data is also information that can be estimated from audio, all of the information necessary for training model N2 can be estimated from audio.
[0056] Using this information (F0 (A1), volume (B1), and voiced / unvoiced_flags (C1)) as targets, model N2 is trained using MIDI data (D1), feature (K1), and feature (F1). Model N2 is a model trained to be able to estimate F0 (A1), volume (B1), and voiced / unvoiced_flags (C1) from MIDI data (D1), feature (K1), and feature (F1).
[0057] An existing method is used to estimate the information from the audio with a relatively high degree of accuracy, and the information is used as true information to train model N2. Model N2 is a model trained using information that has been estimated with a relatively high degree of accuracy from the audio with an existing method as true information.
[0058] The audio is input into a separate neural network, and the information that emerges (F0 (A1), volume (B1), and voiced / unvoiced_flags (C1)) is used as a target to train another neural network. In other words, model N2 is trained to output F0 (A1), volume (B1), and voiced / unvoiced_flags (C1) when MIDI data (D1), feature (K1), and feature (F1) are input.
[0059] Through learning in module M2, it becomes possible to estimate F0 (A1) and other parameters based on the lyrics (T1), MIDI data (D1), and features (K1) specified by the user, thereby enabling singing voice synthesis that is in line with the user's intentions.
[0060] It should be noted that if a reference such as F0 (A1) of the audio is available, the reference may be obtained, in which case module M2 is not required.
[0061] (1-2-3. Details of Learning the Third Model) Next, details of learning the third model will be described. FIG. 4 is an explanatory diagram for explaining details of learning the third model according to the embodiment. FIG. 4 is also an explanatory diagram for explaining details of learning in module M3. In FIG. 4, model N3 (corresponding to the third model) is trained.
[0062] The training in module M3 is to train model N3 using singing voice dataset L3 (which may be the same dataset as dataset L2). Note that dataset L3 is an unlabeled dataset, just like dataset L2.
[0063] Model N31 is a model capable of estimating a feature (K1) from audio. Model N32 is a model capable of estimating a feature (F1) from audio. Model N33 is a model capable of estimating F0 (A1) from audio. Model N34 is a model capable of estimating volume (B1) from audio. Model N35 is a model capable of estimating voiced / unvoiced_flags (C1) from audio. Models N31 to N35 are all models that use audio as input.
[0064] Of these, models N31 to N33 are neural networks, and pre-trained models are used for all of the models N31 to N33.
[0065] As described in the description of module M2, the feature (K1), volume (B1), and voiced / unvoiced_flags (C1) are optional. If the feature (K1), volume (B1), and voiced / unvoiced_flags (C1) are not included in the learning of module M2, they are also unnecessary for module M3. In this case, models N31, N34, and N35 are unnecessary. In FIG. 4, the feature (K1), volume (B1), and voiced / unvoiced_flags (C1) are indicated by shading as optional information. Correspondingly, models N31, N34, and N35 are also indicated by shading. The shading is added for ease of understanding and does not indicate any particular limitation. Below, an example in which the feature (K1), volume (B1), and voiced / unvoiced_flags (C1) are included will be described.
[0066] The input data for model N3 will be described below. The input data is a feature (K1), a feature (F1), F0 (A1), volume (B1), and voiced / unvoiced_flags (C1). First, for the feature (K1), similar to module M2, during training, the feature (K1) estimated from audio is used for training model N3. Meanwhile, during estimation, the feature (K1) specified by the user is input to model N3.
[0067] The features (F1) are the same as those in modules M1 and M2. During training, true HuBERT features that can be estimated from audio are used. Therefore, during training, model N3 is trained using true HuBERT features. In contrast, during estimation, HuBERT features estimated by module M1 are used. HuBERT features estimated using trained model N1 are used.
[0068] F0 (A1) is the same as that of module M2. During training, an F0 that can be estimated with relatively high accuracy from audio is used. Therefore, during training, model N3 is trained using a relatively high-accuracy F0. In contrast, during estimation, the F0 (estimated value) estimated by module M2 is used. The F0 estimated using the trained model N2 is used.
[0069] During training, F0(A1) is estimated by, for example, an F0 estimation algorithm called Harvest. Model N33 is not limited to this example, as long as it is a model that can estimate F0 from audio.
[0070] The volume (B1) is the same as that of module M2. During learning, volume data that can be estimated from audio with a relatively high degree of accuracy is used. Therefore, during learning, model N3 is trained using volume data with a relatively high degree of accuracy. In contrast, during estimation, volume data (estimated value) estimated by module M2 is used. Volume data estimated using the trained model N2 is used.
[0071] During training, the volume (B1) is estimated by, for example, A-weighting. For example, the volume (B1) is estimated by applying A-weighting to frequencies transformed by a Short Time Fourier Transform (STFT). The model N34 is not limited to this example, as long as it is a model capable of estimating volume data from audio.
[0072] The voiced / unvoiced_flags (C1) are the same as those in module M2. During training, voiced / unvoiced_flags that can be estimated with relatively high accuracy from audio are used. Therefore, during training, model N3 is trained using relatively high-accuracy voiced / unvoiced_flags. In contrast, during estimation, the voiced / unvoiced_flags (estimated values) estimated by module M2 are used. The voiced / unvoiced_flags estimated using the trained model N2 are used.
[0073] During training, voiced / unvoiced_flags (C1) is estimated using an existing analysis technique that determines whether a sound is voiced or unvoiced using a numerical value of "0" or "1." Model N35 is not limited to this example, as long as it is a model that can estimate whether a sound is voiced or unvoiced from audio.
[0074] As mentioned above, the features (K1), volume (B1), and voiced / unvoiced_flags (C1) are not required, but including them enables higher quality singing voice synthesis during estimation.
[0075] In this way, input data for training the model N3 is generated. Next, the target information for training the model N3 will be described. The target information is the audio data (Z1) of singing voice.
[0076] The audio data (Z1) is information that can be acquired from the dataset L3. Using the audio data (Z1) as a target, the model N3 is trained using the feature (K1), the feature (F1), the F0 (A1), the volume (B1), and the voiced / unvoiced_flags (C1). The model N3 is trained so that the audio data (Z1) can be generated from the feature (K1), the feature (F1), the F0 (A1), the volume (B1), and the voiced / unvoiced_flags (C1).
[0077] Another neural network is trained using audio data (Z1) as the target. That is, model N3 is trained to output audio data (Z1) when feature (K1), feature (F1), F0 (A1), volume (B1), and voiced / unvoiced_flags (C1) are input. Model N3 is a diffusion model.
[0078] A modified example of module M3 will be described. FIG. 5 is an explanatory diagram for explaining a modified example of module M3 according to the embodiment. Note that the same explanation as in FIG. 4 will be omitted as appropriate. Furthermore, F0 (A1) is used not only for model N3 but also for model N4 (corresponding to the fourth model). Although the layout in the figure is slightly different from FIG. 4, various processes are the same.
[0079] In FIG. 5, audio data (ZZ1) is generated using model N3, and model N4 is introduced to finally generate audio data (Z1).
[0080] The audio data (ZZ1) is a melspectrogram. Model N4 is SiFiGAN. In this way, frequency data of the audio data (Z1) (corresponding to the audio data (ZZ1)) may be generated, and then a singing voice (corresponding to the audio data (Z1)) may be generated (converted). Not only can the audio data (Z1) be generated directly, but by incorporating the generation of a melspectrogram, the audio data (Z1) can be generated in a wider variety of ways.
[0081] In this case, the model N3 is trained using the audio data (ZZ1) as a target. The model N3 is trained so that it can estimate the audio data (ZZ1).
[0082] The details of the learning phase have been explained above using Figures 2 to 5. Next, the details of the estimation phase will be explained.
[0083] (1-3. Details of the Estimation Phase According to the Embodiment) FIG. 6 is an explanatory diagram for explaining the estimation phase according to the embodiment. In the estimation phase, modules M1 to M3 are executed as a series of flows. For ease of explanation, module M3 will be described using the case of FIG. 4 rather than the modified example of FIG. 5.
[0084] A user desires singing voice synthesis and specifies lyrics (T1), timing (P1), MIDI data (D1), and feature (K1). Feature (K1) is specified by specifying the audio. Feature (K1) is a feature of the singing voice of the singer of the specified audio. Information specified by the user in this way is marked with an asterisk. The asterisks are added for ease of understanding and do not indicate any particular limitation.
[0085] The feature (K1) is optional information and can be omitted depending on the user's intention. In the following, an example of singing voice synthesis including the feature (K1) will be described.
[0086] Lyrics T11 are lyrics specified by the user. Timing P11 is timing information indicating which note is pronounced at which part of lyrics T11. MIDI data D11 is MIDI data specified by the user. MIDI data D11 may be any information that expresses musical performance information. For example, it may be data related to a string of notes. Feature K11 is a feature of the singing voice of the singer of the audio specified by the user (the feature may be specified directly). In Figure 6, the feature K11 is shaded to indicate that it is optional information. The shading has been added to make it easier to understand and does not indicate any particular limitation.
[0087] When the user specifies this information and requests singing voice synthesis, the estimation phase is executed. First, lyrics T11 and timing P11 are input to model N1 (model N1 trained in the training phase). Model N1 outputs features F11 from the input information (lyrics T11 and timing P11). Features F11 are HuBERT features. This corresponds to module M1 in the training phase.
[0088] Next, MIDI data D11, feature K11, and feature F11 are input to model N2 (model N2 that has already been trained in the training phase). Model N2 outputs F0A11, volume B11, and voiced / unvoiced_flags C11 from the input information (MIDI data D11, feature K11, and feature F11). This corresponds to module M2 in the training phase.
[0089] The volume (B1) and voiced / unvoiced_flags (C1) are optional information and can be omitted. The following describes an example in which singing voice synthesis is performed using the volume (B1) and voiced / unvoiced_flags (C1) as well. In Figure 6, the volume B11 and voiced / unvoiced_flags C11 are shaded to indicate that they are optional information. The shading has been added for ease of understanding and does not indicate any particular limitations.
[0090] Next, feature K11, feature F11, F0A11, volume B11, and voiced / unvoiced_flags C11 are input to model N3 (model N3 trained in the training phase). Model N3 outputs audio data Z11 from this input information (feature K11, feature F11, F0A11, volume B11, and voiced / unvoiced_flags C11). This corresponds to module M3 in the training phase.
[0091] When inputting the model N3, the feature amount K11, volume B11, and voiced / unvoiced_flags C11 are optional information and can be omitted.
[0092] Although not shown, audio data ZZ11 is output in the modified example of Fig. 5. In this case, the audio data ZZ11 is a singing voice, and the audio data ZZ11 is a mel spectrogram of the singing voice.
[0093] The above is the details of the processing in the estimation phase. Next, the UI during the estimation phase will be described.
[0094] 7A is a diagram (1) showing an example of a UI screen in the estimation phase according to the embodiment. Screen G1 is a screen displayed on the user terminal 10. Screen G1 includes, for example, an operation button B1 for specifying lyrics (T1), an operation button B2 for specifying timing (P1), an operation button B3 for specifying MIDI data (D1), and an operation button B4 for specifying audio for a feature (K1). Operation button B4 clearly indicates that input is "optional."
[0095] Operating (clicking, tapping, etc.) operation button B1 transitions to a lyrics (T1) input screen. Here, the user freely inputs (or selects) lyrics. Operating operation button B2 transitions to a timing (P1) input screen. Here, the user freely inputs (or selects) timing. Operating operation button B3 transitions to a MIDI data (D1) input screen. Here, the user freely inputs (or selects) MIDI data. Operating operation button B4 transitions to an audio input screen. Here, the user freely inputs (or selects) audio.
[0096] The screen G1 also includes an operation button B5 for instructing singing voice synthesis. When the operation button B5 is operated, singing voice synthesis is executed, i.e., the estimation phase is executed.
[0097] FIG. 7B is a diagram (2) showing an example of a UI screen for the estimation phase according to the embodiment. Screen G2 is a screen displayed on the user terminal 10. Screen G2 is a screen showing the results of the estimation phase executed by operating operation button B5 on screen G1. Screen G2 includes an operation button B11 for playing the generated singing voice and an operation button B111 for downloading the audio data. Operating operation button B11 plays the generated singing voice. That is, audio data (Z1) is played. In the example of FIG. 6, audio data Z11 is played.
[0098] By using the operation button B11, the user can confirm what the generated singing voice actually sounds like. If there is an image (including a still image and a moving image) corresponding to the singing voice, that image (which may be a predetermined image or an image generated in parallel with the estimation phase) may also be played back. The image may be displayed or played back in any manner.
[0099] The screen G2 may also include an operation button B12 for evaluating the generated singing voice. Operating the operation button B12 may transition to a screen where the generated singing voice can be evaluated. Here, the user may freely input (or select) an evaluation.
[0100] The screen G2 may also include an operation button B13 for editing the generated singing voice. Operating the operation button B13 may transition to a screen where the generated singing voice can be edited. Here, the user may freely perform editing. For example, editing may be performed on any of the information designated by the user via the screen associated with the operation buttons B1 to B4 shown in FIG. 7A.
[0101] Screen G2 may also include an operation button B14 for executing singing voice synthesis again based on the edited information. When operation button B14 is operated, singing voice synthesis may be executed again based on the edited information edited by the user via the screen linked to operation button B13. In this case, the above-mentioned singing voice synthesis process may be repeated based on the edited information.
[0102] The above is a detailed description of each phase of the information processing system according to the embodiment. The above is an overview of the information processing system according to the embodiment.
[0103] (1-4. Overview of Information Processing Device According to Embodiment) Next, an overview of the information processing device 100 according to the embodiment will be described. The information processing device 100 is an information processing device that aims to promote the improvement of the breadth of applications and the diversity of generation in singing voice synthesis technology, and may be any device as long as it is capable of implementing the processing according to the embodiment. The information processing device 100 is realized by, for example, a server device or a cloud system, and executes the information processing according to the embodiment.
[0104] The information processing device 100 is realized, for example, by a server device or a cloud system that provides or manages a specific service that enables creators and the like to freely create music, etc. The information processing device 100 is realized, for example, by a server device or a cloud system that provides or manages a specific service that supports creators and the like so that they can freely create music, etc. (for example, by enabling the use of various tools for music creation, etc.).
[0105] The information providing device 200 is an information processing device intended to provide information such as teacher data so that the information processing device 100 can appropriately learn the models (models N1 to N3), and may be any device that can realize the processing in the embodiment. The information providing device 200 is realized by, for example, a server device or a cloud system, and executes the information processing according to the embodiment.
[0106] The user terminal 10 is an information processing device used by a user who desires singing voice synthesis in accordance with the user's intentions. The user may, for example, be a user who creates music while paying particular attention to the subtle timbre of each note, and who desires high-quality singing voice synthesis.
[0107] The user terminal 10 may be any device capable of implementing the processes described in the embodiment. The user terminal 10 may be a smartphone, a tablet terminal, a notebook PC, a desktop PC, a mobile phone, a PDA, or other device. Figure 8B shows a case where the user terminal 10 is a smartphone.
[0108] The user terminal 10 is, for example, a smart device such as a smartphone or tablet, and is a mobile terminal device that can communicate with any server device via a wireless communication network such as 4G to 5G (Generations) or LTE (Long Term Evolution). The user terminal 10 may also have a screen such as a liquid crystal display with touch panel functionality, and may accept various operations on displayed data such as content, such as tapping, sliding, and scrolling, performed by the user using a finger or stylus. In FIG. 8B , the user terminal 10 is used by user U1.
[0109] The information processing system according to the embodiment is not limited to being realized by the information processing device 100 and the information providing device 200 being separate devices, but may be realized by integrating the information processing device 100 and the information providing device 200. Furthermore, the information processing system according to the embodiment is not limited to being realized by the information processing device 100 and the user terminal 10 being separate devices, but may be realized by integrating the information processing device 100 and the user terminal 10.
[0110] (1-4-1. Learning Phase of Information Processing Device) The relationship between the information processing of the information processing device 100 and the information providing device 200 will be described using FIG. 8A. FIG. 8A is a sequence diagram (1) showing an overview of the information processing device according to the embodiment. In FIG. 8A, the information processing system 1 includes the information processing device 100 and the information providing device 200. This corresponds to the learning phase.
[0111] (Learning of Module M1) The information processing device 100 acquires a dataset (corresponding to dataset L1) from the information providing device 200 (step S11).
[0112] The information processing device 100 uses the acquired dataset to train a first model (step S12). This dataset includes a singing voice dataset and a speech dataset, of which the speech dataset is a labeled dataset. The information processing device 100 trains the first model using the speech dataset. This makes it possible to handle languages that do not have labels as singing voice datasets.
[0113] In this case, the information processing device 100 estimates the audio HuBERT features from the singing voice dataset, and uses the estimated features as a target to train a first model using the speech dataset. The information processing device 100 also estimates the timing of vocal pronunciation from the singing voice dataset, and uses the timing information together with the speech dataset to train the first model. This corresponds to the training in module M1.
[0114] In this case, the information processing device 100 uses the model N11 to estimate the timing of vocal pronunciation from the vocal dataset, and uses the timing information for learning.
[0115] (Learning of Module M2) The information processing device 100 acquires a dataset (corresponding to dataset L2) from the information providing device 200 (step S13).
[0116] The information processing device 100 uses the acquired data set to train a second model (step S14). This data set includes a singing voice data set. The information processing device 100 uses this singing voice data set to train a second model.
[0117] In this case, the information processing device 100 estimates the F0, volume, and voiced / unvoiced_flags of the audio from the singing voice dataset, and trains the second model using these as targets. The information processing device 100 also estimates the MIDI data of the audio, the features of the singer's singing voice, and the HuBERT features from the singing voice dataset, and trains the second model using these. This corresponds to the training in module M2.
[0118] At this time, the information processing device 100 uses the models N21 to N23 individually to estimate various information from the singing voice data set, and uses the various information for learning.
[0119] In addition, it is not necessary to include either or both of the volume and the voiced / unvoiced_flags in the learning process. Furthermore, it is also not necessary to include the features of the singer's singing voice in the learning process. If these features are not included in the learning process, there is no need to estimate this information.
[0120] (Learning of Module M3) The information processing device 100 acquires a dataset (corresponding to dataset L3) from the information providing device 200 (step S15).
[0121] The information processing device 100 uses the acquired data set to train a third model (step S16). This data set includes a singing voice data set. The information processing device 100 uses this singing voice data set to train a third model.
[0122] In this case, the information processing device 100 estimates the features, HuBERT features, F0, volume, and voiced / unvoiced_flags of the singer of the audio from the singing voice dataset, and uses these to train the third model with the audio as the target. This corresponds to the training in module M3.
[0123] At this time, the information processing device 100 uses the models N31 to N35 individually to estimate various information from the singing voice data set, and uses the various information for learning.
[0124] Furthermore, depending on the training of the second model, one or both of the volume and the voiced / unvoiced_flags may not be included in the training. For example, if the volume is not included in the training of the second model, the volume may also not be included in the training of the third model. Furthermore, depending on the training of the second model, the features of the singer's singing voice may also not be included in the training. For example, if the features of the singer's singing voice are not included in the training of the second model, the same may be true for the training of the third model. If they are not included in the training, estimation of that information also becomes unnecessary.
[0125] The information processing device 100 stores various trained models that have been trained in this manner for use in estimation.
[0126] (1-4-2. Estimation Phase of Information Processing Device) Next, the relationship between the information processing of the information processing device 100 and the user terminal 10 will be described using FIG. 8B. FIG. 8B is a sequence diagram (2) showing an overview of the information processing device according to the embodiment. In FIG. 8B, the information processing system 1 includes the information processing device 100 and the user terminal 10. This corresponds to the estimation phase.
[0127] (Trigger of estimation phase) The information processing device 100 acquires a request to execute singing voice synthesis from the user terminal 10 (step S21). For example, the information processing device 100 acquires the request to execute singing voice synthesis transmitted from the user terminal 10 when the user U1 operates the operation button B5 shown in Fig. 7A.
[0128] At this time, the information processing device 100 acquires information specified by the user U1 on screen G1 shown in Fig. 7A, etc. For example, the information processing device 100 acquires information about the lyrics specified by the user U1, information about the timing of the notes, MIDI data, and information about the audio. Of these, the audio is specified to synthesize the singing voice using the singer's singing voice, and is not necessary unless the user U1 specifically requests otherwise.
[0129] (Execution of Module M1) When the information processing device 100 receives a request to execute singing voice synthesis, it executes the estimation phase. First, the information processing device 100 inputs information about the lyrics specified by the user U1 and information about the timing of the notes into the first model (step S22). Then, the first model outputs HuBERT features.
[0130] (Execution of module M2) Then, the information processing device 100 inputs the MIDI data specified by the user U1, the vocal features of the singer identified based on the audio information, and the HuBERT features output from the first model into the second model (step S23). The second model then outputs the F0, volume, and voiced / unvoiced_flag of the vocal.
[0131] In this case, if user U1 has not specified audio, it is not necessary to input the features of the singer's singing voice. Furthermore, due to the training of the second model, there are cases where one or both of the volume and the voiced / unvoiced_flags are not output.
[0132] (Execution of module M3) Then, the information processing device 100 inputs the features of the singing voice of the singer identified based on the information about the audio specified by the user U1, the HuBERT features output from the first model, and the F0, volume, and voiced / unvoiced_flags output from the second model to the third model (step S24). The third model then outputs audio data of the singing voice.
[0133] In this case, as with module M2, if user U1 has not specified audio, there is no need to input the features of the singer's singing voice. Also, there is no need to input information that was not output from the second model.
[0134] (Providing singing voice) The information processing device 100 provides audio data of the singing voice output from the third model to the user U1 (step S25). For example, the information processing device 100 transmits information for displaying the screen G2 shown in Fig. 7B to the user terminal 10. The generated singing voice is played back when the user U1 operates, for example, the operation button B11 shown in Fig. 7B.
[0135] Furthermore, when the information processing device 100 receives an editing request from the user U1, the information processing device 100 may execute singing voice synthesis again based on the edited information. For example, if the user U1 edits any of the information designated by the user U1 after checking the generated singing voice, the information processing device 100 may transmit information to the user terminal 10 to enable singing voice synthesis to be executed again based on the edited information.
[0136] (1-5. Configuration of user terminal according to embodiment) Next, the configuration of the user terminal 10 according to the embodiment will be described using Fig. 9. Fig. 9 is a diagram showing an example configuration of the user terminal 10 according to the embodiment. As shown in Fig. 9, the user terminal has a communication unit 11, an input unit 12, an output unit 13, and a control unit 14.
[0137] The communication unit 11 is realized by, for example, a network interface card (NIC) or a network interface controller. The communication unit 11 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the information processing device 100, etc., via the network N. The network N is realized by, for example, a wireless communication standard or method such as Bluetooth (registered trademark), the Internet, Wi-Fi (registered trademark), UWB (Ultra Wide Band), or LPWA (Low Power Wide Area).
[0138] The input unit 12 accepts various operations from the user. In FIG. 8B , various operations are accepted from user U1. For example, the input unit 12 may accept various operations from the user via a display screen using a touch panel function. The input unit 12 may also accept various operations from buttons provided on the user terminal 10 or a keyboard or mouse connected to the user terminal 10.
[0139] The output unit 13 is a display screen of a tablet terminal or the like realized by, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display, and is a display device for displaying various information. For example, the output unit 13 displays information transmitted from the information processing device 100.
[0140] The control unit 14 is realized by, for example, a central processing unit (CPU), a micro processing unit (MPU), a graphics processing unit (GPU), etc. executing a program stored inside the user terminal 10 using a random access memory (RAM) etc. as a working area. The control unit 14 is also a controller, and may be realized by, for example, an integrated circuit such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a micro controller unit (MCU).
[0141] As shown in FIG. 9, the control unit 14 has a receiving unit 141 and a transmitting unit 142, and realizes or executes the information processing operations described below.
[0142] The receiving unit 141 receives various information from other information processing devices such as the information processing device 100. For example, the receiving unit 141 receives lyrics designated, input, or selected by a user or the like, information regarding the timing of the sounds, MIDI data, and audio data generated by singing voice synthesis based on the features of the singing voice of the singer of the audio. Also, for example, the receiving unit 141 receives information for displaying a UI screen including operation buttons for playing the singing voice thus generated.
[0143] The transmitting unit 142 transmits various information to another information processing device such as the information processing device 100. For example, the transmitting unit 142 transmits information about lyrics specified, input, or selected by a user, information about the timing of the sounds, MIDI data, and information about audio. Note that the MIDI data may be, for example, data created by a user typing a melody or the like into a keyboard.
[0144] Furthermore, for example, the transmitting unit 142 transmits operation information of a user or the like on a predetermined UI screen. For example, the transmitting unit 142 transmits a request to execute singing voice synthesis based on an operation of a user or the like on a predetermined UI screen.
[0145] Furthermore, for example, the transmitting unit 142 transmits editing information for information designated, input, or selected by a user or the like on a predetermined UI screen. Furthermore, for example, the transmitting unit 142 transmits a request to execute singing voice synthesis again based on the editing information for information designated, input, or selected by a user or the like on a predetermined UI screen.
[0146] (1-6. Configuration of Information Providing Device According to Embodiment) Next, the configuration of the information providing device 200 according to the embodiment will be described with reference to FIG. 10. FIG. 10 is a diagram showing an example of the configuration of the information providing device 200 according to the embodiment. As shown in FIG. 10, the information providing device 200 has a communication unit 210, a storage unit 220, and a control unit 230. Note that the information providing device 200 may also have an input unit (e.g., a keyboard or a mouse) that accepts various operations from an administrator of the information providing device 200, and a display unit (e.g., a liquid crystal display) that displays various information.
[0147] As described above, the information providing device 200 may be realized by being integrated with the information processing device 100 described later, or may be a part of the information processing device 100 .
[0148] The communication unit 110 is realized by, for example, a NIC, a network interface controller, etc. The communication unit 110 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the information processing device 100, etc. via the network N. The network N is realized by, for example, a wireless communication standard or method such as Bluetooth, the Internet, Wi-Fi, UWB, or LPWA.
[0149] The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk, an SSD (Solid State Drive), an optical disk, etc. As shown in Fig. 10 , the storage unit 120 has a singing voice data storage unit 221 and a speech data storage unit 222.
[0150] The singing voice data storage unit 221 stores information related to singing voice data, such as audio data of the singing voice and information related to the singer of the singing voice.
[0151] The speech data storage unit 222 stores information about speech data such as utterances, conversations, etc. The speech data storage unit 222 stores, for example, text information indicating the content of speech such as utterances and conversations, and voice information of speeches.
[0152] The control unit 230 is realized by, for example, a CPU, an MPU, a GPU, or the like executing a program stored inside the information providing device 200 using a RAM or the like as a work area. The control unit 230 is also a controller, and may be realized by, for example, an integrated circuit such as an ASIC, an FPGA, or an MCU.
[0153] 10, the control unit 230 has a receiving unit 231 and a transmitting unit 232, and realizes or executes the information processing action described below. Note that the internal configuration of the control unit 230 is not limited to the configuration shown in FIG. 10, and other configurations may be used as long as they perform the information processing described below.
[0154] The receiving unit 231 receives various information from other information processing devices such as the information processing device 100. For example, the receiving unit 231 receives a transmission request for singing voice data, speech data, or the like.
[0155] The transmitting unit 232 transmits various information to other information processing devices such as the information processing device 100. For example, the transmitting unit 232 transmits singing voice data, speech data, etc. based on a transmission request for singing voice data, speech data, etc.
[0156] (1-7. Configuration of information processing device according to embodiment) Next, the configuration of the information processing device 100 according to the embodiment will be described with reference to FIG. 11. FIG. 11 is a diagram showing an example of the configuration of the information processing device 100 according to the embodiment. As shown in FIG. 11, the information processing device 100 has a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing device 100 may have an input unit that accepts various operations from an administrator of the information processing device 100, and a display unit that displays various information.
[0157] The communication unit 110 is realized by, for example, a NIC, a network interface controller, etc. The communication unit 110 is connected to a network N by wire or wirelessly, and transmits and receives information to and from the user terminal 10, the information providing device 200, etc. via the network N. The network N is realized by, for example, a wireless communication standard or method such as Bluetooth, the Internet, Wi-Fi, UWB, or LPWA.
[0158] The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk, an SSD, an optical disk, etc. As shown in Fig. 11 , the storage unit 120 has a model storage unit 121 and a singing voice synthesis result storage unit 122.
[0159] The model storage unit 121 stores information about models to be applied in the estimation phase. For example, the model storage unit 121 stores information about a trained first model (corresponding to model N1) trained by module M1, information about a trained second model (corresponding to model N2) trained by module M2, and information about a trained third model (corresponding to model N3) trained by module M3.
[0160] Furthermore, for example, the model storage unit 121 stores information about the models (corresponding to models N11, N21 to N23, and N31 to N35) required to obtain input data when learning and estimating the first model, when learning and estimating the second model, and when learning and estimating the third model.
[0161] An example of the model storage unit 121 according to the embodiment is shown in Fig. 12. As shown in Fig. 12, the model storage unit 121 has items such as "model ID", "type", and "model".
[0162] "Model ID" indicates identification information for identifying a model. "Type" indicates what the model outputs. "Model" indicates a model. In the example shown in FIG. 12, conceptual information such as "Model #1" and "Model #2" is stored in "Model," but in reality, model parameters and the like are stored. For example, information contained in a dataset used to train the model is stored.
[0163] The singing synthesis result storage unit 122 stores information about the audio data of the singing voice generated by singing synthesis. Fig. 13 shows an example of the singing synthesis result storage unit 122 according to the embodiment. As shown in Fig. 13, the singing synthesis result storage unit 122 has fields such as "singing voice ID," "user ID," "input data," and "singing voice audio data."
[0164] "Singing voice ID" indicates identification information for identifying the audio data of the singing voice generated by singing voice synthesis. "User ID" indicates identification information for identifying the user who performed the singing voice synthesis. "Input data" indicates the input data applied in the estimation phase. In the example shown in Figure 13, conceptual information such as "input data #1" and "input data #2" is stored in "input data," but in reality, text information indicating the input data is stored.
[0165] "Singing voice audio data" refers to the audio data of the singing voice generated in the estimation phase. In the example shown in Fig. 13, conceptual information such as "singing voice audio data #1" and "singing voice audio data #2" is stored in "singing voice audio data," but in reality, frequency data of the singing voice audio data and the like is stored.
[0166] The control unit 130 is realized by, for example, a CPU, an MPU, a GPU, or the like executing a program (for example, an information processing program according to the present disclosure) stored inside the information processing device 100 using a RAM or the like as a work area. The control unit 130 is also a controller, and may be realized by, for example, an integrated circuit such as an ASIC, an FPGA, or an MCU.
[0167] 11 , the control unit 130 has an acquisition unit 131, a first learning unit 132, a second learning unit 133, a third learning unit 134, a first estimating unit 135, a second estimating unit 136, a generating unit 137, and a providing unit 138, and realizes or executes the information processing functions described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 11 , and may be any other configuration that performs the information processing described below.
[0168] The acquisition unit 131 acquires various types of information. For example, the acquisition unit 131 acquires information transmitted from the user terminal 10, the information providing device 200, or an external information processing device such as a server device or cloud system that provides or manages a specific service.
[0169] In the learning phase, the acquiring unit 131 acquires, for example, information related to singing voice data and speech data, such as information included in a singing voice data set and a speech data set.
[0170] In the learning phase, the acquiring unit 131 acquires, for example, information contained in a labeled dataset (corresponding to dataset L1) that includes singing voice data and speech data, and also acquires, for example, information contained in an unlabeled dataset (corresponding to datasets L2 and L3) that includes singing voice data (but does not include speech data).
[0171] In the learning phase, the acquisition unit 131 acquires, for example, information estimated based on information contained in these data sets.
[0172] In the learning phase, the acquisition unit 131 acquires, for example, information included in a data set corresponding to each module or information estimated based on information included in the data set.
[0173] In the estimation phase, the acquisition unit 131 acquires, for example, information designated, input, or selected by a user, etc. For example, the acquisition unit 131 acquires information about lyrics designated, input, or selected by a user, etc., information about the timing of the sounds, MIDI data, and information about audio.
[0174] In the estimation phase, the acquisition unit 131 acquires information that is estimated based on the information specified, input, or selected by the user, for example.
[0175] In the estimation phase, the acquisition unit 131 acquires, for example, the estimation result of each module. For example, for module M2, the acquisition unit 131 acquires the estimation result of module M1. For module M3, the acquisition unit 131 acquires the estimation results of modules M1 and M2.
[0176] The first training unit 132 trains a first model based on, for example, information included in the labeled dataset (corresponding to dataset L1) acquired by the acquisition unit 131. For example, the first training unit 132 trains the first model based on HuBERT features estimated from the audio of singing voice and speech (labels), using the HuBERT features as true HuBERT features (i.e., targets) so that the first model can estimate the HuBERT features when lyrics are input.
[0177] The second learning unit 133 trains a second model based on, for example, information included in the unlabeled dataset (corresponding to dataset L2) acquired by the acquisition unit 131. For example, the second learning unit 133 trains the second model based on MIDI data estimated from the audio of singing voice, features of the singer's singing voice, HuBERT features, F0, volume, and voiced / unvoiced_flags, so that the second model can estimate F0, volume, and voiced / unvoiced_flags from the MIDI data, features of the singer's singing voice, and HuBERT features.
[0178] The third learning unit 134 trains a third model based on, for example, information included in the unlabeled dataset (corresponding to dataset L3) acquired by the acquisition unit 131. For example, based on features of a singer's singing voice estimated from audio of the singing voice, HuBERT features, F0, volume, voiced / unvoiced_flags, and audio data of the singing voice, the third learning unit 134 trains the third model so that audio data of the singing voice can be generated from the features of the singer's singing voice, HuBERT features, F0, volume, and voiced / unvoiced_flags.
[0179] The first estimation unit 135 estimates the first feature amount from the information acquired by the acquisition unit 131 using, for example, the first model trained by the first learning unit 132 .
[0180] The first estimation unit 135 estimates the first feature from specified lyrics (such as lyrics specified, input, or selected by the user) using, for example, a trained first model that has been trained in advance using the first feature estimated from the singing voice as a target.
[0181] The first estimation unit 135 estimates the first feature amount from predetermined lyrics, for example, using a first model that has been trained using predetermined text information as labels.
[0182] The first estimation unit 135 estimates a first feature from predetermined lyrics, for example, using a first model that has been trained using timing information estimated from the singing voice together with predetermined text information as input data.
[0183] The second estimation unit 136 estimates the second feature from the information acquired by the acquisition unit 131 and the first feature estimated by the first estimation unit 135, for example, using a learned second model learned by the second learning unit 133.
[0184] The second estimation unit 136 estimates the second feature from predetermined MIDI data (such as MIDI data specified, input, or selected by the user) and the first feature estimated from predetermined lyrics, for example, using a trained second model that has been trained in advance using the second feature estimated from the singing voice as a target.
[0185] The second estimation unit 136 estimates the second feature from the specified MIDI data and the first feature estimated from the specified lyrics, for example, using a second model that has been trained using the MIDI data and the first feature estimated from the singing voice as input data.
[0186] The second estimation unit 136 estimates the second feature from the feature of a specified singer's singing voice, specified MIDI data, and the first feature estimated from specified lyrics, for example, using a second model that has been trained with the feature of a singer's singing voice estimated from the singing voice together with MIDI data estimated from the singing voice and the first feature as input data.
[0187] The generation unit 137 generates singing voice audio data from the information acquired by the acquisition unit 131, the first feature estimated by the first estimation unit 135, and the second feature estimated by the second estimation unit 136, for example, using a trained third model trained by the third learning unit 134.
[0188] The generation unit 137 generates a singing voice from a first feature estimated from specified lyrics and a second feature estimated from the specified lyrics and specified MIDI data, for example, using a pre-trained third model.
[0189] The generation unit 137 generates a singing voice from a first feature estimated from specified lyrics and a second feature estimated from the specified lyrics and specified MIDI data, for example, using a third model that has been trained using the first feature and second feature estimated from the singing voice as input data.
[0190] The generation unit 137 generates singing voice from the first feature estimated from specified lyrics and the second feature estimated from the specified lyrics and specified MIDI data, for example, using a third model that has been trained with the feature of the singer's singing voice estimated from the singing voice as input data together with the first feature and second feature estimated from the singing voice.
[0191] The generating unit 137 generates singing voices using, for example, a third model that uses a mel spectrogram of singing voices as output data, and a fourth model that converts the mel spectrogram into singing voices.
[0192] The generating unit 137 generates a singing voice to be provided to a user who desires singing voice synthesis by, for example, specifying, inputting or selecting predetermined lyrics and predetermined MIDI data.
[0193] The providing unit 138 provides, for example, information regarding the audio data of the singing voice generated by the generating unit 137. For example, the providing unit 138 provides information regarding the audio data of the singing voice to a user who has requested singing voice synthesis. For example, the providing unit 138 transmits information regarding the audio data of the singing voice to the user terminal 10 of the user who has requested singing voice synthesis. For example, the providing unit 138 transmits information for displaying a UI screen including operation buttons for playing the generated singing voice and operation buttons for downloading the generated singing voice.
[0194] The providing unit 138 may, for example, transmit the generated singing voice audio data directly to the user terminal 10 of the user who requested singing voice synthesis, or may transmit it to a server device or cloud system of a specific service that can be downloaded by the user.
[0195] Next, the processing of each unit constituting the information processing device 100 will be described in detail according to the flow with reference to Figures 14 to 17. Figures 14 to 16 show the flow of the learning processing in the control unit 130, and Figure 17 shows the flow of the estimation processing in the control unit 130.
[0196] 14 is a flowchart showing the flow of the learning process (module M1) in the control unit 130. The information processing device 100 acquires a labeled dataset of singing voice and speech (step S101).
[0197] The information processing device 100 estimates HuBERT features from the singing voice using the HuBERT model (step S102).
[0198] The information processing device 100 trains a first model using the speech (label) with the estimated HuBERT features as targets (step S103). Then, the information processing device 100 stores the trained first model.
[0199] 15 is a flowchart showing the flow of the learning process (module M2) in the control unit 130. The information processing device 100 acquires an unlabeled data set of singing voice (step S201).
[0200] The information processing device 100 estimates MIDI data, features of the singer's singing voice, HuBERT features, F0, volume, and voiced / unvoiced_flags from the singing voice using various models (step S202).
[0201] The information processing device 100 trains a second model so that F0, volume, and voiced / unvoiced_flags can be estimated from the MIDI data, the features of the singer's singing voice, and the HuBERT features (step S203).The information processing device 100 then stores the trained second model.
[0202] 16 is a flowchart showing the flow of the learning process (module M3) in the control unit 130. The information processing device 100 acquires an unlabeled data set of singing voice (step S301).
[0203] The information processing device 100 estimates the features of the singer's singing voice, the HuBERT features, F0, volume, and voiced / unvoiced_flags from the singing voice using various models (step S302).
[0204] The information processing device 100 trains a third model so that audio data of the singing voice can be generated from the features of the singer's singing voice, the HuBERT features, F0, volume, and voiced / unvoiced_flags (step S303).The information processing device 100 then stores the trained third model.
[0205] 17 is a flowchart showing the flow of the estimation process (modules M1 to M3) in the control unit 130. The information processing device 100 acquires information designated, input, or selected by a user or the like (step S401).
[0206] The information processing device 100 estimates a first feature amount using a first model based on the acquired information (step S402).
[0207] The information processing apparatus 100 estimates the second feature amount using the second model based on the acquired information and the estimated first feature amount (step S403).
[0208] The information processing device 100 generates audio data of the singing voice using the third model based on the acquired information, the estimated first feature amount, and the estimated second feature amount (step S404).
[0209] The information processing device 100 provides information about the generated audio data of the singing voice to the user who has requested singing voice synthesis (step S405).
[0210] (2. Other Embodiments) The processing according to each of the above-described embodiments may be implemented in various different forms other than the above-described embodiments.
[0211] Among the processes described in the above embodiments of the present disclosure, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the process procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the illustrated information.
[0212] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0213] Furthermore, the above-described embodiments of the present disclosure can be combined as appropriate within the scope of the processing content without causing inconsistencies. Furthermore, the order of the steps shown in the sequence diagrams or flowcharts of the present embodiments can be changed as appropriate. For example, the steps may be processed in chronological order, repeatedly, or partially in parallel.
[0214] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0215] (3. Hardware Configuration) The information processing device 100 and the like according to the embodiments of the present disclosure described above are realized by, for example, a computer 1000 configured as shown in FIG. 18 . The information processing device 100 will be described as an example. FIG. 18 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the information processing device 100. The computer 1000 has a processing circuitry 1100, a RAM 1200, a ROM 1300, a secondary storage device 1400, a communication interface 1500, an input / output interface 1600, a display unit 1700, a camera unit 1800, a microphone 1900, and a speaker 2000. The components of the computer 1000 are connected by a bus 1050.
[0216] The processing circuit 1100 operates and controls each unit based on programs stored in the ROM 1300 or the secondary storage device 1400. For example, the processing circuit 1100 loads the programs stored in the ROM 1300 or the secondary storage device 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0217] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the processing circuit 1100 when the computer 1000 is started up, and programs that depend on the hardware of the computer 1000 .
[0218] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the processing circuit 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records programs for each process of the information processing device 100 according to an embodiment of the present disclosure, which are examples of program data 1450.
[0219] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550. The communication interface 1500 corresponds to the communication unit 110 provided in the information processing device 100. For example, the processing circuit 1100 receives data from other devices and transmits data generated by the processing circuit 1100 to other devices via the communication interface 1500.
[0220] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the processing circuit 1100 receives data from an input device such as a microphone 1900 or a touch panel via the input / output interface 1600. The processing circuit 1100 also transmits data to an output device such as a display unit 1700 or a speaker 2000 via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of the media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.
[0221] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electroluminescence display (EL display). The display unit 1700 may also be a touch panel display device or a video projection device.
[0222] The camera unit 1800 is an interface through which the computer 1000 captures images. The microphone 1900 is an interface through which the computer 1000 captures audio. The speaker 2000 is an interface through which the computer 1000 outputs audio processed by the computer 1000. The components of the computer 1000 are connected by a bus 1050. The interfaces do not necessarily need to be provided inside the computer 1000, but may be provided outside the computer 1000 via a network or the like. Furthermore, the components constituting the computer 1000 may be controlled by a circuit different from the processing circuit 1100. For example, the display unit 1700 may be controlled not by the processing circuit 1100 but by a circuit dedicated to display processing provided in the display unit 1700.
[0223] For example, when the computer 1000 functions as the information processing device 100 according to an embodiment of the present disclosure, the processing circuit 1100 of the computer 1000 functions as the control unit 130 by executing a program loaded onto the RAM 1200. The secondary storage device 1400 stores the information processing program according to the present disclosure and various data stored in the storage device 120. The processing circuit 1100 reads and executes program data 1450 from the secondary storage device 1400. Alternatively, the processing circuit 1100 may obtain these programs from another device via an external network 1550. That is, the secondary storage device 1400 does not need to be located inside the computer 1000, but may also be located outside the computer 1000. The processing circuit 1100 is an example of an integrated circuit, and a CPU, an MPU, a GPU, an APU, an ASIC, and an FPGA can all be considered to be integrated circuits.
[0224] The present technology can also be configured as follows. (1) An information processing device comprising: a first estimation unit that estimates a first feature from lyrics; a second estimation unit that estimates a second feature from performance data that digitizes music performance information and the first feature; and a generation unit that generates singing voice based on the first feature and the second feature. (2) The information processing device described in (1), wherein the first estimation unit estimates the first feature from lyrics specified by a user using a trained first model that has been trained in advance using the first feature estimated from singing voice as a target. (3) The information processing device described in (2), wherein the first estimation unit estimates the first feature using the first model that has been trained using labels for the lyrics. (4) The information processing device described in (3), wherein the first estimation unit estimates the first feature using the first model that has been trained including timing information estimated from singing voice. (5) The information processing device described in any one of (1) to (4), wherein the second estimation unit estimates the second feature from the performance data specified by a user and the first feature estimated from lyrics specified by the user, using a trained second model that has been trained in advance with the second feature estimated from singing voice as a target. (6) The information processing device described in (5), wherein the second estimation unit estimates the second feature using the second model that has been trained using the performance data estimated from singing voice and the first feature. (7) The information processing device described in (6), wherein the second estimation unit estimates the second feature using the second model that has been trained including feature of the singer's singing voice estimated from the singing voice. (8) The information processing device described in any one of (1) to (7), wherein the generation unit generates singing voice from the first feature estimated from lyrics specified by a user and the second feature estimated from the lyrics specified by the user and the performance data, using a trained third model that has been trained in advance. (9) The information processing device according to (8), wherein the generation unit generates singing voice using the third model trained using the first feature amount and the second feature amount estimated from singing voice.(10) The information processing device according to (9), wherein the generation unit generates a singing voice using the third model trained including features of the singer's singing voice estimated from the singing voice. (11) The information processing device according to any one of (8) to (10), wherein the generation unit generates a singing voice using the third model whose output data is a mel spectrogram of the singing voice and a fourth model that converts the mel spectrogram into a singing voice. (12) The information processing device according to any one of (8) to (11), wherein the generation unit generates a singing voice to be provided to the user who requests singing voice synthesis. (13) The information processing device according to any one of (1) to (12), wherein the first features are HuBERT features. (14) The information processing device according to any one of (1) to (13), wherein the second features are features including at least F0, which represents the frequency of a sound. (15) An information processing method executed by a computer, comprising: a first estimating step of estimating a first feature quantity from lyrics, a second estimating step of estimating a second feature quantity from performance data obtained by digitizing music performance information and the first feature quantity, and a generating step of generating a singing voice based on the first feature quantity and the second feature quantity. (16) An information processing program for causing a computer to execute: a first estimating procedure of estimating a first feature quantity from lyrics, a second estimating procedure of estimating a second feature quantity from performance data obtained by digitizing music performance information and the first feature quantity, and a generating procedure of generating a singing voice based on the first feature quantity and the second feature quantity.
[0225] REFERENCE SIGNS LIST 1 Information processing system 10 User terminal 11 Communication unit 12 Input unit 13 Output unit 14 Control unit 100 Information processing device 110 Communication unit 120 Memory unit 121 Model memory unit 122 Singing voice synthesis result memory unit 130 Control unit 131 Acquisition unit 132 First learning unit 133 Second learning unit 134 Third learning unit 135 First estimation unit 136 Second estimation unit 137 Generation unit 138 Provision unit 141 Reception unit 142 Transmission unit 200 Information providing device 210 Communication unit 220 Memory unit 221 Singing voice data memory unit 222 Speech data memory unit 230 Control unit 231 Reception unit 232 Transmission unit N Network
Claims
1. An information processing device comprising: a first estimation unit that estimates a first feature from lyrics; a second estimation unit that estimates a second feature from performance data that has been digitized from music performance information and the first feature; and a generation unit that generates a singing voice based on the first feature and the second feature.
2. The information processing device described in claim 1, wherein the first estimation unit estimates the first feature from lyrics specified by a user using a trained first model that has been trained in advance using the first feature estimated from singing voice as a target.
3. The information processing device according to claim 2, wherein the first estimation unit estimates the first feature amount using the first model trained using labels for the lyrics.
4. The information processing device according to claim 3, wherein the first estimation unit estimates the first feature using the first model trained to include timing information estimated from singing voice.
5. The information processing device described in claim 1, wherein the second estimation unit estimates the second feature from the performance data specified by the user and the first feature estimated from lyrics specified by the user, using a trained second model that has been trained in advance using the second feature estimated from singing voice as a target.
6. The information processing device according to claim 5, wherein the second estimation unit estimates the second feature using the second model trained using the performance data estimated from the singing voice and the first feature.
7. The information processing device according to claim 6, wherein the second estimation unit estimates the second feature using the second model trained to include features of the singer's singing voice estimated from the singing voice.
8. The information processing device described in claim 1, wherein the generation unit uses a pre-trained third model to generate singing voices from the first feature estimated from lyrics specified by a user and the second feature estimated from the lyrics specified by the user and the performance data.
9. The information processing device according to claim 8, wherein the generation unit generates singing voice using the third model trained using the first feature and the second feature estimated from singing voice.
10. The information processing device according to claim 9, wherein the generation unit generates the singing voice using the third model that has been trained to include features of the singer's singing voice estimated from the singing voice.
11. The information processing device according to claim 8, wherein the generation unit generates singing voice using the third model, which uses a mel spectrogram of singing voice as output data, and a fourth model, which converts the mel spectrogram into singing voice.
12. The information processing device according to claim 8, wherein the generation unit generates a singing voice to be provided to the user who requests singing voice synthesis.
13. The information processing device according to claim 1, wherein the first feature is a HuBERT feature.
14. The information processing device according to claim 1, wherein the second feature is a feature including at least F0 representing the frequency of a sound.
15. An information processing method executed by a computer, comprising: a first estimation step of estimating a first feature quantity from lyrics; a second estimation step of estimating a second feature quantity from performance data that has been converted into digital form information about musical performance and the first feature quantity; and a generation step of generating a singing voice based on the first feature quantity and the second feature quantity.
16. An information processing program for causing a computer to execute the following steps: a first estimation procedure for estimating a first feature quantity from lyrics; a second estimation procedure for estimating a second feature quantity from performance data that has been converted into digital form information about musical performance and the first feature quantity; and a generation procedure for generating a singing voice based on the first feature quantity and the second feature quantity.
Citation Information
Patent Citations
Generating device for singing voice synthetic data
JP1991007996A
Singing voice synthesizing device
JP1994337690A