Learning device, conversion device, learning method, conversion method, and program
The system improves voice conversion accuracy by training a mathematical model with speaker and F0 labels, enhancing mel spectrogram estimation and speech conversion.
Patent Information
- Application Number
- JP2024548851
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2026-01-07
- Estimated Expiration
- 2042-09-27
AI Technical Summary
Existing voice conversion techniques do not provide sufficient accuracy in converting voice signals.
A system that processes speech signals using a mathematical model trained with speaker labels, mel spectrogram information, and F0 labels, incorporating processes for speaker and F0 label generation, encoding, and decoding to improve accuracy.
Enhances the accuracy of voice conversion by refining the mel spectrogram estimation, leading to improved speech conversion results.
Smart Images

Figure 0007795138000011 
Figure 0007795138000012 
Figure 0007795138000013
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning device, a conversion device, a learning method, a conversion method, and a program. [Background technology]
[0002] Voice conversion technology is a method for converting a voice signal into a signal with different properties in terms of certain characteristics. For example, one such voice conversion method is a method for converting a voice signal from one speaker into a voice signal from a different speaker. In recent years, speaker conversion methods based on a deep learning framework have been proposed (Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, Nobukatsu Hojo. “ACVAE-VC: Non-Parallel Voice Conversion With Auxiliary Classifier Variational Autoencoder.” IEEE / ACM Transactions on Audio, Speech and Language Processing. Vol. 27, No. 9. pp. 1432-1443. 2019 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the techniques proposed so far have not always provided good speech conversion accuracy.
[0005] In view of the above circumstances, an object of the present invention is to provide a technique for improving the accuracy of voice conversion. [Means for solving the problem]
[0006] One aspect of the present invention is a system for processing a speech signal including a control unit that performs training of a mathematical model of a training target based on a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of a speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech, wherein the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model, and the mathematical model performs a speaker label generation process that generates a speaker label indicating a speaker determined based on a predetermined rule, an F0 label generation process that generates an F0 label indicating an F0 pattern determined based on the predetermined rule, an encoding process that extracts features of the mel spectrogram, and an encoding process that extracts features of the mel spectrogram and an F0 label indicating an F0 pattern of the speech. and a decoding process which is a process for a decoding target and which estimates a mel spectrogram indicated by the decoding target by decoding the decoding target, wherein the mathematical model is executed in the learning, and the decoding process is executed in the learning for a first decoding target which is a decoding target including the features, the speaker label, and the F0 label input to the mathematical model, and a second decoding target which is a decoding target including the features, a result of the speaker label generation process, and a result of the F0 label generation process, and the control unit updates the mathematical model based on the result of execution of the mathematical model.
[0007] One aspect of the present invention provides a system including an acquisition process for acquiring a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech; and a control unit for training a mathematical model of a training object based on the speaker label, the mel spectrogram information, and the F0 label, wherein the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model, and the mathematical model includes a speaker label generation process for generating a speaker label indicating a speaker determined based on a predetermined rule, an F0 label generation process for generating an F0 label indicating an F0 pattern determined based on the predetermined rule, an encoding process for extracting features of the mel spectrogram, and a decoding process for processing a decoding object that is a processing object including at least the features and that estimates a mel spectrogram indicated by the decoding object by decoding the decoding object. and a decoding process for decoding a mel spectrogram indicated by the mel spectrogram information obtained in the acquisition process into a mel spectrogram in which the speaker is identified by the speaker label obtained in the acquisition process and the F0 label is identified by the F0 label obtained in the acquisition process, the mel spectrogram being a mel spectrogram in which the speaker is identified by the speaker label obtained in the acquisition process and the F0 label is identified by the F0 label obtained in the acquisition process, the control unit updates the mathematical model based on the result of execution of the mathematical model.
[0008] One aspect of the present invention includes a control step of training a mathematical model of a training object based on a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of a speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech, wherein the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model, and the mathematical model includes a speaker label generation process of generating a speaker label indicating a speaker determined based on a predetermined rule, an F0 label generation process of generating an F0 label indicating an F0 pattern determined based on the predetermined rule, an encoding process of extracting a feature of the mel spectrogram, and an encoding process of extracting a feature of the mel spectrogram from a processing object including at least the feature. a decoding process which is a process for a decoding target and which estimates a mel spectrogram indicated by the decoding target by decoding the decoding target, wherein the mathematical model is executed in the learning, and the decoding process is executed in the learning for a first decoding target which is a decoding target including the features, the speaker label, and the F0 label input to the mathematical model, and a second decoding target which is a decoding target including the features, a result of the speaker label generation process, and a result of the F0 label generation process, and the control step updates the mathematical model based on the result of execution of the mathematical model.
[0009] One aspect of the present invention includes an acquisition process for acquiring a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech; and a control step for training a mathematical model of a training object based on the speaker label, the mel spectrogram information, and the F0 label, wherein the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model, and the mathematical model includes a speaker label generation process for generating a speaker label indicating a speaker determined based on a predetermined rule, an F0 label generation process for generating an F0 label indicating an F0 pattern determined based on the predetermined rule, an encoding process for extracting features of the mel spectrogram, and a decoding process for processing a decoding object that is a processing object including at least the features and that estimates a mel spectrogram indicated by the decoding object by decoding the decoding object. and a conversion control step of executing a conversion process using the encoding process and the decoding process in the trained mathematical model obtained by the training method, wherein the mathematical model is executed in the training, and the decoding process is executed on a first decoding target that is a decoding target including the features, the speaker label and the F0 label input to the mathematical model, and a second decoding target that is a decoding target including the features, a result of the speaker label generation process, and a result of the F0 label generation process, and the control step updates the mathematical model based on the result of execution of the mathematical model.
[0010] One aspect of the present invention is a program for causing a computer to function as either the learning device or the conversion device described above. [Effects of the Invention]
[0011] The present invention makes it possible to improve the accuracy of voice conversion. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a conversion system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of a learning device according to an embodiment. [Figure 3] 10 is a flowchart showing an example of a flow of processing executed by a learning device according to an embodiment. [Figure 4] FIG. 2 is a diagram illustrating an example of a hardware configuration of a conversion device according to an embodiment. [Figure 5] 10 is a flowchart showing an example of a flow of processing executed by a conversion device according to an embodiment. [Figure 6] FIG. 10 is a diagram showing an example of a result of a first test in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] (Embodiment) FIG. 1 is a diagram illustrating an example of the configuration of a conversion system 100 according to an embodiment. The conversion system 100 includes a learning device 1 and a conversion device 2. The learning device 1 updates a mathematical model to be learned (hereinafter referred to as a "learning target model") through learning. Specifically, the learning target model is a mathematical model including a mathematical model used for voice conversion.
[0014] Note that the voice conversion in the conversion system 100 specifically refers to the conversion of a mel spectrogram. Based on the converted mel spectrogram, a vocoder can be used to estimate the waveform of the converted voice. Therefore, the higher the accuracy of the mel spectrogram conversion, the higher the accuracy of the voice conversion. Therefore, improving the accuracy of the mel spectrogram conversion leads to improved accuracy of the voice conversion. Furthermore, in the conversion system 100, the mathematical model used for voice conversion is, more specifically, the mathematical model used for the conversion of a mel spectrogram.
[0015] The learning device 1 continues learning of the model to be trained until a predetermined termination condition for learning (hereinafter referred to as the "learning termination condition") is satisfied. The learning termination condition may be, for example, a condition that the number of updates of the model to be trained reaches a predetermined number, or a condition that the change in the model to be trained due to the update is smaller than a predetermined change.
[0016] The conversion device 2 converts the mel spectrogram using part of the processing included in the trained learning model. "Trained" means that the learning termination condition has been met. Therefore, a trained mathematical model means a mathematical model at the time that the learning termination condition has been met.
[0017] <About Learning Device 1> The learning device 1 executes a learning process. The learning process is a process in which a pair of a speaker label, mel spectrogram information, and an F0 label is used as a processing target, and a learning target model is learned based on the processing target until a learning termination condition is met. The speaker label is information indicating a speaker. The mel spectrogram information is information indicating a mel spectrogram of a speech uttered by a speaker. The mel spectrogram indicated by the mel spectrogram information may be a logarithmic mel spectrogram. The F0 label is information indicating the F0 pattern of a speech uttered by a speaker.
[0018] In the training process, a training target model is executed. The speaker label, mel spectrogram information, and F0 label, which are the processing targets of the training process, are input to the training target model. Hereinafter, the speaker label input to the training target model will be referred to as the first speaker label. Hereinafter, the F0 label input to the training target model will be referred to as the first F0 label.
[0019] The training model includes a speaker label generation process, an F0 label generation process, an encoding process, and a decoding process. When the training model is executed, the speaker label generation process, the F0 label generation process, the encoding process, and the decoding process are executed. Therefore, the training process executes the speaker label generation process, the F0 label generation process, the encoding process, and the decoding process.
[0020] The speaker label generation process is a process for generating a speaker label indicating a speaker determined based on a predetermined rule. Hereinafter, the result of the speaker label generation process will be referred to as a second speaker label. That is, the second speaker label is a speaker label indicating a speaker determined based on a predetermined rule by executing the speaker label generation process. The predetermined rule in the speaker label generation process is, for example, a rule that a speaker is randomly determined from among speaker candidates prepared in advance.
[0021] The F0 label generation process is a process for generating an F0 label that indicates an F0 pattern determined based on a predetermined rule. Hereinafter, the result of the F0 label generation process will be referred to as a second F0 label. In other words, the second F0 label is an F0 label that indicates an F0 pattern determined based on a predetermined rule by executing the F0 label generation process.
[0022] The predetermined rule in the F0 label generation process is, for example, a rule that the F0 label generated in the F0 label generation process is an F0 label randomly determined from F0 label candidates prepared in advance. The predetermined rule in the F0 label generation process is, for example, a rule that the F0 label generated in the F0 label generation process indicates the result of adding random noise to the F0 pattern indicated by the first F0 label.
[0023] The encoding process is a process for extracting the features of the mel spectrogram indicated by the mel spectrogram information input to the learning target model. The decoding process is a process for the decoding target, which is the processing target, and is a process for estimating the mel spectrogram indicated by the decoding target by decoding the decoding target. The decoding target is information that includes at least the features obtained by the encoding process (hereinafter referred to as "mel features").
[0024] In the training process, the decoding process is performed on multiple decoding targets. One of the decoding targets includes Mel features, a first speaker label, and a first F0 label. Hereinafter, the decoding target including Mel features, a first speaker label, and a first F0 label will be referred to as the first decoding target.
[0025] The other decoding target includes Mel features, a second speaker label, and a second F0 label. Hereinafter, the decoding target including Mel features, a second speaker label, and a second F0 label will be referred to as a second decoding target.
[0026] In the learning process, a process of updating the learning object model (hereinafter referred to as "update process") is performed based on the results of execution of the learning object model. The learning object model is updated so as to satisfy at least the first update condition and the second update condition.
[0027] The first update condition is a condition that the learning object model is updated so that the difference between the result of the decoding process on the first decoding object and the mel spectrogram indicated by the mel spectrogram information input to the learning object model becomes small.
[0028] The second update condition is a condition that, when the difference condition is satisfied, the training target model is updated so that the result of the decoding process for the first decoding target differs from the result of the decoding process for the second decoding target. The difference condition is a condition that at least one of the following conditions is satisfied: the first speaker label and the second speaker label are different; and the first F0 label and the second F0 label are different.
[0029] The encoding process and the decoding process are processes in which the contents are updated by the update process.
[0030] <First example of update processing> A first example of the update process will be described. The learning target model further includes a speaker label estimation process and an F0 label estimation process. The speaker label estimation process is a process in which the processing target is the result of a decoding process and estimates a speaker label based on the processing target. The F0 label estimation process is a process in which the processing target is the result of a decoding process and estimates an F0 label based on the processing target.
[0031] In a first example of the update process, a speaker label estimation process is performed on the second decoding result, and an F0 label estimation process is performed on the second decoding result. The second decoding result is the result of the decoding process on the second decoding target.
[0032] In a first example of the update process, the training model is updated based on a loss function including at least the reconstruction error, the second speaker label loss function, and the second F0 label loss function.
[0033] The reconstruction error indicates the difference between the first decoding result and the mel spectrogram information input to the learning target model. The first decoding result is the result of the decoding process for the first decoding target.
[0034] The second speaker label loss function indicates the difference between the result of the speaker label estimation process using the second decoding result as the processing target and the second speaker label. The second F0 label loss function indicates the difference between the result of the F0 label estimation process using the second decoding result as the processing target and the second F0 label.
[0035] In the first example of the update process, the training model is updated so that the difference between the reconstruction error, the difference between the second speaker label loss function, and the difference between the second F0 label loss function becomes smaller. In this way, the first example of the update process performs an update that satisfies at least the first and second update conditions.
[0036] <Second example of update processing> A second example of the update process will be described. As in the first example of the update process, the training model further includes a speaker label estimation process and an F0 label estimation process. In the second example of the update process, a speaker label estimation process is performed on the second decoding result, and an F0 label estimation process is performed on the second decoding result.
[0037] In the second example of the update process, a speaker label estimation process is further performed in which the first decoding result is processed, and a speaker label estimation process is performed in which the mel spectrogram information input to the training model is processed.
[0038] In a second example of the update process, the training model is updated based on a loss function including at least the reconstruction error, the second speaker label loss function, the second F0 label loss function, and the first speaker label loss function.
[0039] The first speaker label loss function is the sum of the first sub-difference and the second sub-difference. The first sub-difference is the difference between the result of the speaker label estimation process that processes the first decoding result and the first speaker label. The second sub-difference is the difference between the result of the speaker label estimation process that processes the mel spectrogram information input to the training model and the first speaker label.
[0040] In the second example of the update process, the training model is updated so that the difference between the reconstruction error, the difference between the second speaker label loss function, the difference between the second F0 label loss function, and the difference between the first speaker label loss function becomes smaller. In this way, the second example of the update process performs an update that satisfies at least the first and second update conditions.
[0041] <Third example of update processing> A third example of the update process will be described. As in the first example of the update process, the training model further includes a speaker label estimation process and an F0 label estimation process. In the third example of the update process, a speaker label estimation process is performed on the second decoding result, and an F0 label estimation process is performed on the second decoding result.
[0042] In the third example of the update process, an F0 label estimation process is further performed, in which the first decoding result is processed, and an F0 label estimation process is performed, in which the mel spectrogram information input to the model to be trained is processed.
[0043] In a third example of the update process, the training model is updated based on a loss function including at least the reconstruction error, the second speaker label loss function, the second F0 label loss function, and the first F0 label loss function.
[0044] The first F0 label loss function is the sum of the third and fourth sub-differences. The third sub-difference is the difference between the first F0 label and the result of the F0 label estimation process that processes the first decoding result. The fourth sub-difference is the difference between the first F0 label and the result of the F0 label estimation process that processes the mel spectrogram information input to the model being trained.
[0045] In the third example of the update process, the training model is updated so that the difference between the reconstruction error, the difference between the second speaker label loss function, the difference between the second F0 label loss function, and the difference between the first F0 label loss function becomes smaller. In this way, the third example of the update process performs an update that satisfies at least the first and second update conditions.
[0046] <Fourth example of update processing> A fourth example of the update process will be described. As in the first example of the update process, the training model further includes a speaker label estimation process and an F0 label estimation process. In the fourth example of the update process, a speaker label estimation process is performed on the second decoding result, and an F0 label estimation process is performed on the second decoding result.
[0047] In a fourth example of the update process, a speaker label estimation process is further performed using the first decoding result, and a speaker label estimation process is further performed using the mel spectrogram information input to the training model.In the fourth example of the update process, an F0 label estimation process is further performed using the first decoding result, and a speaker label estimation process is further performed using the mel spectrogram information input to the training model.
[0048] In a fourth example of the update process, the training model is updated based on a loss function including at least the reconstruction error, the second speaker label loss function, the second F0 label loss function, the first speaker label loss function, and the first F0 label loss function.
[0049] In the fourth example of the update process, the training model is updated so that the difference between the reconstruction error, the difference between the second speaker label loss function, the difference between the second F0 label loss function, the difference between the first speaker label loss function, and the difference between the first F0 label loss function becomes smaller. In this way, the fourth example of the update process performs an update that satisfies at least the first and second update conditions.
[0050] The speaker label estimation process and the F0 label estimation process may be included in the training model. In this case, the contents of the speaker label estimation process and the F0 label estimation process may be updated when the training model is updated by the update process.
[0051] In this way, in all of the first to fourth examples of the update process, updating is performed using a loss function including the reconstruction error, the second speaker label loss function, and the second F0 label loss function.
[0052] <About Converter 2> As described above, the conversion device 2 performs conversion of the mel spectrogram using part of the processing included in the trained training model. Specifically, part of the processing included in the trained training model used by the conversion device 2 is encoding processing and decoding processing in the trained training model. Hereinafter, the encoding processing in the trained training model will be referred to as trained encoding processing. Hereinafter, the decoding processing in the trained training model will be referred to as trained decoding processing.
[0053] The conversion device 2 executes a conversion process. The conversion process is a process of converting a source mel spectrogram into a destination mel spectrogram based on the source mel spectrogram, a speaker label, and an F0 label. The destination mel spectrogram is converted into a mel spectrogram whose F0 pattern is the F0 pattern indicated by the F0 label and whose speaker is the speaker indicated by the speaker label. The source mel spectrogram is the source mel spectrogram.
[0054] The speaker of a mel spectrogram refers to the speaker who uttered the voice indicated by the mel spectrogram.
[0055] In the conversion process, a trained encoding process is performed on the source mel spectrogram. In the conversion process, a trained decoding process is performed on a decoding target including the result of the trained encoding process on the source mel spectrogram, the speaker label, and the F0 label. The result of the trained decoding process is a converted mel spectrogram obtained by the conversion process.
[0056] 2 is a diagram showing an example of the hardware configuration of a learning device 1 according to an embodiment. The learning device 1 includes a control unit 11 having a processor 91, such as a CPU (Central Processing Unit), and a memory 92 connected via a bus, and executes a program. By executing the program, the learning device 1 functions as a device including the control unit 11, an input / output interface 12, and a memory unit 13.
[0057] More specifically, the processor 91 reads out a program stored in the storage unit 13 and stores the read out program in the memory 92. When the processor 91 executes the program stored in the memory 92, the learning device 1 functions as a device including the control unit 11, the input / output interface 12, and the storage unit 13.
[0058] The control unit 11 controls the operation of various functional units included in the learning device 1. The control unit 11 executes, for example, a learning process. The control unit 11 controls, for example, the operation of the input / output interface 12 to acquire information input to the input / output interface 12. The control unit 11 controls, for example, the operation of the input / output interface 12 to output the results of the learning process to a predetermined output destination. The results of the learning process that the control unit 11 controls the operation of the input / output interface 12 to output to a predetermined output destination are, for example, a computer program that indicates part of the processing included in the trained learning model and that is used to convert a mel spectrogram by the conversion device 2.
[0059] Input / output interface 12 inputs and outputs various types of information. Input / output interface 12 is an interface that outputs information and is configured to include a display device such as a CRT (Cathode Ray Tube) display, a liquid crystal display, or an organic EL (Electro-Luminescence) display. The interface that outputs information may be configured as an interface that connects these display devices to learning device 1.
[0060] Input / output interface 12 is an interface for receiving information input and is configured to include input devices such as a mouse, keyboard, touch panel, etc. An interface for outputting information may be configured as an interface for connecting these input devices to study device 1.
[0061] The input / output interface 12 also includes a communication device. The input / output interface 12 communicates with an external device via wired or wireless communication using the communication device. The external device is, for example, the conversion device 2. The input / output interface 12 transmits, via communication using the communication device, to the conversion device 2 a computer program that indicates part of the processing included in the trained model to be trained and that is used to convert a mel spectrogram by the conversion device 2. Therefore, the conversion device 2 is an example of a predetermined output destination.
[0062] The external device may be, for example, an external storage device such as a hard disk drive or a solid state drive (SSD). A computer program that indicates part of the processing included in the trained learning model and that is used to convert a mel spectrogram by the conversion device 2 may be output to the external storage device. Therefore, the external storage device is an example of a predetermined output destination.
[0063] The input / output interface 12 receives a processing target for the learning process via an input device or a communication device.
[0064] The storage unit 13 is configured using a computer-readable storage medium (non-transitory computer-readable recording medium) such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 13 stores various information related to the learning device 1. The storage unit 13 stores, for example, the results of processing executed by the control unit 11. Therefore, the storage unit 13 stores, for example, the results of the learning processing. The storage unit 13 stores, for example, the processing target of the learning processing.
[0065] FIG. 3 is a flowchart showing an example of the flow of processing executed by the learning device 1 in the embodiment.
[0066] The control unit 11 inputs the first speaker label, mel spectrogram information, and first F0 label to the training model (step S101). That is, the first speaker label, mel spectrogram information, and first F0 label are input to the training model. The control unit 11 acquires the first speaker label, mel spectrogram information, and first F0 label via, for example, the input / output interface 12.
[0067] Next, the control unit 11 executes a speaker label generation process (step S102). By executing the process of step S102, a second speaker label is obtained. Next, the control unit 11 executes an F0 label generation process (step S103). By executing the process of step S103, a second F0 label is obtained. Next, the control unit 11 executes an encoding process (step S104) for the mel spectrogram information acquired in step S101 as the processing target. By executing step S104, mel features are obtained.
[0068] Next, the control unit 11 executes a decoding process on a decoding target (i.e., a first decoding target) including the Mel feature obtained in step S104 and the first speaker label and first F0 label obtained in step S101 (step S105). Next, the control unit 11 executes a decoding process on a decoding target (i.e., a second decoding target) including the Mel feature obtained in step S104, the second speaker label obtained in step S102, and the second F0 label obtained in step S103 (step S106).
[0069] Next, the control unit 11 updates the learning target model based on the results of the processing in step S105 and the results of the processing in step S106 so as to satisfy at least the first update condition and the second update condition (step S107). Next, the control unit 11 determines whether the learning termination condition is satisfied (step S108).
[0070] If the learning termination condition is not satisfied (step S108: NO), the process returns to step S101. On the other hand, if the learning termination condition is satisfied (step S108: YES), the process ends. The model to be trained at the time the process ends is the trained model.
[0071] The process of step S105 may be executed at any timing as long as it is before the process of step S107 and after the process of step S104. The process of step S106 may be executed at any timing as long as it is after the process of steps S102, S103, and S104 and before the process of step S107.
[0072] Furthermore, the process of step S104 may be executed at any timing after execution of step S101 and before execution of the process of step S105 or step S106, whichever is executed first. The process of step S103 may be executed at any timing after execution of step S101 and before execution of step S106.
[0073] The process of step S102 may be executed at any timing after the process of step S101 and before the process of step S106. In this way, the processes of steps S101 to S108 may be executed at any timing as long as they do not violate the law of causality.
[0074] 4 is a diagram showing an example of the hardware configuration of the conversion device 2 according to the embodiment. The conversion device 2 includes a control unit 21 having a processor 93 such as a CPU and a memory 94 connected via a bus, and executes a program. By executing the program, the conversion device 2 functions as a device including the control unit 21, an input / output interface 22, and a storage unit 23.
[0075] More specifically, the processor 93 reads the program stored in the storage unit 23 and stores the read program in the memory 94. The processor 93 executes the program stored in the memory 94, causing the conversion device 2 to function as a device including the control unit 21, the input / output interface 22, and the storage unit 23.
[0076] The control unit 21 controls the operation of various functional units included in the conversion device 2. The control unit 21 executes, for example, a conversion process. The control unit 21 acquires information input to the input / output interface 22. The control unit 21 controls, for example, the operation of the input / output interface 22 to output the results of the conversion process.
[0077] The input / output interface 22 inputs and outputs various types of information. The input / output interface 22 is an interface that outputs information and includes a display device such as a CRT display, a liquid crystal display, or an organic EL display. The interface that outputs information may be configured as an interface that connects these display devices to the conversion device 2.
[0078] The input / output interface 22 is an interface for receiving input of information and is configured to include input devices such as a mouse, keyboard, touch panel, etc. An interface for outputting information may be configured as an interface for connecting these input devices to the conversion device 2.
[0079] The input / output interface 22 also includes a communication device. The input / output interface 22 communicates with an external device via wired or wireless communication using the communication device. The external device is, for example, the learning device 1. The input / output interface 22 acquires, via communication using the communication device, a computer program from the learning device 1 that indicates part of the processing included in the trained learning object model and that is used for converting a mel spectrogram by the conversion device 2.
[0080] The external device may be, for example, an external storage device such as a hard disk drive or a solid state drive (SSD). A computer program that indicates part of the processing included in the trained learning model and that is used for converting a mel spectrogram by the conversion device 2 may be read from the external storage device.
[0081] A source mel spectrogram, a speaker label, and an F0 label are input to the input / output interface 22 via an input device or a communication device. The control unit 21 acquires the source mel spectrogram, the speaker label, and the F0 label input to the input / output interface 22.
[0082] The storage unit 23 is configured using a computer-readable storage medium device (non-transitory computer-readable recording medium) such as a magnetic hard disk device or a semiconductor storage device. The storage unit 23 stores various information related to the conversion device 2. The storage unit 23 stores, for example, the results of processing executed by the control unit 21. Therefore, the storage unit 23 stores, for example, the results of conversion processing. The storage unit 23 stores, for example, the processing target of the conversion processing.
[0083] FIG. 5 is a flowchart showing an example of the flow of processing executed by the conversion device 2 in the embodiment.
[0084] The control unit 21 executes an acquisition process (step S201). The acquisition process is a process for acquiring a source mel spectrogram, a speaker label, and an F0 label. The source mel spectrogram, the speaker label, and the F0 label are input to the input / output interface 22 by, for example, a user. In this case, the control unit 21 executes the acquisition process to acquire the source mel spectrogram, the speaker label, and the F0 label input to the input / output interface 22.
[0085] Next, control unit 21 executes a conversion process based on the information obtained in step S201 (step S202). By executing the conversion process, a mel spectrogram is obtained based on the source mel spectrogram, the F0 pattern of which is the F0 pattern indicated by the F0 label and the speaker is the speaker indicated by the speaker label. Next, control unit 21 controls the operation of input / output interface 22 to cause input / output interface 22 to output the mel spectrogram obtained in step S202 (step S203).
[0086] <Experimental Results> An example of experimental results using the conversion system 100 is shown below. In the experiment, a mini-batch consisting of 16 samples was used in each learning step, and the parameters of the entire neural network model were updated using the backpropagation algorithm. In the experiment, a set of mel-spectrogram information indicating logarithmic mel-spectrograms defined from the speech of four speakers, F0 labels, and speaker labels was used as one sample. In the experiment, training was performed for 5,000 epochs. Parallel WaveGAN (see Reference 1) was used as a vocoder to convert the test data.
[0087] Reference 1: R. Yamamoto, E. Song, and JM. Kim. Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In ICASSP, pages 6199-6203, 2020.
[0088] The experiments were conducted using the CMU Arctic speech database (see Reference 2) consisting of the above four speakers.
[0089] Reference 2: J. Kominek and A.W. Black. The CMU Arctic speech databases. In SSW5-2004, pages 223-224, 2004.
[0090] In the experiment, a test was conducted in which the logarithmic F0 pattern was transformed by a specified amount β. The test data used in the experiment was test data that had not been used for training.
[0091] Two types of experiments were conducted: one where β = -log(1.5) and the other where β = log(1.5). In the experiments, mel-spectrogram transformations were performed for each value of β on the speech of each of the four speakers. The experimental results shown in Figure 6, which will be described later, show the average values of the results for β = -log(1.5) and β = log(1.5) for all test data from the four speakers. Therefore, the experimental results in Figure 6 show the average values of eight types of results.
[0092] More specifically, experiments were conducted for each of the four speakers using the two types of β, and 132 test data sets were used for each speaker and β pair. In other words, a total of 2 × 4 × 132 = 1056 experiments were conducted. The experimental results in Figure 6 show the average values of the 1056 experimental results. Furthermore, in the experiments, the speaker indicated by the speaker label was the speaker of the mel spectrogram indicated by the mel spectrogram information. Therefore, no speaker conversion was performed in the experiments conducted this time.
[0093] In the experiment, the mel spectrogram used was a mel spectrogram of speech from which information on silent intervals at the beginning and end of the time axis had been removed.
[0094] In the experiments, the Vector Quantized-Variational AutoEncoder (VQ-VAE) (see Reference 3) was used for encoding.
[0095] Reference 3: A. van den Oord, O. Vinyals, and K. Kavukcuoglu. Neural discrete representation learning. In NeurIPS, volume 30, pages 6309-6318, 2017.
[0096] In the experiment, the F0 label generation process was performed using the following random resampling method. i Let,p i For time i where p = 0, the logarithm F0 label after random resampling is set to zero, and p i For time i that does not satisfy =0, the logarithmic F0 label after random resampling is log(p i ) + β0, where β0 was a random value sampled from a uniform distribution in the interval 0.3 to 3 (inclusive).
[0097] In the experiment, the speaker label generation process involved sampling from a discrete uniform distribution on the set {1, 2, 3, 4}, where "1" represents one of the four speakers, "2" represents one of the remaining three speakers, "3" represents one of the remaining two speakers, and "4" represents the last speaker.
[0098] In the experiments, the root mean square error Δf0 between two logarithmic F0 contours was used as an objective measure of F0 contour transformation. Mel-cepstral distortion (MCD) was also used as an objective measure of voice quality transformation (i.e., speaker transformation). In both evaluations, endpoint-free Dynamic Time Warping (DTW) was applied to eliminate the time misalignment between the target features and the transformed features. These indices were calculated after the time alignment was achieved by applying DTW.
[0099] In the experiment, Δf0 was measured as the error between the target logarithmic F0 contour and the logarithmic F0 contour of the converted speech, which was determined from the converted speech using the WORLD pitch extractor (see Reference 4).
[0100] Reference 4: M. Morise, F. Yokomori, and K. Ozawa. WORLD: A vocoder-based high-quality speech synthesis system for real-time applications. IEICE Transactions on Information and Systems E99.D(7):1877-1884, 2016.
[0101] In the experiment, the MCD between the speech before and after conversion by the decoding process was obtained. The converted speech is the result of the decoding process.
[0102] 6 is a diagram showing an example of the results of an experiment in an embodiment. "DisC" refers to a conversion process using the learned encoding process and decoding process obtained by the learning device 1. "DisC-NoAux" refers to a conversion process using the learned encoding process and decoding process resulting from learning by the learning device 1 using only the reconstruction error as the loss function. Therefore, "DisC-NoAux" is an example of a comparison target for evaluating the performance of "DisC."
[0103] The results in Figure 6 show that the accuracy of mel spectrogram conversion (i.e., audio conversion) by "DisC" is higher than that of mel spectrogram conversion by "DisC-NoAux".
[0104] The learning device 1 in this embodiment configured as described above trains a learning target model using F0 labels, and the training is performed to produce results that indicate differences in F0 labels. As a result, the learning device 1 can obtain a mathematical model that obtains a mel spectrogram that more appropriately reflects changes in the F0 labels than in training without using F0 labels. The training also involves learning an encoding process that extracts mel features, which are feature quantities of the mel spectrogram. Therefore, the learning device 1 can obtain a mathematical model that estimates a mel spectrogram using mel features that are more appropriate than mel features obtained manually.
[0105] If more appropriate mel features are used, the accuracy of mel spectrogram estimation by decoding processing will also be high. For these reasons, the learning device 1 can improve the accuracy of mel spectrogram conversion. As described above, improving the accuracy of mel spectrogram conversion improves the accuracy of speech conversion. Therefore, the fact that the learning device 1 can improve the accuracy of mel spectrogram conversion means that the learning device 1 can improve the accuracy of speech conversion.
[0106] Furthermore, the conversion device 2 in this embodiment configured as above obtains a mel spectrogram whose F0 pattern is the F0 pattern indicated by the destination F0 label, based on the source mel spectrogram, using the learned encoding and decoding processes obtained by the learning device 1. Therefore, the conversion device 2 can improve the accuracy of speech conversion.
[0107] The conversion system 100 of the embodiment configured as described above also includes the learning device 1. Therefore, the conversion system 100 can improve the accuracy of speech conversion.
[0108] (Variation) In addition, a zero label may also be used in the learning process. The zero label is information indicating an F0 pattern with a value of 0. In other words, the zero label is information indicating silence. In such a case, in the learning process, a decoding process is performed on a decoding target (hereinafter referred to as the "0th decoding target") that includes Mel features, a first speaker label, and a zero label. In such a case, in the learning process, an F0 label estimation process is performed on the 0th decoding result. The 0th decoding result is the result of the decoding process on the 0th decoding target. In such a case, the loss function may include, for example, a function indicating the difference between the result of the decoding process on the 0th decoding target and the zero label (hereinafter referred to as the "0th F0 label loss function").
[0109] <Example of loss function> An example of the loss function is shown in the formula below. The loss function is expressed, for example, as the following formula (1).
[0110]
number
[0111]
number
[0112]
number
[0113]
number
[0114]
number
[0115]
number
[0116]
number
[0117]
number
[0118]
number
[0119]
number
[0120] CE(·,·) represents the cross-entropy loss. F represents the size of the input mel spectrogram in the frequency direction. T represents the size of the input mel spectrogram in the time direction. i in equation (2) corresponds to the i-th element in the frequency direction. j in equation (2) corresponds to the j-th element in the time direction. Therefore, V with a hat symbol ij represents the i-th element in the frequency direction and the j-th element in the time direction of the matrix V that represents the variance of the output from the decoder. ij represents the i-th element in the frequency direction and the j-th element in the time direction of the input logarithmic mel spectrogram. ij represents the i-th element in the frequency direction and the j-th element in the time direction of the matrix M that represents the average of the decoder output. s is any natural number.
[0121] p with a circumflex rand represents the result of the F0 label estimation process using the second decoding result as the processing target. rand represents the second F0 label. p0 with a hat symbol represents the result of the F0 label estimation process that processes the 0th decoding result. p0 without a hat symbol represents the zero label.
[0122] The p with a hat symbol represents the result of the F0 label estimation process, which processes the mel spectrogram information input to the model to be trained. The i-th element of p with a hat symbol is p i p without hats represents the F0 label input to the training model. The i-th element of p without hats is expressed as p without hats. i It is expressed as:
[0123] p with a circumflex rec represents the result of the F0 label estimation process that processes the first decoding result.
[0124] s with a caret rand represents the result of the speaker label estimation process using the second decoding result as the processing target. rand represents the second speaker label. s with a hat symbol represents the result of the speaker label estimation process that processes the mel spectrogram information input to the training model. s without a hat symbol represents the speaker label input to the training model.
[0125] s with a caret rec represents the result of the speaker label estimation process for the first decoding result.
[0126] Equation (2) is an example of a reconstruction error. The first term in the Σ symbol on the right side of equation (5) is an example of a 2F0 label loss function. The second term in the Σ symbol on the right side of equation (5) is an example of a 0F0 label loss function. Equation (6) is an example of a 2nd speaker label loss function.
[0127] Equation (7) is an example of the first F0 label loss function. The first term in the Σ symbol on the right side of equation (7) is an example of the fourth subdifference. The second term in the Σ symbol on the right side of equation (7) is an example of the third subdifference.
[0128] Equation (8) is an example of a first speaker label loss function. The first term in the Σ symbol on the right side of Equation (8) is an example of a second subdifference. The second term in the Σ symbol on the right side of Equation (8) is an example of a first subdifference.
[0129] It should be noted that equations (1) to (10) are merely examples, and for example, in equation (1), each function on the right side may be multiplied by a predetermined weight.
[0130] Each of the learning device 1 and the conversion device 2 may be implemented using a plurality of information processing devices communicably connected via a network. In this case, the functional units of each of the learning device 1 and the conversion device 2 may be distributed and implemented among the plurality of information processing devices.
[0131] The control unit 21 included in the conversion device 2 is an example of a conversion control unit.
[0132] All or part of the functions of the learning device 1 and the conversion device 2 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into computer systems. The program may be transmitted via a telecommunications line.
[0133] Although an embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Explanation of symbols]
[0134] 100...conversion system, 1...learning device, 2...conversion device, 11...control unit, 12...input / output interface, 13...storage unit, 21...control unit, 22...input / output interface, 23...storage unit, 91...processor, 92...memory, 93...processor, 94...memory
Claims
1. a control unit that performs learning of a mathematical model of a learning target based on a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of a speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech; Equipped with the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model; the mathematical model includes: a speaker label generation process for generating a speaker label indicating a speaker determined based on a predetermined rule; an F0 label generation process for generating an F0 label indicating an F0 pattern determined based on a predetermined rule; an encoding process for extracting features of the mel spectrogram; and a decoding process for a decoding target, which is a processing target including at least the features, and for estimating a mel spectrogram indicated by the decoding target by decoding the decoding target, The learning is performed by executing the mathematical model; the decoding process in the learning is performed on a first decoding target that is a decoding target including the feature, the speaker label, and the F0 label input to the mathematical model, and a second decoding target that is a decoding target including the feature, a result of the speaker label generation process, and a result of the F0 label generation process; the control unit updates the mathematical model based on a result of the execution of the mathematical model. Learning device.
2. The control unit updates the mathematical model by a first update condition that the mathematical model is updated so that a difference between a result of the decoding process on the first decoding target and a mel spectrogram indicated by the mel spectrogram information becomes small; and a second update condition for updating the mathematical model so that the result of the decoding process for the first decoding target and the result of the decoding process for the second decoding target differ when at least one of a condition that the speaker label input to the mathematical model differs from a result of the speaker label generation process and a condition that the F0 label input to the mathematical model differs from a result of the F0 label generation process is satisfied; To fulfill the above, The learning device according to claim 1 .
3. The mathematical model further comprises: a speaker label estimation process in which a processing object is a result of the decoding process and a speaker label is estimated based on the processing object; an F0 label estimation process in which a processing object is a result of the decoding process and an F0 label is estimated based on the processing object; Including, The control unit updates the mathematical model by a difference indicated by a reconstruction error indicating a difference between a first decoding result, which is a result of the decoding process on the first decoding target, and the mel spectrogram information input to the mathematical model; a difference indicated by a second speaker label loss function indicating a difference between a result of the speaker label estimation process and a result of the speaker label generation process, the second decoding result being a result of the decoding process for the second decoding target; a difference indicated by a second F0 label loss function indicating a difference between a result of the F0 label estimation process and a result of the F0 label generation process, the second decoding result being a result of the decoding process for the second decoding target; To reduce the The learning device according to claim 1 .
4. an acquisition process for acquiring a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of a speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech; a control unit that learns a mathematical model of a learning target based on a speaker label, mel spectrogram information, and an F0 label, wherein the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model, and the mathematical model includes: a speaker label generation process that generates a speaker label indicating a speaker determined based on a predetermined rule; an F0 label generation process that generates an F0 label indicating an F0 pattern determined based on the predetermined rule; an encoding process that extracts features of the mel spectrogram; and a decoding process that is a process on a decoding target that is a processing target including at least the features, and that estimates a mel spectrogram indicated by the decoding target by decoding the decoding target, and the learning is performed by the mathematical model. a conversion process of converting a mel spectrogram indicated by the mel spectrogram information obtained in the acquisition process into a mel spectrogram in which the speaker is the speaker indicated by the speaker label obtained in the acquisition process and the F0 pattern is the F0 pattern indicated by the F0 label obtained in the acquisition process, using the encoding process and the decoding process in the trained mathematical model obtained by the learning device, wherein the decoding process in the training is performed on a first decoding target which is a decoding target including the feature amount, the speaker label and the F0 label input to the mathematical model, and a second decoding target which is a decoding target including the feature amount, a result of the speaker label generation process, and a result of the F0 label generation process, and the control unit updates the mathematical model based on a result of execution of the mathematical model; A conversion control unit that executes A conversion device comprising:
5. a control step of learning a mathematical model of a learning object based on a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of a speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech; and the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model; the mathematical model includes: a speaker label generation process for generating a speaker label indicating a speaker determined based on a predetermined rule; an F0 label generation process for generating an F0 label indicating an F0 pattern determined based on a predetermined rule; an encoding process for extracting features of the mel spectrogram; and a decoding process for a decoding target, which is a processing target including at least the features, and for estimating a mel spectrogram indicated by the decoding target by decoding the decoding target, The learning is performed by executing the mathematical model; the decoding process in the learning is performed on a first decoding target that is a decoding target including the feature, the speaker label, and the F0 label input to the mathematical model, and a second decoding target that is a decoding target including the feature, a result of the speaker label generation process, and a result of the F0 label generation process; the control step updates the mathematical model based on a result of the execution of the mathematical model. How to learn.
6. an acquisition process for acquiring a speaker label indicating a speaker, mel spectrogram information indicating a mel spectrogram of a speech uttered by the speaker, and an F0 label indicating an F0 pattern of the speech; and a control step of learning a mathematical model of a learning target based on a speaker label, mel spectrogram information, and an F0 label, wherein the speaker label, the mel spectrogram information, and the F0 label are input to the mathematical model, and the mathematical model includes: a speaker label generation process of generating a speaker label indicating a speaker determined based on a predetermined rule; an F0 label generation process of generating an F0 label indicating an F0 pattern determined based on the predetermined rule; an encoding process of extracting features of the mel spectrogram; and a decoding process of estimating a mel spectrogram indicated by a decoding target, which is a processing target including at least the features, by decoding the decoding target. a conversion process of converting a mel spectrogram indicated by the mel spectrogram information obtained in the acquisition process into a mel spectrogram in which the speaker is the speaker indicated by the speaker label obtained in the acquisition process and the F0 pattern is the F0 pattern indicated by the F0 label obtained in the acquisition process, using the encoding process and the decoding process in the trained mathematical model obtained by the learning method, wherein the decoding process in the training is performed on a first decoding target which is a decoding target including the feature amount, the speaker label and the F0 label input to the mathematical model, and a second decoding target which is a decoding target including the feature amount, a result of the speaker label generation process, and a result of the F0 label generation process, and the control step updates the mathematical model based on a result of execution of the mathematical model; a conversion control step that performs A conversion method comprising:
7. A program for causing a computer to function as either the learning device according to any one of claims 1 to 3 or the conversion device according to claim 4.
Citation Information
Patent Citations
VAE (Variational Autoencoder)-based voice conversion method under non-parallel corpus training
CN108777140A
Learning method of conversion model and learning device of conversion model
JP2019040123A
Voice conversion learning device, voice conversion device, method and program
JP2019144402A
Text-to-speech synthesis in target speaker voice using neural networks
JP2021524063A
Voice quality conversion device, voice quality conversion method, and program
JP2022127898A