Conversion device, conversion method, and recording medium

The conversion device and method address the challenge of achieving natural voice quality by aligning the pitch range of a speaker's voice with a voice conversion model, ensuring accurate and natural voice output.

WO2026094564A1PCT designated stage Publication Date: 2026-05-07DOWANGO KK
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
DOWANGO KK
Filing Date
2025-10-07
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing voice conversion technologies struggle to produce natural voice quality when the source and destination voice qualities are dissimilar.

Method used

A conversion device and method that acquires and converts the pitch range of a speaker's utterance data to match the pitch range of a voice conversion model, removing irrelevant sounds and adjusting the pitch range to ensure natural voice quality output.

Benefits of technology

Enables natural voice quality conversion regardless of the original voice quality, enhancing accuracy and usability of voice conversion systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025035565_07052026_PF_FP_ABST
    Figure JP2025035565_07052026_PF_FP_ABST
Patent Text Reader

Abstract

This conversion device 1 comprises: an acquisition unit 21 that acquires a vocal range of a voice quality conversion model and a vocal range of a speaker; and a conversion unit 23 that converts a vocal range of utterance data 14 of a speaker P into the vocal range of the voice quality conversion model and outputs the vocal range–converted utterance data 17 to the voice quality conversion model.
Need to check novelty before this filing date? Find Prior Art

Description

Conversion device, conversion method, and recording medium

[0001] The present invention relates to a conversion device, a conversion method, and a recording medium.

[0002] With the recent development of AI (Artificial Intelligence) technology, there is a voice conversion technology that can accurately imitate voice quality. The voice conversion model used for conversion in voice conversion technology is trained with the voices of characters such as voice actors. When the voice quality of the source of voice conversion is similar to the voice quality of the character of the voice conversion model, the voice conversion technology can output natural voice.

[0003] Patent Document 1 performs voice quality conversion on the voice in which the characteristics are reflected when characteristic voice is input into the voice corresponding to the speaker information of the conversion destination.

[0004] Japanese Unexamined Patent Application Publication No. 2024-018197

[0005] When the voice quality of the source of conversion and the voice quality of the speaker after conversion are not similar, the voice conversion technology may output unnatural voice quality.

[0006] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technology that can convert to natural voice quality regardless of the voice quality of the source of conversion.

[0007] A conversion device according to an aspect of the present invention includes an acquisition unit that acquires the pitch range of a voice conversion model and the pitch range of a speaker, and converts the pitch range of the utterance data of the speaker to the pitch range of the voice conversion model, and outputs the utterance data after pitch range conversion to the voice conversion model.

[0008] A conversion method according to an aspect of the present invention is that a computer acquires the pitch range of a voice conversion model and the pitch range of a speaker, converts the pitch range of the utterance data of the speaker to the pitch range of the voice conversion model, and outputs the utterance data after pitch range conversion to the voice conversion model.

[0009] A recording medium according to one aspect of the present invention is a computer-readable recording medium that records a program that causes a computer to function as an acquisition unit that acquires the pitch range of a voice conversion model and the pitch range of a speaker, and a conversion unit that converts the pitch range of the speaker's speech data to the pitch range of the voice conversion model and outputs the speech data after pitch range conversion to the voice conversion model.

[0010] According to the present invention, it is possible to provide a technology that can convert a voice to a natural voice quality regardless of the original voice quality.

[0011] Figure 1 is a diagram illustrating the system configuration of the conversion system of this disclosure. Figure 2 is a diagram illustrating the conversion method in the conversion system of this disclosure. Figure 3 is a diagram illustrating the functional blocks of the conversion device. Figure 4 is a diagram illustrating an example of the data structure of pitch range data. Figure 5 is a diagram illustrating an example of the data structure of character pitch range data. Figure 6 is a diagram illustrating an example of pitch range compression by the conversion device. Figure 7 is a flowchart illustrating the conversion process by the conversion device. Figure 8 is a diagram illustrating the system configuration of a modified conversion system. Figure 9 is a diagram illustrating the functional blocks of the practice device. Figure 10 is a diagram illustrating the conversion method in a modified conversion system. Figure 11 is a diagram illustrating the hardware configuration of a computer used in devices such as a conversion device.

[0012] Embodiments of the present invention will be described below with reference to the drawings. In the drawings, identical parts are denoted by the same reference numerals and their descriptions are omitted.

[0013] (Conversion System) As shown in Figure 1, the conversion system 7 according to this disclosure inputs the utterance of speaker P into a voice quality conversion model, and uses the voice quality conversion model to convert the utterance of speaker P into a predetermined voice quality and output it. The voice quality conversion model may also be called, for example, an AI voice changer.

[0014] In this disclosure, vocal range refers to the range of pitches. Voice quality refers to the characteristics or resonance of a voice. In this disclosure, voice quality may exclude vocal range or may include vocal range.

[0015] The converted voice quality of a voice conversion model can be any voice quality, such as the voice quality of a character voiced by a voice actor, or a voice quality generated programmatically. The voice conversion model is a model that has learned from any voice quality, such as the voice quality of a character voiced by a voice actor, or a voice quality generated through experiments, etc.

[0016] In this disclosure, the voice conversion model is a model learned from the voice qualities of characters. A voice conversion model is provided for each character identifier. The characters may include virtual characters corresponding to voice qualities generated in experiments, etc.

[0017] As shown in Figure 1, the conversion system 7 comprises a conversion device 1 and a voice quality conversion device 3.

[0018] The voice conversion device 3 uses a voice conversion model to convert the voice quality of the input speech data to a predetermined voice quality. The voice conversion device 3 may have multiple voice conversion models and may convert and output the input voice quality using a specified voice conversion model.

[0019] In this disclosure, for each voice conversion model used by the voice conversion device 3, the pitch range of the converted voice of that voice conversion model is predetermined. The pitch range of the voice conversion model is the pitch range of the character etc. learned by the voice conversion model, and is a pitch range that can guarantee the quality of voice conversion.

[0020] The vocal range of the voice conversion model is shared with conversion device 1. For each voice conversion model used by voice conversion device 3, conversion device 1 stores the vocal range of the converted voice as character vocal range data 12.

[0021] In this disclosure, the conversion device 1 processes the speaker P's speech data so that the voice conversion device 3 can convert it into a natural voice using the voice conversion model. The conversion device 1 converts the pitch range of the speech data input to the voice conversion device 3 to match the pitch range of the voice conversion model. First, the conversion device 1 identifies the pitch range of the voice conversion model character and the pitch range of speaker P. The conversion device 1 converts the pitch range of speaker P's speech data 14 to the pitch range of the voice conversion model character and inputs the pitch-range converted speech data to the voice conversion device 3.

[0022] The voice conversion device 3 receives data in which the pitch range of speaker P's speech data 14 has been converted to the character's pitch range. The voice conversion device 3 uses a voice conversion model to output data in which the input data has been converted to the character's voice quality. The voice conversion device 3 outputs data in which speaker P's speech is the same in content, speed, and pitch, but the voice quality has been converted to the voice quality of the voice conversion model.

[0023] In this way, the conversion device 1 inputs the speech data converted to the pitch range of the voice quality to be converted to the voice quality conversion device 3. As a result, even if the voice quality of speaker P and the voice quality converted by the voice quality conversion model are not similar, the conversion system 7 can output speech in which the voice quality of speaker P's utterances has been naturally converted to the voice quality of the voice quality conversion model.

[0024] Referring to Figure 2, the conversion method using the conversion system 7 will be explained.

[0025] In step S1, the conversion device 1 obtains the vocal range of the character in the voice conversion model used by the voice conversion device 3 and the vocal range of speaker P. For example, the conversion device 1 obtains the character identifier specified by speaker P and obtains the vocal range of the specified character identifier from the character vocal range data 12. The conversion device 1 has speaker P read a sample sentence or the like to obtain speaker P's speech data and obtains speaker P's vocal range from that speech data.

[0026] In step S2, the conversion device 1 acquires the speech data 14 of speaker P. The speech data 14 acquired here is the target of conversion by the voice quality conversion model.

[0027] In step S3, the conversion device 1 removes the unwanted data 15 from the speech data 14. The unwanted data 15 is data that is not subject to conversion by the voice conversion model used by the voice conversion device 3. The unwanted data 15 includes, for example, laughter, growls, sneezes, etc., which are sounds made by a person but are not spoken voices.

[0028] In step S4, the conversion device 1 outputs the unwanted data 15, which has been removed from the speech data 14, to the voice quality conversion device 3.

[0029] In step S5, the conversion device 1 converts the pitch range of the speech data 16, from which the irrelevant data 15 has been removed, to the pitch range of the character acquired in step S1, thereby generating speech data 17 after pitch range conversion. The conversion device 1 converts the pitch range using methods such as shift conversion or pitch range compression.

[0030] In step S6, the conversion device 1 outputs the speech data 17 generated in step S5 after pitch range conversion to the voice quality conversion device 3.

[0031] In step S7, the voice conversion device 3 converts the speech data 17, which has been converted in pitch range and input from the conversion device 1, into the character's voice using a voice conversion model. In step S8, the voice conversion device 3 synthesizes the non-target data 15, which was input in step S4, with the data converted into the character's voice, and plays it back.

[0032] The process shown in Figure 2 is just one example and is not limited to this.

[0033] For example, the conversion system 7 described herein describes a case in which a conversion device 1 converts the pitch range of a set of speech data 14 input by speaker P, and a voice quality conversion device 3 converts the voice quality of the data whose pitch range has been converted by the conversion device 1, but is not limited to this case.

[0034] The speech of speaker P and the voice conversion of that speech may be performed in real time. For example, the conversion device 1 sequentially acquires speech data 14 while speaker P is speaking and identifies the unit data to be processed from the speech data 14. For each of the unit data, the conversion device 1 sequentially repeats the processing from steps S3 to S8. The unit data is a single data item obtained by dividing the speech data 14 into predetermined units such as phonemes, syllables, or moras.

[0035] This allows the conversion system 7 to perform speaker P's utterance and voice quality conversion of that utterance in real time. By performing voice quality conversion in real time, the conversion system 7 can be used for real-time processing such as voice chat with other people.

[0036] The conversion system 7 described herein involves a conversion device 1 that converts the pitch range of speaker P's speech data 14 to match the pitch range of a voice quality conversion model, and outputs the converted data to a voice quality conversion device 3. As a result, the voice quality conversion device 3 only needs to convert the voice quality of the data in the pitch range that matches the pitch range of the voice quality conversion model character, thus enabling accurate voice quality conversion and output of natural-sounding speech data.

[0037] Furthermore, sounds that are emitted by a person but are not spoken voices, such as the excluded data 15, can be a factor in degrading the accuracy of conversion by the voice quality conversion model. Therefore, in the conversion system 7 according to this disclosure, the conversion device 1 removes the excluded data 15 from the speech data 14 and outputs the excluded data 15 to the voice quality conversion device 3, and also outputs speech data 17 with a converted pitch range, obtained by converting the pitch range of the speech data 14, to the voice quality conversion device 3. The voice quality conversion device 3 converts the voice quality of the speech data 17 with a converted pitch range, which is input from the conversion device 1, using the voice quality conversion model, and then superimposes the excluded data 15 input from the conversion device 1 to reproduce it. In this way, the conversion system 7 can reproduce each of the speaker P's utterances while ensuring the accuracy of voice quality conversion by the voice quality conversion model.

[0038] (Conversion device) The conversion device 1 according to this disclosure will be described with reference to Figure 3.

[0039] As shown in Figure 3, the conversion device 1 includes the following data: pitch range data 11, character pitch range data 12, character identifier 13, speech data 14, excluded data 15, removed speech data 16, and speech data 17 after pitch range conversion, as well as the functions of acquisition unit 21, removal unit 22, and conversion unit 23. Each data is stored in a storage device such as memory 902 or storage 903. Each function is implemented in the CPU 901.

[0040] The pitch range data 11 is the data of the pitch range of the speaker P. For example, the speaker P is made to speak by having the speaker P read a sample sentence in advance. The conversion device 1 identifies the pitch range of the speaker P from the speech of the speaker P and sets it in the pitch range data 11. The pitch range data 11 identifies the lowest pitch and the highest pitch as the pitch range of the speaker P, for example, as shown in FIG. 4. The pitch range data 11 may be set by the speaker P. Also, the pitch range data 11 may set the pitch range estimated from the attribute information such as the age, gender, and height of the speaker P.

[0041] The character pitch range data 12 is the data for identifying the pitch range of the voice conversion model. The pitch range of the voice conversion model is the pitch range that can ensure the quality of voice conversion.

[0042] When there are voice conversion models for each of a plurality of characters, the pitch range of each character is identified. The character pitch range data 12 associates the lowest pitch and the highest pitch with the character identifier, for example, as shown in FIG. 5.

[0043] The character identifier 13 identifies the character of one of the plurality of voice conversion models used by the voice conversion device 3.

[0044] The speech data 14 is the data of the speech of the original speaker P. The speech data 14 has a data format that can be processed by a computer.

[0045] The non-target data 15 is the data of the sound that is not speech in the speech data 14. The non-target data 15 is generated by the removal unit 22.

[0046] The speech data 16 after removal is the data after the non-target data 15 is removed from the speech data 14. The speech data 16 after removal is the data that is the processing target of the conversion unit 23 among the speech data 14. The speech data 16 after removal is generated by the removal unit 22.

[0047] The speech data 17 after pitch conversion is the data obtained by converting the pitch range of the speech data 16 after removal to the pitch range of the voice conversion model. The speech data 17 after pitch conversion is generated by the conversion unit 23.

[0048] The acquisition unit 21 acquires the pitch range of the voice conversion model, the pitch range of the speaker P, and the utterance data 14 of the speaker P.

[0049] The pitch range of the voice conversion model is obtained, for example, when the speaker P designates the character of the voice conversion model. In the present disclosure, the conversion device 1 holds the pitch range of each character of the plurality of voice conversion models used by the voice conversion device 3 as character pitch range data 12. The acquisition unit 21 specifies one character identifier 13 among the plurality of characters defined by the character pitch range data 12 by an operation by the speaker P or the like, and acquires the pitch range corresponding to the specified character identifier 13 from the character pitch range data 12.

[0050] The acquisition unit 21 acquires the pitch range of the speaker P in advance as pitch range data 11.

[0051] After the acquisition unit 21 acquires the pitch range of the voice conversion model and the pitch range of the speaker P, the acquisition unit 21 acquires the utterance data 14 of the speaker P. The acquisition unit 21 causes the speaker P to utter the content to be converted into the voice conversion model, and acquires the utterance data 14 of the speaker P.

[0052] The removal unit 22 removes the non-target data 15 that is not the target of conversion by the voice conversion model from the utterance data 14 of the speaker P. The removal unit 22 specifies, as the non-target data 15, data in the utterance data 14 that is laughter, groans, sneezes, etc., which are sounds made by a person but not speech. The removal unit 22 deletes the non-target data 15 from the utterance data 14 and generates the utterance data 16 after removal.

[0053] The removal unit 22 further outputs the non-target data 15 to the voice conversion model of the voice conversion device 3. The non-target data is synthesized into the data after conversion by the voice conversion model.

[0054] The conversion unit 23 converts the pitch range of the utterance data 14 of the speaker P into the pitch range of the voice conversion model, and generates the utterance data 17 after pitch range conversion. The conversion unit 23 outputs the utterance data 17 after pitch range conversion to the voice conversion model of the voice conversion device 3.

[0055] The conversion unit 23 shifts the vocal range of speaker P in the speech data 14 to the vocal range of the voice quality conversion model. For example, the conversion unit 23 calculates the shift amount from a reference value of speaker P's vocal range and a reference value of the voice quality conversion model's vocal range. The reference value is, for example, the median of the vocal range. The conversion unit 23 shifts the vocal range of the speech data 14 according to the calculated shift amount.

[0056] Consider the case where the vocal range of the model to be converted is C4-C5, and speaker P is male with a vocal range of C3-E4. Since speaker P's vocal range is lower than that of the model to be converted, a pitch correction of +12 (1 octave) is applied. As a result, the vocal range of speaker P's speech data 14 is converted from C3-E4 to C4-E5.

[0057] In this disclosure, we will explain the case in which the shift amount is calculated from the vocal range of speaker P, but the shift amount may also be set by speaker P.

[0058] The conversion unit 23 compresses the pitch range of the speech data 14 after shift conversion that falls outside the pitch range of the voice quality conversion model to the pitch range of the voice quality conversion model.

[0059] The pitch range after shift conversion may fall outside the pitch range of the voice conversion model. To bring the pitch range after shift conversion within the pitch range of the voice conversion model, a certain range of pitches that falls outside the pitch range of the voice conversion model is pushed into a certain upper and lower range of the voice conversion model's pitch range.

[0060] Figure 6 shows an example of compression by the conversion unit 23. The vocal range of speaker P after shift conversion is wider in the high and low ranges compared to the vocal range of the voice quality conversion model. Therefore, the compressed vocal range is set to two semitones above and below the vocal range of the voice quality conversion model. The conversion unit 23 pushes the vocal range that falls outside the vocal range of the voice quality conversion model into the compressed vocal range. Here, the arrangement of each note within the compressed vocal range is processed using moving averages or exponential moving averages to smooth the changes in pitch.

[0061] The compressed frequency range may be determined by the number of sounds that fall outside the frequency range of the voice conversion model within the shifted frequency range of speaker P. The compressed frequency range may be pre-set for each voice conversion model. The compressed frequency range may also be determined by adjusting the pre-set frequency range for the voice conversion model using the number of sounds that fall outside the frequency range of the voice conversion model.

[0062] If the speech data 14 contains unwanted data 15, the conversion unit 23 converts the pitch range of the data from which the unwanted data 15 has been removed from the speech data 14 into the pitch range of the voice quality conversion model, thereby generating speech data 17 after pitch range conversion. The conversion unit 23 outputs the speech data 17 after pitch range conversion to the voice quality conversion model of the voice quality conversion device 3.

[0063] In this disclosure, the conversion unit 23 has described a case in which, after shifting the pitch range, there are pitch ranges in the shifted speech data that fall outside the pitch range of the voice quality conversion model, and the conversion unit 23 compresses them. However, the disclosure is not limited to this case. If the pitch range of speaker P and the pitch range of the voice quality conversion model are the same or close, the conversion unit 23 may, without using shift conversion, compress the pitch range of speaker P in the speech data 14 that falls outside the pitch range of the voice quality conversion model to the pitch range of the voice quality conversion model. For example, if the reference value of speaker P's pitch range and the reference value of the voice quality conversion model's pitch range are the same or within a predetermined value, the conversion unit 23 may perform only pitch range compression without shift conversion.

[0064] In this disclosure, the conversion unit 23 is described in the case of converting the pitch range, but it is not limited to this. If the pitch range of the speaker P's speech data 14 is within the pitch range of the voice quality conversion model, the conversion unit 23 does not need to convert the pitch range of the speech data 14.

[0065] The conversion process by the conversion unit 23 will be explained with reference to Figure 7.

[0066] In step S101, the conversion unit 23 obtains speech data 16 after removing the unwanted data 15. In step S102, the conversion unit 23 shifts the pitch range of the speech data 16 after removal.

[0067] In step S103, the conversion unit 23 determines whether there are any pitch ranges that fall outside the pitch range of the voice quality conversion model after the shift conversion in step S102. If there are pitch ranges that fall outside the range, the process proceeds to step S104. If there are no pitch ranges that fall outside the range, the process proceeds to step S105.

[0068] In step S104, the conversion unit 23 compresses the frequency ranges that fall outside the frequency range of the voice quality conversion model into the frequency range of the voice quality conversion model.

[0069] The conversion unit 23 generates speech data 17 after pitch range conversion by shift conversion in step S102, or by shift conversion in step S102 and compression in step S104. In step S105, the conversion unit 23 outputs the speech data 17 after pitch range conversion to the voice quality conversion device 3.

[0070] The conversion system 7 described herein converts the vocal range of speaker P's speech data into the vocal range of the voice conversion model before inputting it into the voice conversion model. Therefore, it can convert to a natural voice regardless of the original voice quality.

[0071] (Modified Version) A modified version of the conversion system 7a will be described with reference to Figure 8. The conversion system 7a shown in Figure 8 differs from the conversion system 7 shown in Figure 1 in that it includes a practice device 5.

[0072] Generally, voice conversion models can be difficult for beginners to use. This is because they may not know whether the microphone is on, whether they are speaking at the appropriate volume, what pitch is necessary to convert to a natural voice, or what speaking style will make the target character speak naturally.

[0073] Therefore, the practice device 5 supports speaker P's practice so that they can master the voice quality conversion model.

[0074] In the conversion system 7a, the conversion device 1 converts the speaker P's vocal range to the vocal range of the character in the voice quality conversion model, and the voice quality conversion device 3 converts the utterance converted to the character's vocal range to the character's voice quality. The practice device 5 is used for speaker P to practice the character's utterances other than the character's vocal range and voice quality.

[0075] The practice device 5 includes dialogue data 51, personality data 52, speech data 53, and vocal range data 54, as well as the functions of a display unit 61, an acquisition unit 62, a playback unit 63, and a specific unit 64. Each piece of data is stored in a storage device such as a memory 902 or storage 903. Each function is implemented in the CPU 901.

[0076] The dialogue data 51 identifies characteristic lines of a character used in the voice quality conversion model. The dialogue data 51 contains one or more lines of dialogue. The dialogue data 51 may also include symbols that identify the intonation, accent, and other nuances of speech when the character speaks those lines.

[0077] Personality data 52 identifies the character's personality or attributes. Personality data 52 identifies the personality or attributes that influence the character's speech. Personality data 52 may be the character's personality, attributes, or data containing attributes, or it may be data that identifies events that describe the character's personality or attributes.

[0078] Speech data 53 is data from an utterance by speaker P.

[0079] The vocal range data 54 is data that identifies the vocal range of speaker P. The vocal range data 54 is, for example, the data shown in Figure 4. The vocal range data 54 is generated by the identification unit 64.

[0080] The display unit 61 displays a message to the speaker P, the source of the voice quality conversion model, prompting them to utter the dialogue data 51. The display unit 61 further displays personality data 52. The display unit 61 can prompt speaker P to utter the character's lines after understanding the character's personality or attributes. The conversion device 1 and conversion unit 23 convert speaker P's vocal range and voice quality to the character's vocal range and voice quality. The conversion system 7a makes it easier to convert speaker P's utterances into the character's utterances.

[0081] The acquisition unit 62 acquires the speech data 53 of speaker P.

[0082] The playback unit 63 plays back the acquired speech data 53 of speaker P. By playing back the speaker's speech data 14, speaker P can check their own speech and practice to make it sound more like the character's speech.

[0083] In addition to the speaker P's speech data 53, the playback unit 63 may also play back speech data that has been converted from the speech data 14 to the character's vocal range and voice quality by the conversion device 1 and the voice quality conversion device 3. The playback unit 63 may further play back the speech of the character of the voice quality conversion model.

[0084] The playback unit 63 may repeatedly play back the speech data 53 from speaker P, the speech data 14 from which the pitch range and voice quality have been converted, and the character's speech. Speaker P can compare each speech and analyze them to make them closer to the character's speech.

[0085] The playback unit 63 may not only play back each utterance data, but also analyze the speaker P's utterance data 14 to analyze characteristics such as speaking speed, volume, and intonation. The playback unit 63 may also display differences between the characteristics of the speaker P's utterance and those of the character as items that need to be corrected. Furthermore, if the characteristics of speaker P's utterance data 14 and those of the character's utterance are the same or within a predetermined value, the playback unit 63 may play back audio data in which the character praises the speaker's utterance.

[0086] The identification unit 64 identifies the vocal range of speaker P from the speech data 53 and generates vocal range data 54. The vocal range data 54 is transmitted to the conversion device 1 and is referenced when converting the vocal range of speaker P's speech data. The identification unit 64 may also identify the shift amount when shifting the vocal range and share the identified shift amount with the conversion device 1. The identification unit 64 may also identify the compressed vocal range when compressing the vocal range and share the identified compressed vocal range with the conversion device 1. In addition, the identification unit 64 may identify parameters in vocal range conversion or voice quality conversion and share them with the conversion device 1 or voice quality conversion device 3.

[0087] The identification unit 64 may identify a microphone gain suitable for speaker P, in addition to the frequency range, and set it to the identified gain. The identification unit 64 may also analyze speaker P's speech data 14 to identify the ratio of mixing speaker P's voice with the voice after frequency range and voice quality conversion, and the amount of noise cancellation. The identified ratio and noise cancellation amount are shared and referenced by the conversion device 1 or the voice quality conversion device 3.

[0088] When speaker P selects a voice conversion model to which they wish to change their voice, the practice device 5 displays a line of dialogue that matches the selected voice conversion model and plays a sample audio file of the dialogue being read aloud in the target voice quality. The device converts the pitch range and voice quality of speaker P's speech as they read the dialogue and records it, then plays back the recorded audio. The practice device 5 may also provide opportunities to repeat this practice multiple times.

[0089] The practice method relating to this disclosure will be explained with reference to Figure 10.

[0090] In step S51, the training device 5 displays the dialogue data 51. In step S52, the training device 5 displays the personality data 52. At this time, the training device 5 may also play back the audio of the character speaking the lines from the dialogue data 51.

[0091] In step S53, the training device 5 acquires speech data 14 from speaker P. In step S54, the training device 5 may play back the speech data 14. The training device 5 may also play back the audio of a character speaking the lines from the dialogue data 51. The training device 5 may play back speaker P's speech data 53 and the voice from the character simultaneously. The training device 5 may play back speaker P's speech data 53 and then the voice from the character, or play back the voice from the character and then play back speaker P's speech data 53.

[0092] The practice device 5 also identifies the vocal range of speaker P from the speech data 53. The identification unit 64 may also identify parameters used for converting vocal range or voice quality, in addition to the vocal range of speaker P. The practice device 5 transmits the vocal range data and parameters to the conversion device 1 or the voice quality conversion device.

[0093] The training device 5, which relates to this modified version, provides an opportunity for speaker P to practice so that their speech data 53 approaches the speech of a character. Speaker P can become familiar with using the voice conversion model and practice to make their speech sound more like the speech of a character. The training device 5 can contribute to the widespread adoption of voice conversion models.

[0094] (Technology A) A training device comprising a storage device that stores dialogue data that identifies characteristic lines of a character used in the conversion of a voice quality conversion model, and a display unit that displays a message prompting the speaker of the source of the voice quality conversion model to utter the dialogue data.

[0095] (Technology B) The training device according to Technology A, wherein the display unit further displays personality data that identifies the character's characteristics or attributes.

[0096] The devices described above—the conversion device 1, the voice quality conversion device 3, and the practice device 5—are each part of a general-purpose computer system comprising, for example, a CPU (Central Processing Unit, processor) 901, memory 902, storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), communication device 904, input device 905, and output device 906. In this computer system, the CPU 901 executes a program loaded onto the memory 902, thereby realizing the respective functions of each device.

[0097] Each device may be implemented on a single computer, or on multiple computers. Furthermore, each device may be a virtual machine implemented on a computer.

[0098] The programs for each device can be stored on computer-readable storage media such as HDDs, SSDs, USB (Universal Serial Bus) memory, CDs (Compact Discs), and DVDs (Digital Versatile Discs), or they can be distributed over a network.

[0099] Any part or all of the functional components described in this disclosure may be implemented by program. The programs referred to in this disclosure may be distributed non-temporarily on a computer-readable recording medium, distributed via communication lines such as the Internet (including wireless communication), or distributed installed on any terminal. While those skilled in the art may conceive of additional effects and various modifications of the present invention based on the above description, the aspects of this disclosure are not limited to the individual embodiments described above. Various additions, modifications, and partial deletions are possible without departing from the conceptual idea and spirit of the present invention derived from the claims and their equivalents. For example, what is described in this disclosure as a single device (or component, hereinafter the same) (including what is depicted as a single device in the drawings) may be implemented by multiple devices. Conversely, what is described in this disclosure as multiple devices (including what is depicted as multiple devices in the drawings) may be implemented by a single device. Alternatively, some or all of the means or functions included in one device (e.g., a server) may be included in another device (e.g., a user terminal). Furthermore, a "system" may consist of one device, or it may consist of two or more devices (for example, a server and a user terminal, or multiple user terminals).

[0100] Furthermore, not all matters described in this disclosure are mandatory requirements. In particular, matters described in this disclosure but not in the claims can be considered optional additional matters.

[0101] It should be noted that the applicant is only aware of the prior art inventions described in the "Prior Art Documents" section of this disclosure, and this disclosure is not necessarily intended to solve the problems described in those prior art inventions. The problems that this disclosure aims to solve should be determined by considering the disclosure as a whole. For example, if this disclosure describes that a certain effect is achieved by a specific configuration, it can also be said that the problem that is the inverse of that specified effect is solved. However, this does not necessarily mean that such a specific configuration is an essential requirement.

[0102] 1 Conversion device 3 Voice quality conversion device 5 Practice device 7 Conversion system 11, 54 Vocal range data 12 Character vocal range data 13 Character identifier 14, 53 Speech data 15 Data to exclude 16 Speech data after removal 17 Speech data after vocal range conversion 21, 62 Acquisition unit 22 Removal unit 23 Conversion unit 51 Dialogue data 52 Personality data 61 Display unit 63 Playback unit 64 Identification unit 901 CPU 902 Memory 903 Storage 904 Communication device 905 Input device 906 Output device P Speaker

Claims

1. A conversion device comprising: an acquisition unit that acquires the vocal range of a voice quality conversion model and the vocal range of a speaker; and a conversion unit that converts the vocal range of the speaker's speech data to the vocal range of the voice quality conversion model and outputs the speech data after vocal range conversion to the voice quality conversion model.

2. The conversion device according to claim 1, wherein the conversion unit shifts and converts the speaker's vocal range in the speech data to the vocal range of the voice quality conversion model.

3. The conversion device according to claim 2, wherein the conversion unit compresses the range of the speech data after shift conversion that falls outside the range of the voice quality conversion model to the range of the voice quality conversion model.

4. The conversion device according to claim 1, wherein the conversion unit compresses the range of the speaker's voice in the speech data that falls outside the range of the voice quality conversion model to the range of the voice quality conversion model.

5. The conversion device according to claim 1, further comprising a removal unit that removes data from which the voice quality conversion model does not apply from the speaker's speech data, wherein the conversion unit converts the pitch range of the data from which the data to be converted has been removed from the speech data to the pitch range of the voice quality conversion model.

6. The conversion device according to claim 5, wherein the removal unit outputs the excluded data to the voice quality conversion model, and the excluded data is synthesized with the data converted by the voice quality conversion model.

7. A conversion method comprising: a computer acquiring the vocal range of a voice conversion model and the vocal range of a speaker; converting the vocal range of the speaker's speech data to the vocal range of the voice conversion model; and outputting the speech data after vocal range conversion to the voice conversion model.

8. A computer-readable recording medium that stores a program that causes a computer to function as an acquisition unit that acquires the pitch range of a voice conversion model and the pitch range of a speaker, and a conversion unit that converts the pitch range of the speaker's speech data to the pitch range of the voice conversion model and outputs the speech data after pitch range conversion to the voice conversion model.

Citation Information

Patent Citations

  • Voice quality conversion system

    JP2018005048A

  • Computer program, server device, terminal device and voice signal processing method

    JP2023123694A

  • Voice Conversion Training and Data Collection

    US20080255827A1