Conversion device, conversion method, and program

The conversion device and method address unnatural voice quality issues by aligning the pitch range of a speaker's voice with a voice quality conversion model, ensuring natural voice output.

JP2026079662APending Publication Date: 2026-05-15DOWANGO KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DOWANGO KK
Filing Date
2025-03-14
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing voice conversion technologies produce unnatural voice quality when the source and conversion speaker qualities are dissimilar.

Method used

A conversion device and method that acquires and converts the pitch range of a speaker's speech data to match the pitch range of a voice quality conversion model, ensuring natural voice quality output regardless of the original voice quality.

Benefits of technology

Enables natural voice quality conversion by aligning the pitch range of the speaker's voice with the voice quality conversion model, resulting in accurate and natural-sounding speech data output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026079662000001_ABST
    Figure 2026079662000001_ABST
Patent Text Reader

Abstract

This invention provides a conversion device, conversion method, and program that convert voices to a natural quality regardless of the source voice quality. [Solution] The conversion device 1 includes an acquisition unit 21 that acquires the vocal range of the voice quality conversion model and the vocal range of the speaker, and a conversion unit 23 that converts the vocal range of the speaker P's speech data 14 to the vocal range of the voice quality conversion model and outputs the speech data 17 after vocal range conversion to the voice quality conversion model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a conversion device, a conversion method, and a program.

Background Art

[0002] With the recent development of AI (Artificial Intelligence) technology, there is a voice conversion technology that can accurately imitate voice quality. The voice conversion model used for voice conversion in voice conversion technology is trained with the voices of characters such as voice actors. When the voice quality of the source of voice conversion is similar to the voice quality of the character of the voice conversion model, the voice conversion technology can output a natural voice.

[0003] Patent Document 1 performs voice quality conversion on the voice in which the characteristics are reflected into the voice corresponding to the speaker information of the conversion destination when characteristic voice is input.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] When the voice quality of the source of conversion and the voice quality of the speaker after conversion are not similar, the voice conversion technology may output an unnatural voice quality.

[0006] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a technology capable of converting to a natural voice quality regardless of the voice quality of the source of conversion.

Means for Solving the Problems

[0007] A conversion device according to one aspect of the present invention comprises an acquisition unit that acquires the pitch range of a voice quality conversion model and the pitch range of a speaker, and a conversion unit that converts the pitch range of the speaker's speech data to the pitch range of the voice quality conversion model and outputs the speech data after pitch range conversion to the voice quality conversion model.

[0008] In one embodiment of the present invention, a conversion method is provided in which a computer acquires the pitch range of a voice conversion model and the pitch range of a speaker, converts the pitch range of the speaker's speech data to the pitch range of the voice conversion model, and outputs the speech data after pitch range conversion to the voice conversion model.

[0009] A program according to one aspect of the present invention causes a computer to function as an acquisition unit that acquires the pitch range of a voice conversion model and the pitch range of a speaker, and a conversion unit that converts the pitch range of the speaker's speech data to the pitch range of the voice conversion model and outputs the speech data after pitch range conversion to the voice conversion model. [Effects of the Invention]

[0010] According to the present invention, it is possible to provide a technology that can convert a voice to a natural voice quality regardless of the original voice quality. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a diagram illustrating the system configuration of the conversion system of this disclosure. [Figure 2] Figure 2 is a diagram illustrating the conversion method in the conversion system of this disclosure. [Figure 3] Figure 3 is a diagram illustrating the functional blocks of the conversion device. [Figure 4] Figure 4 illustrates an example of a data structure for pitch range data. [Figure 5] Figure 5 illustrates an example of a data structure for character vocal range data. [Figure 6] Figure 6 illustrates an example of frequency range compression by a conversion device. [Figure 7] Figure 7 is a flowchart illustrating the conversion process performed by the conversion device. [Figure 8] FIG. 8 is a diagram for explaining the system configuration of the conversion system of the modified example. [Figure 9] FIG. 9 is a diagram for explaining the functional blocks of the training device. [Figure 10] FIG. 10 is a diagram for explaining the conversion method in the conversion system of the modified example. [Figure 11] FIG. 11 is a diagram for explaining the hardware configuration of a computer used in a device such as a conversion device.

Embodiments for Carrying Out the Invention

[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In the description of the drawings, the same reference numerals are given to the same parts and the description thereof is omitted.

[0013] (Conversion System) As shown in FIG. 1, the conversion system 7 according to the present disclosure inputs the utterance of the speaker P into the voice quality conversion model, and converts and outputs the utterance of the speaker P into a predetermined voice quality using the voice quality conversion model. The voice quality conversion model may be, for example, also called an AI voice changer in some cases.

[0014] In the present disclosure, the pitch range is the range of the pitch of sound. The voice quality is the characteristic or timbre of the voice. In the present disclosure, the voice quality may exclude the pitch range or may include the pitch range.

[0015] The voice quality after conversion by the voice quality conversion model is, for example, an arbitrary voice quality such as the voice quality of a character played by a voice actor, a programmatically generated voice quality, etc. The voice quality conversion model is a model that has learned an arbitrary voice quality such as the voice quality of a character played by a voice actor, a voice quality generated by experiments, etc.

[0016] In the present disclosure, the voice quality conversion model is a model learned from the voice quality of a character. A voice quality conversion model is prepared for each character identifier. The character may include a virtual character for the voice quality generated in experiments, etc.

[0017] As shown in FIG. 1, the conversion system 7 includes a conversion device 1 and a voice quality conversion device 3.

[0018] The voice quality conversion device 3 uses a voice quality conversion model to convert the voice quality of the input utterance data into a predetermined voice quality. The voice quality conversion device 3 may have a plurality of voice quality conversion models and use the specified voice quality conversion model to convert and output the input voice quality.

[0019] For each voice quality conversion model used by the voice quality conversion device 3 in the present disclosure, the pitch range of the voice quality after conversion of the voice quality conversion model is predetermined. The pitch range of the voice quality conversion model is the pitch range of a character or the like learned by the voice quality conversion model, and is a pitch range that can ensure the quality of voice quality conversion.

[0020] The pitch range of the voice quality conversion model is shared with the conversion device 1. The conversion device 1 holds, as character pitch range data 12, the pitch range of the voice quality after conversion for each voice quality conversion model used by the voice quality conversion device 3.

[0021] In the present disclosure, the conversion device 1 processes the utterance data of the speaker P so that the voice quality conversion device 3 can convert the voice quality into a natural voice quality using the voice quality conversion model. The conversion device 1 converts the pitch range of the utterance data input to the voice quality conversion device 3 to match the pitch range of the voice quality conversion model. First, the conversion device 1 identifies the pitch range of the character of the voice quality conversion model and the pitch range of the speaker P. The conversion device 1 converts the pitch range of the utterance data 14 of the speaker P into the pitch range of the character of the voice quality conversion model, and inputs the utterance data after pitch range conversion to the voice quality conversion device 3.

[0022] Data in which the pitch range of the utterance data 14 of the speaker P is converted into the pitch range of the character is input to the voice quality conversion device 3. The voice quality conversion device 3 outputs data obtained by converting the input data into the voice quality of the character using the voice quality conversion model. The voice quality conversion device 3 outputs data in which the utterance of the speaker P has the same content, speed, and pitch of the voice of the speaker P, and the voice quality is converted into the voice quality of the voice quality conversion model.

[0023] In this way, the conversion device 1 inputs the speech data converted to the pitch range of the voice quality to be converted to the voice quality conversion device 3. As a result, even if the voice quality of speaker P and the voice quality converted by the voice quality conversion model are not similar, the conversion system 7 can output speech in which the voice quality of speaker P's utterances has been naturally converted to the voice quality of the voice quality conversion model.

[0024] The conversion method using the conversion system 7 will be explained with reference to Figure 2.

[0025] In step S1, the conversion device 1 obtains the vocal range of the character in the voice conversion model used by the voice conversion device 3, and the vocal range of speaker P. For example, the conversion device 1 obtains the character identifier specified by speaker P and obtains the vocal range of the specified character identifier from the character vocal range data 12. The conversion device 1 has speaker P read a sample sentence or the like to obtain speaker P's speech data, and obtains speaker P's vocal range from that speech data.

[0026] In step S2, the conversion device 1 acquires the speech data 14 of speaker P. The speech data 14 acquired here is the target of conversion by the voice quality conversion model.

[0027] In step S3, the conversion device 1 removes the unwanted data 15 from the speech data 14. The unwanted data 15 is data that is not subject to conversion by the voice conversion model used by the voice conversion device 3. The unwanted data 15 includes, for example, laughter, growls, sneezes, etc., which are sounds made by a person but are not spoken voices.

[0028] In step S4, the conversion device 1 outputs the unwanted data 15, which has been removed from the speech data 14, to the voice quality conversion device 3.

[0029] In step S5, the conversion device 1 converts the pitch range of the speech data 16, from which the irrelevant data 15 has been removed, to the pitch range of the character acquired in step S1, thereby generating speech data 17 after pitch range conversion. The conversion device 1 converts the pitch range using methods such as shift conversion or pitch range compression.

[0030] In step S6, the conversion device 1 outputs the speech data 17 generated in step S5 after pitch range conversion to the voice quality conversion device 3.

[0031] In step S7, the voice conversion device 3 converts the speech data 17, which has been converted to a different pitch range and input from the conversion device 1, into the character's voice using a voice conversion model. In step S8, the voice conversion device 3 synthesizes the data that was not relevant in step S4, which was input in step S4, with the data that has been converted into the character's voice, and plays it back.

[0032] The process shown in Figure 2 is just one example and is not limited to this.

[0033] For example, the conversion system 7 described herein describes a case in which a conversion device 1 converts the pitch range of a set of speech data 14 input by speaker P, and a voice quality conversion device 3 converts the voice quality of the data whose pitch range has been converted by the conversion device 1, but is not limited to this case.

[0034] The utterance of speaker P and the voice conversion of that utterance may be performed in real time. For example, the conversion device 1 sequentially acquires utterance data 14 while speaker P is uttering and identifies the unit data to be processed from the utterance data 14. For each unit data, the conversion device 1 sequentially repeats the processing from steps S3 to S8. The unit data is a single data item obtained by dividing the utterance data 14 into predetermined units such as phonemes, syllables, or moras.

[0035] This allows the conversion system 7 to perform speaker P's utterance and voice quality conversion of that utterance in real time. By performing voice quality conversion in real time, the conversion system 7 can be used for real-time processing such as voice chat with other people.

[0036] The conversion system 7 described herein involves a conversion device 1 converting the pitch range of speaker P's speech data 14 to match the pitch range of a voice quality conversion model, and outputting the converted data to a voice quality conversion device 3. As a result, the voice quality conversion device 3 only needs to convert the voice quality of the data in the pitch range that matches the pitch range of the voice quality conversion model character, thus enabling accurate voice quality conversion and outputting natural-sounding speech data.

[0037] Furthermore, sounds that are produced by a person but are not spoken voices, such as the excluded data 15, can be a factor in degrading the accuracy of conversion by the voice quality conversion model. Therefore, in the conversion system 7 according to this disclosure, the conversion device 1 removes the excluded data 15 from the speech data 14 and outputs the excluded data 15 to the voice quality conversion device 3, and also outputs speech data 17 with a converted pitch range, obtained by converting the pitch range of the speech data 14, to the voice quality conversion device 3. The voice quality conversion device 3 converts the pitch range converted speech data 17 input from the conversion device 1 using the voice quality conversion model, and then superimposes the excluded data 15 input from the conversion device 1 to reproduce it. In this way, the conversion system 7 can reproduce each of the speaker P's utterances while ensuring the accuracy of voice quality conversion by the voice quality conversion model.

[0038] (Converter) The conversion device 1 relating to this disclosure will be explained with reference to Figure 3.

[0039] As shown in Figure 3, the conversion device 1 includes the following data: pitch range data 11, character pitch range data 12, character identifier 13, speech data 14, excluded data 15, removed speech data 16, and speech data 17 after pitch range conversion, as well as the functions of acquisition unit 21, removal unit 22, and conversion unit 23. Each data is stored in a storage device such as memory 902 or storage 903. Each function is implemented in the CPU 901.

[0040] The vocal range data 11 is the vocal range data of speaker P. For example, speaker P is made to speak by having them read a sample sentence beforehand. The conversion device 1 identifies speaker P's vocal range from their speech and sets it in the vocal range data 11. The vocal range data 11 identifies the lowest and highest notes as speaker P's vocal range, for example, as shown in Figure 4. The vocal range data 11 may also be set by speaker P. Alternatively, the vocal range data 11 may be set based on an estimated vocal range from speaker P's attribute information such as age, gender, and height.

[0041] Character vocal range data 12 is data that identifies the vocal range of the voice conversion model. The vocal range of the voice conversion model is the range that can guarantee the quality of the voice conversion.

[0042] If there is a voice conversion model for each of the multiple characters, the vocal range of each character is identified. The character vocal range data 12 associates the lowest and highest notes with the character identifier, for example, as shown in Figure 5.

[0043] The character identifier 13 identifies the character of one of the multiple voice conversion models used by the voice conversion device 3.

[0044] The speech data 14 is the data of the utterance of the original speaker P. The speech data 14 has a computer-processable data format.

[0045] The excluded data 15 consists of sound data from the speech data 14 that does not represent spoken voice. The excluded data 15 is generated by the removal unit 22.

[0046] The removed speech data 16 is the data remaining after the unwanted data 15 has been removed from the speech data 14. The removed speech data 16 is the data from the speech data 14 that is subject to processing by the conversion unit 23. The removed speech data 16 is generated by the removal unit 22.

[0047] The speech data 17 after pitch range conversion is data obtained by converting the pitch range of the removed speech data 16 to the pitch range of the voice quality conversion model. The speech data 17 after pitch range conversion is generated by the conversion unit 23.

[0048] The acquisition unit 21 acquires the vocal range of the voice quality conversion model, the vocal range of speaker P, and the speech data 14 of speaker P.

[0049] The vocal range of the voice conversion model is obtained, for example, by speaker P specifying a character of the voice conversion model. In this disclosure, the conversion device 1 stores the vocal range of each of the multiple vocal conversion model characters used by the voice conversion device 3 as character vocal range data 12. The acquisition unit 21 identifies one character identifier 13 from among the multiple characters defined in the character vocal range data 12 through operations by speaker P, etc., and obtains the vocal range corresponding to the identified character identifier 13 from the character vocal range data 12.

[0050] The acquisition unit 21 acquires the speaker P's vocal range in advance as vocal range data 11.

[0051] The acquisition unit 21 acquires the vocal range of the voice conversion model and the vocal range of speaker P, and then acquires speaker P's speech data 14. The acquisition unit 21 has speaker P speak the content that it wants to be converted into the voice conversion model, and then acquires speaker P's speech data 14.

[0052] The removal unit 22 removes excluded data 15 from the speaker P's speech data 14, which are not subject to conversion by the voice quality conversion model. The removal unit 22 identifies excluded data 15 from the speech data 14, such as laughter, growls, and sneezes, which are sounds made by a person but not spoken language. The removal unit 22 deletes the excluded data 15 from the speech data 14 to generate the speech data 16 after the exclusion.

[0053] The removal unit 22 further outputs the unwanted data 15 to the voice conversion model of the voice conversion device 3. The unwanted data is then combined with the data converted by the voice conversion model.

[0054] The conversion unit 23 converts the pitch range of speaker P's speech data 14 to the pitch range of the voice quality conversion model to generate speech data 17 after pitch range conversion. The conversion unit 23 outputs the speech data 17 after pitch range conversion to the voice quality conversion model of the voice quality conversion device 3.

[0055] The conversion unit 23 shifts the vocal range of speaker P in the speech data 14 to the vocal range of the voice quality conversion model. For example, the conversion unit 23 calculates the shift amount from a reference value of speaker P's vocal range and a reference value of the voice quality conversion model's vocal range. The reference value is, for example, the median of the vocal range. The conversion unit 23 shifts the vocal range of the speech data 14 according to the calculated shift amount.

[0056] Consider the case where the vocal range of the model to be converted is C4-C5, speaker P is male, and speaker P's vocal range is C3-E4. Since speaker P's vocal range is lower than the vocal range of the model to be converted, a pitch correction of +12 (1 octave) is applied. As a result, the vocal range of speaker P's speech data 14 is converted from C3-E4 to C4-E5.

[0057] In this disclosure, we will explain the case where the shift amount is calculated from the speaker P's vocal range, but the shift amount may also be set by speaker P.

[0058] The conversion unit 23 compresses the pitch range of the shift-converted speech data 14 that falls outside the pitch range of the voice quality conversion model to the pitch range of the voice quality conversion model.

[0059] The pitch range after shift conversion may fall outside the pitch range of the voice conversion model. To bring the pitch range after shift conversion within the pitch range of the voice conversion model, a certain range of pitches that falls outside the pitch range of the voice conversion model is pushed into a certain upper and lower range of the voice conversion model's pitch range.

[0060] Figure 6 shows an example of compression by the conversion unit 23. The vocal range of speaker P after shift conversion is wider in the high and low ranges compared to the vocal range of the voice quality conversion model. Therefore, the compressed vocal range is set to two semitones above and below the vocal range of the voice quality conversion model. The conversion unit 23 pushes the vocal range that falls outside the vocal range of the voice quality conversion model into the compressed vocal range. Here, the arrangement of each note within the compressed vocal range is processed using moving averages or exponential moving averages to smooth the changes in pitch.

[0061] The compressed frequency range may be determined by the number of sounds that fall outside the frequency range of the voice conversion model within the shifted frequency range of speaker P. The compressed frequency range may be pre-set for each voice conversion model. The compressed frequency range may also be determined by adjusting the pre-set frequency range for the voice conversion model using the number of sounds that fall outside the frequency range of the voice conversion model.

[0062] If the speech data 14 contains unwanted data 15, the conversion unit 23 converts the pitch range of the data from which the unwanted data 15 has been removed from the speech data 14 to the pitch range of the voice quality conversion model, thereby generating speech data 17 after pitch range conversion. The conversion unit 23 outputs the speech data 17 after pitch range conversion to the voice quality conversion model of the voice quality conversion device 3.

[0063] In this disclosure, the conversion unit 23 has described a case in which, after shifting the pitch range, there are pitch ranges in the shifted speech data that fall outside the pitch range of the voice quality conversion model, and the conversion unit 23 compresses them. However, the disclosure is not limited to this case. If the pitch range of speaker P and the pitch range of the voice quality conversion model are the same or close, the conversion unit 23 may compress the pitch range of speaker P in the speech data 14 that falls outside the pitch range of the voice quality conversion model to the pitch range of the voice quality conversion model without using shift conversion. For example, if the reference value of speaker P's pitch range and the reference value of the voice quality conversion model's pitch range are the same or within a predetermined value, the conversion unit 23 may perform only pitch range compression without shift conversion.

[0064] In this disclosure, the conversion unit 23 is described in the case of converting the pitch range, but is not limited to this. If the pitch range of the speaker P's speech data 14 is within the pitch range of the voice quality conversion model, the conversion unit 23 does not need to convert the pitch range of the speech data 14.

[0065] The conversion process by the conversion unit 23 will be explained with reference to Figure 7.

[0066] In step S101, the conversion unit 23 obtains speech data 16 after removing the unwanted data 15. In step S102, the conversion unit 23 performs a shift conversion on the pitch range of the speech data 16 after removal.

[0067] In step S103, the conversion unit 23 determines whether there are any vocal ranges that fall outside the vocal range of the voice quality conversion model after the shift conversion in step S102. If there are vocal ranges that fall outside the range, the process proceeds to step S104. If there are no vocal ranges that fall outside the range, the process proceeds to step S105.

[0068] In step S104, the conversion unit 23 compresses the frequency ranges that fall outside the range of the voice conversion model into the range of the voice conversion model.

[0069] The conversion unit 23 generates speech data 17 after pitch range conversion by shift conversion in step S102, or by shift conversion in step S102 and compression in step S104. In step S105, the conversion unit 23 outputs the speech data 17 after pitch range conversion to the voice quality conversion device 3.

[0070] The conversion system 7 described herein converts the vocal range of speaker P's speech data into the vocal range of the voice conversion model before inputting it into the voice conversion model. Therefore, it can convert to a natural voice regardless of the original voice quality.

[0071] (modified version) Referring to Figure 8, a modified example of the conversion system 7a will be described. The conversion system 7a shown in Figure 8 differs from the conversion system 7 shown in Figure 1 in that it includes a practice device 5.

[0072] Generally, voice conversion models can be difficult for beginners to use. This is because they may not know whether the microphone is on, whether they are speaking at the appropriate volume, what pitch is necessary to convert to a natural voice, or what speaking style will make the target character speak naturally.

[0073] Therefore, the practice device 5 supports speaker P's practice so that they can master the voice quality conversion model.

[0074] In the conversion system 7a, the conversion device 1 converts the speaker P's vocal range to the vocal range of the character in the voice quality conversion model, and the voice quality conversion device 3 converts the utterance converted to the character's vocal range to the character's voice quality. The practice device 5 is used for speaker P to practice the character's utterances other than the character's vocal range and voice quality.

[0075] The practice device 5 includes dialogue data 51, personality data 52, speech data 53, and vocal range data 54, as well as the functions of a display unit 61, an acquisition unit 62, a playback unit 63, and a specific unit 64. Each piece of data is stored in a storage device such as memory 902 or storage 903. Each function is implemented in the CPU 901.

[0076] Dialogue data 51 identifies characteristic lines of dialogue used in the voice quality conversion model. Dialogue data 51 contains one or more lines of dialogue. Dialogue data 51 may also include symbols that identify the intonation, accent, and other nuances of speech when the character speaks those lines.

[0077] Personality data 52 identifies the character's personality or attributes. Personality data 52 identifies the personality or attributes that influence the character's speech. Personality data 52 may be the character's personality, attributes, or data containing attributes, or it may be data that identifies events that describe the character's personality or attributes.

[0078] Speech data 53 is data from an utterance by speaker P.

[0079] The vocal range data 54 is data that identifies the vocal range of speaker P. The vocal range data 54 is, for example, the data shown in Figure 4. The vocal range data 54 is generated by the identification unit 64.

[0080] The display unit 61 displays a message to the speaker P, the source of the voice quality conversion model, prompting them to utter the dialogue data 51. The display unit 61 further displays personality data 52. The display unit 61 can prompt speaker P to utter the character's lines after understanding the character's personality or attributes. The conversion device 1 and conversion unit 23 convert speaker P's vocal range and voice quality to the character's vocal range and voice quality. The conversion system 7a makes it easier to convert speaker P's utterances into character utterances.

[0081] The acquisition unit 62 acquires the speech data 53 of speaker P.

[0082] The playback unit 63 plays back the acquired speech data 53 of speaker P. By playing back the speaker's speech data 14, speaker P can check their own speech and practice to make it sound more like the character's speech.

[0083] The playback unit 63 may also play back the speech data 53 of speaker P, as well as speech data converted from the speech data 14 to the character's vocal range and voice quality by the conversion device 1 and the voice quality conversion device 3. The playback unit 63 may further play back the speech of the character of the voice quality conversion model.

[0084] The playback unit 63 may repeatedly play back the speech data 53 from speaker P, the speech data 14 from which the pitch range and voice quality have been converted, and the character's speech. Speaker P can compare each speech and analyze them to make them closer to the character's speech.

[0085] The playback unit 63 may not only play back each utterance data, but also analyze the speaker P's utterance data 14 to analyze characteristics such as speaking speed, volume, and intonation. The playback unit 63 may also display differences between the characteristics of the speaker P's utterance and those of the character as items that need to be corrected. Furthermore, if the characteristics of speaker P's utterance data 14 and those of the character's utterance are identical or within a predetermined value, the playback unit 63 may play back audio data in which the character praises the speaker's utterance.

[0086] The identification unit 64 identifies the vocal range of speaker P from the speech data 53 and generates vocal range data 54. The vocal range data 54 is transmitted to the conversion device 1 and is referenced when converting the vocal range of speaker P's speech data. The identification unit 64 may also identify the shift amount when shifting the vocal range and share the identified shift amount with the conversion device 1. The identification unit 64 may also identify the compressed vocal range when compressing the vocal range and share the identified compressed vocal range with the conversion device 1. In addition, the identification unit 64 may identify parameters in vocal range conversion or voice quality conversion and share them with the conversion device 1 or voice quality conversion device 3.

[0087] The identification unit 64 may identify a microphone gain suitable for speaker P, in addition to the frequency range, and set it to the identified gain. The identification unit 64 may also analyze speaker P's speech data 14 to identify the ratio of mixing speaker P's voice with the voice after frequency range and voice quality conversion, and the amount of noise cancellation. The identified ratio and noise cancellation amount are shared and referenced by the conversion device 1 or the voice quality conversion device 3.

[0088] When speaker P selects a voice conversion model to which they wish to change their voice, the practice device 5 displays a line of dialogue that matches the selected voice conversion model and plays a sample audio file of the dialogue being read aloud in the target voice. The device converts the pitch range and voice quality of speaker P's speech as they read the dialogue and records it, then plays back the recorded audio. The practice device 5 may also provide opportunities to repeat this practice multiple times.

[0089] The practice method relating to this disclosure will be explained with reference to Figure 10.

[0090] In step S51, the training device 5 displays the dialogue data 51. In step S52, the training device 5 displays the personality data 52. At this time, the training device 5 may also play back the audio of the character speaking the lines from the dialogue data 51.

[0091] In step S53, the training device 5 acquires speech data 14 from speaker P. In step S54, the training device 5 may play back the speech data 14. The training device 5 may also play back the audio of a character speaking the lines from the dialogue data 51. The training device 5 may play back speaker P's speech data 53 and the audio from the character simultaneously. The training device 5 may play back speaker P's speech data 53 and then the audio from the character, or play back the audio from the character and then play back speaker P's speech data 53.

[0092] The practice device 5 also identifies the vocal range of speaker P from the speech data 53. The identification unit 64 may also identify parameters used for converting vocal range or voice quality, in addition to the vocal range of speaker P. The practice device 5 transmits the vocal range data and parameters to the conversion device 1 or the voice quality conversion device.

[0093] The training device 5, which relates to this modified version, provides an opportunity for speaker P to practice so that their speech data 53 approaches the character's speech. Speaker P can become familiar with using the voice conversion model and practice to make their speech sound more like the character's. The training device 5 can contribute to the widespread adoption of voice conversion models.

[0094] (Technology A) A memory device that stores dialogue data to identify characteristic lines of a character used in the conversion of a voice quality model, A display unit that displays a message prompting the speaker of the voice conversion model to speak the dialogue data. A training device equipped with the following features.

[0095] (Technology B) The display unit further displays personality data that identifies the character's characteristics or attributes. The training device described in Technical A.

[0096] The devices described above—the conversion device 1, the voice quality conversion device 3, and the practice device 5—are each part of a general-purpose computer system comprising, for example, a CPU (Central Processing Unit, processor) 901, memory 902, storage 903 (HDD: Hard Disk Drive, SSD: Solid State Drive), communication device 904, input device 905, and output device 906. In this computer system, the CPU 901 executes a program loaded onto the memory 902, thereby realizing the respective functions of each device.

[0097] Each device may be implemented on a single computer, or on multiple computers. Furthermore, each device may be a virtual machine implemented on a computer.

[0098] The programs for each device are stored on HDD, SSD, USB (Universal Serial Bus) memory, and CD (Compact). It can be stored on computer-readable recording media such as Discs and DVDs (Digital Versatile Discs), or distributed over a network.

[0099] Any part or all of the functional components described in this disclosure may be implemented by program. The programs referred to in this disclosure may be distributed non-temporarily on a computer-readable recording medium, distributed via communication lines such as the Internet (including wireless communication), or distributed installed on any terminal. While those skilled in the art may conceive of additional effects and various modifications of the present invention based on the above description, the aspects of this disclosure are not limited to the individual embodiments described above. Various additions, modifications, and partial deletions are possible without departing from the conceptual idea and spirit of the present invention derived from the claims and their equivalents. For example, what is described in this disclosure as a single device (or component, hereinafter the same) (including what is depicted as a single device in the drawings) may be implemented by multiple devices. Conversely, what is described in this disclosure as multiple devices (including what is depicted as multiple devices in the drawings) may be implemented by a single device. Alternatively, some or all of the means or functions included in one device (e.g., a server) may be included in another device (e.g., a user terminal). Furthermore, a "system" may consist of one device, or it may consist of two or more devices (for example, a server and a user terminal, or multiple user terminals).

[0100] Furthermore, not all matters described in this disclosure are mandatory requirements. In particular, matters described in this disclosure but not in the claims can be considered optional additional matters.

[0101] It should be noted that the applicant is only aware of the prior art inventions described in the "Prior Art Documents" section of this disclosure, and this disclosure is not necessarily intended to solve the problems described in those prior art inventions. The problems that this disclosure aims to solve should be determined by considering the disclosure as a whole. For example, if this disclosure describes that a certain effect is achieved by a specific configuration, it can also be said that the problem that is the inverse of that specified effect is solved. However, this does not necessarily mean that such a specific configuration is an essential requirement. [Explanation of Symbols]

[0102] 1. Conversion device 3. Voice quality conversion device 5 Exercise equipment 7 Conversion System 11.54 Vocal Range Data 12 Character Vocal Range Data 13. Identifier 14, 53 speech data 15. Excluded data 16. Speech data after removal 17. Speech data after pitch range conversion 21, 62 Acquisition Department 22 Removal part 23 Conversion section 51 Dialogue Data 52 Personality Data 61 Display section 63 Playback Department 64 Specific part 901 CPU 902 memory 903 Storage 904 Communication equipment 905 Input device 906 Output device P Speaker

Claims

1. The voice quality conversion model's vocal range and the speaker's vocal range are acquired by the acquisition unit. A conversion unit that converts the pitch range of the speaker's speech data to the pitch range of the voice conversion model, and outputs the speech data after pitch range conversion to the voice conversion model, A conversion device equipped with the following features.

2. The conversion unit shifts the speaker's vocal range in the speech data to the vocal range of the voice quality conversion model. The conversion device according to claim 1.

3. The conversion unit compresses the range of the speech data after shift conversion that falls outside the range of the voice conversion model to the range of the voice conversion model. The conversion device according to claim 2.

4. The conversion unit compresses the vocal range of the speaker in the speech data that falls outside the vocal range of the voice conversion model to the vocal range of the voice conversion model. The conversion device according to claim 1.

5. The system further includes a removal unit that removes data from the speaker's speech data that is not subject to conversion by the voice quality conversion model, The conversion unit converts the pitch range of the data from which the unwanted data has been removed from the speech data into the pitch range of the voice quality conversion model. The conversion device according to claim 1.

6. The removal unit outputs the excluded data to the voice quality conversion model. The data converted by the voice quality conversion model is then combined with the excluded data. The conversion device according to claim 5.

7. Computers The vocal range of the voice conversion model and the speaker's vocal range are obtained. The pitch range of the speaker's speech data is converted to the pitch range of the voice conversion model, and the speech data after pitch range conversion is output to the voice conversion model. Conversion method.

8. Computers, The voice quality conversion model's vocal range and the speaker's vocal range are acquired by the acquisition unit. A conversion unit that converts the pitch range of the speaker's speech data to the pitch range of the voice conversion model, and outputs the speech data after pitch range conversion to the voice conversion model. A program that makes it function as such.