Voice providing device, voice providing method, and program

The voice providing device addresses individual differences in user reactions by offering personalized model voices based on self-evaluation, enhancing speaking practice efficiency through tailored voice output.

WO2025158663A1PCT designated stage Publication Date: 2025-07-31NT T INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/002493
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing voice practice systems fail to account for individual differences in user reactions to their own voice, leading to inconsistent practice effectiveness due to factors like personality and listening habits, making a universal approach ineffective for all users.

Method used

A voice providing device that determines a user's self-evaluation value and, based on this, provides a model voice tailored to the user's comfort level by outputting a teacher's voice of the same sex or a synthesized voice with the user's voice quality, addressing individual differences in user reactions.

Benefits of technology

Enhances practice effectiveness by providing personalized model voices that cater to users' comfort levels, improving speaking practice efficiency by considering personality traits and listening habits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024002493_31072025_PF_FP_ABST
    Figure JP2024002493_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The objective of the present invention is to achieve a greater practice effect by solving the problem that each individual user has a different response to the user's own voice. To this end, the present invention provides a voice providing device for providing a model voice, the device comprising: a voice personalization unit 22 that determines whether a self-evaluation value indicating a user's skill level in speaking according to the user is equal to or greater than a first threshold value, and outputs a voice of a teacher of the same sex as the user as the model voice when the self-evaluation value is less than the first threshold value; and a voice synthesis unit 11 that outputs a voice combining the voice quality of the user and the utterance style of the teacher when the self-evaluation value is equal to or greater than the first threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Audio providing device, audio providing method, and program

[0001] The present disclosure relates to a technology for providing (proposing) model voices that users can use to practice speaking techniques for presentations and the like.

[0002] In recent years, research has been conducted into improving speaking practice performance by listening to model voices. Among these, it has been suggested that listening to a user's own voice as a model voice can improve foreign language speaking ability. For example, Non-Patent Document 1 shows that using a user's own voice is effective in practicing foreign language pronunciation. Research is also being conducted on tools that support speaking practice, not only for foreign language pronunciation but also for native languages. For example, Non-Patent Document 2 shows the effectiveness of speaking practice using teaching materials that allow users to listen to model voices. Even in such cases, listening to the user's own voice as a model voice can make the user more aware of the difference between the model voice and their current situation, which is expected to improve the effectiveness of practice.

[0003] Ding, Shaojin, et al. "Golden speaker builder—An interactive tool for pronunciation training." Speech Communication 115 (2019): 51-66. Hirano, Miho, and Shibata, Yoshiaki. "Evaluation of individualized speaking learning materials focusing on paralinguistic skills." Lifelong Learning and Career Education Research 11 (2015): 47-51.

[0004] However, depending on the user's personality and habits, some users may feel uncomfortable listening to a recording of their own voice, so one's own voice may not necessarily be suitable for all users. Thus, in the past, due to individual differences in users' reactions to their own voices, it was not possible to expect a high level of practice effect.

[0005] The present disclosure has been made in consideration of the above circumstances, and aims to achieve a higher practice effect by resolving individual differences in reactions to a user's own voice.

[0006] In order to achieve the above object, the present disclosure provides a voice providing device that provides model voice, and includes: a voice personalization unit that determines whether a self-evaluation value indicating the level of skill a user has in their speaking style is equal to or greater than a first threshold, and if the self-evaluation value is less than the first threshold, outputs the voice of a teacher of the same gender as the user as the model voice; and a voice synthesis unit that outputs voice that combines the user's voice quality with the teacher's speaking style if the self-evaluation value is equal to or greater than the first threshold.

[0007] As described above, according to the present disclosure, by resolving individual differences in reactions to the user's own voice, it is possible to expect a higher practice effect.

[0008] FIG. 1 is a diagram illustrating an example of a hardware configuration of a voice providing device according to an embodiment. FIG. 2 is a diagram illustrating a functional configuration of a voice providing device according to a first embodiment. FIG. 3 is a flowchart illustrating a voice providing method according to the first embodiment. FIG. 4 is a diagram illustrating a functional configuration of a voice providing device according to a second embodiment. FIG. 5 is a flowchart illustrating a voice providing method according to the second embodiment. FIG. 6 is a diagram illustrating a functional configuration of a voice providing device according to a third embodiment. FIG. 7 is a flowchart illustrating a voice providing method according to the third embodiment. FIG. 8 is a diagram illustrating a functional configuration of a voice providing device according to a fourth embodiment. FIG. 9 is a flowchart illustrating a voice providing method according to the fourth embodiment.

[0009] Overview of the Voice Providing Device First, an overview of the voice providing device according to the embodiment will be described.

[0010] Factors that influence differences in a user's reaction to their own voice include personality traits such as self-esteem and extroversion, and habits of listening to their own voice. In this embodiment, the presentation practice tool is equipped with self-evaluation items, a function for viewing and reviewing one's own presentation, and a function for chatting with other practice participants and instructors, etc., and these are used to infer the user's personality traits and habits and determine whether to present a model voice using the user's own voice. This makes it possible to use a model voice using the user's own voice only for people who do not have a negative reaction to the user's voice, which is thought to enable more efficient speaking practice. Note that the voice providing device 10 of the first embodiment described below has basic functions that are premised on addressing individual differences in reactions to the user's own voice.

[0011] [Hardware Configuration of the Voice Providing Device] Next, the electrical hardware configuration of the voice providing device will be described with reference to Fig. 1. Fig. 1 is a diagram showing the electrical hardware configuration of the voice providing device. Note that the configuration shown in Fig. 1 is common to the voice providing device 10 according to the first embodiment, the voice providing device 20 according to the second embodiment, the voice providing device 30 according to the third embodiment, and the voice providing device 40 according to the fourth embodiment.

[0012] The computer serving as the audio providing device in FIG. 1 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected by a bus 1010.

[0013] The program that realizes the processing on the computer is provided by a recording medium 1001, such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via the communication network 100. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.

[0014] The memory device 1003 reads and stores the program from the auxiliary storage device 1002 when an instruction to start the program is received. The CPU 1004 realizes the functions related to the device in accordance with the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a communication network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.

[0015] First Embodiment Next, a first embodiment will be described with reference to FIGS. 2 and 3. FIG.

[0016] In the first embodiment, the recorded voice of the user and the recorded voice of the teacher who is the model are used as inputs to the voice synthesis unit 11, and a model voice tailored to the user is output.

[0017] [Functional Configuration of First Embodiment] The functional configuration of the first embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing the functional configuration of the voice providing device according to the first embodiment. As shown in Fig. 2, the voice providing device 10 according to the first embodiment has a voice synthesis unit 11. This is a function or means realized by instructions from the CPU 1004 shown in Fig. 1. In addition, a user's recorded voice DB (Data Base) 91 and a teacher's recorded voice DB 92 are constructed in the recording medium 1001, the auxiliary storage device 1002, or the memory device 1003.

[0018] Data of the user's recorded voice is stored in the user's recorded voice DB 91. The user's recorded voice is the voice of the user requesting that several dozen sentences, each about several seconds long, be read aloud before starting to practice the content that the user wants to practice, and the content of the sentences is recorded.

[0019] Data on recorded voices of model teachers is stored in the teacher's recorded voice DB 92. The recorded voices of teachers are preferably those of professional lecturers or other people with an established reputation for their speaking style.

[0020] The user's recorded voice DB 91 and the teacher's recorded voice DB 92 may be constructed in a database server or the like outside the voice providing device 10. In this case, the voice providing device 10 accesses the database server or the like to transmit and receive data to and from the recorded voice DB 91 or the teacher's recorded voice DB 92.

[0021] The speech synthesis unit 11 acquires the user's recorded speech from the user's recorded speech DB 91 and the teacher's speech from the teacher's recorded speech DB 92. By inputting the user's speech and the teacher's speech, the speech synthesis unit 11 generates and outputs a model speech that has the "user's voice quality and the teacher's speaking style." As a speech synthesis method, for example, a speech synthesis technology such as that shown in Reference 1 can be used.

[0022] <Reference 1> Fujita, Kenichi, et al. "Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model." arXiv preprint arXiv:2304.11976 (2023). [Processing of First Embodiment] A speech provision method according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing the speech provision method according to the first embodiment.

[0023] S11: The voice synthesis unit 11 synthesizes the user's voice quality v based on the user's recorded voice acquired from the user's recorded voice DB 91 and the teacher's voice acquired from the teacher's recorded voice DB 92. u and the teacher's speaking style t A voice that combines both voices is a model voice v e Let's say.

[0024] S12: The voice synthesis unit 11 receives the sample voice v e Output.

[0025] Effect of First Embodiment As described above, the voice providing device 10 according to the first embodiment can provide a voice that combines the user's voice quality with the teacher's speaking style as a model voice.

[0026] Second Embodiment Next, a second embodiment will be described with reference to FIGS.

[0027] In the first embodiment, a method for proposing a model voice using the user's own voice was described. In the first embodiment, it was assumed that using the user's own voice is effective for everyone. However, there is a possibility that some users find their own voice unpleasant, which may hinder the efficiency of practice.

[0028] Experiments have shown that factors that affect individual differences in users' impression evaluations of their own voices include personality traits such as self-esteem or extroversion, or habits of listening to one's own voice. Specifically, approximately 200 participants were asked to record their own voices, listen to multiple voices including their own recordings, and evaluate their impressions of the voices, such as attractiveness and familiarity, and the relationship between the impression evaluation results and the participants' personal characteristics (age, gender, personality, habits, values, etc.) was analyzed.

[0029] Based on the experimental results, in the second embodiment, the user is asked to self-evaluate their own speaking style, and the level of self-affirmation is examined based on the level of the evaluation, thereby taking into consideration individual differences.

[0030] [Functional Configuration of Second Embodiment] The functional configuration of the second embodiment will be described with reference to Fig. 4. Fig. 4 is a diagram showing the functional configuration of a voice providing device according to the second embodiment. Note that in Fig. 4, the same functions as those in Fig. 2 according to the first embodiment are denoted by the same reference numerals, and their description will be omitted.

[0031] As shown in Fig. 4, the voice providing device 20 according to the second embodiment has a voice synthesis unit 11, a user information acquisition unit 21, and a voice personalization unit 22. These are functions or means realized by instructions from the CPU 1004 shown in Fig. 1. In addition, a user's recorded voice DB 91, a teacher's recorded voice DB 92, and a user information DB 93 are constructed in the recording medium 1001, the auxiliary storage device 1002, or the memory device 1003.

[0032] The user's recorded voice DB 91, the teacher's recorded voice DB 92, and the user information DB 93 may be constructed in a database server or the like outside the voice providing device 20. In this case, the voice providing device 20 accesses the database server or the like to transmit and receive data to and from the recorded voice DB 91, the teacher's recorded voice DB 92, or the user information DB 93. In addition, the teacher's recorded voice DB 92 in the second embodiment stores data on recorded voices by male teachers and recorded voices by female teachers.

[0033] The user information acquisition unit 21 acquires user information (the user's recorded voice, the user's self-evaluation value (score), and the user's gender) to be used for personalization, and temporarily stores the information in the user information DB 93. As in the first embodiment, the user information acquisition unit 21 acquires the user's recorded voice from the user's recorded voice DB 91, and also acquires, from a questionnaire, information on the user's gender and a "self-evaluation value" that indicates the level of skill (ability) the user has in their speaking style.

[0034] The user information DB 93 temporarily stores the user information acquired by the user information acquisition unit 21 .

[0035] The voice personalizing unit 22 determines self-affirmation based on the user's self-evaluation value from the user information (the user's recorded voice, the user's self-evaluation value, and the user's gender) acquired from the user information DB 93. The voice personalizing unit 22 determines that the self-evaluation value is high if it is equal to or greater than a threshold T1, and that the self-affirmation value is low if it is less than the threshold T1. Note that the threshold T1 is an example of a first threshold.

[0036] If the self-evaluation value is less than the threshold T1, the voice personalization unit 22 outputs voice data of a teacher of the same gender from the teacher's recorded voice DB as a model voice based on the user's gender information.

[0037] On the other hand, when the self-evaluation value is equal to or greater than the threshold T1, and the self-esteem is high, the voice synthesis unit 11 uses the user's recorded voice and a recorded voice of a teacher of the same sex as the user obtained from the teacher's recorded voice DB 92 to output a voice having both the user's voice quality and the teacher's speaking style as a model voice. Note that in this case, the voice synthesis unit 11 may also use a recorded voice of a teacher of the opposite sex to the user.

[0038] [Processing of Second Embodiment] A sound providing method according to the second embodiment will be described with reference to Fig. 5. Fig. 5 is a flowchart showing the sound providing method according to the second embodiment.

[0039] S21: The user information acquisition unit 21 acquires the user's recorded voice from the user's recorded voice DB 91 as user information, acquires the self-evaluation value and the user's gender from the questionnaire, and temporarily stores them in the user information DB 93.

[0040] S22: The voice personalization unit 22 determines (judges) whether the user's self-evaluation value among the user information (user's recorded voice, user's self-evaluation value, user's gender) obtained from the user information DB 93 is greater than or equal to the threshold value T1.

[0041] S23: If the self-evaluation value is equal to or greater than the threshold value T1 (S22; YES), the voice synthesis unit 11 calculates the user's voice quality v based on the user's recorded voice acquired from the user's recorded voice DB 91 and the voice of a teacher of the same or opposite sex as the user acquired from the teacher's recorded voice DB 92. u and the teacher's speaking style t A voice that combines both voices is a model voice v e Let's say.

[0042] S24: On the other hand, if the self-evaluation value is less than (not greater than) the threshold value T1 (S22; NO), the voice personalization unit 22 selects the voice of a teacher of the same gender as the user as the model voice v e Let's say.

[0043] S25: The voice synthesis unit 11 or the voice personalization unit 22 receives the model voice v e Output.

[0044] As described above, the voice providing device 20 according to the second embodiment switches between outputting a voice that combines the user's voice quality and the teacher's speaking style as a model voice, or outputting a voice of a teacher of the same gender as the user as a model voice, depending on the user's self-esteem determined based on a "self-evaluation value" that indicates the level of skill (ability) of the user's speaking style. This provides the effect of achieving a higher practice effect by addressing individual differences in reactions to the user's own voice.

[0045] Third Embodiment Next, a third embodiment will be described with reference to FIGS.

[0046] In the second embodiment, the voice providing device 20 determines the self-affirmation level of the user from the self-evaluation value. However, it does not distinguish whether the user's original skill level is high or whether the user's self-evaluation level is high. Therefore, the voice providing device 30 according to the third embodiment calculates the difference between the evaluation value of the user's speaking skill (e.g., evaluation by a lecturer) and the self-evaluation, and determines the self-affirmation level from the difference.

[0047] [Functional Configuration of Third Embodiment] The functional configuration of the third embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing the functional configuration of a voice providing device according to the third embodiment. Note that in Fig. 6, the same functions as those in Fig. 2 according to the first embodiment and Fig. 4 according to the second embodiment are denoted by the same reference numerals, and description thereof will be omitted.

[0048] As shown in Fig. 6, the voice providing device 30 according to the third embodiment has a voice synthesis unit 11, a user information acquisition unit 31, and a voice personalization unit 32. These are functions or means realized by instructions from the CPU 1004 shown in Fig. 1. In addition, a user's recorded voice DB 91, a teacher's recorded voice DB 92, and a user information DB 93 are constructed in the recording medium 1001, the auxiliary storage device 1002, or the memory device 1003.

[0049] The user's recorded voice DB 91, the teacher's recorded voice DB 92, and the user information DB 93 may be constructed in a database server or the like outside the voice providing device 30. In this case, the voice providing device 30 accesses the database server or the like to transmit and receive data to and from the recorded voice DB 91, the teacher's recorded voice DB 92, or the user information DB 93. In addition, the teacher's recorded voice DB 92 in the third embodiment stores data on recorded voices by male teachers and recorded voices by female teachers.

[0050] The user information acquisition unit 31 acquires user information (the user's recorded voice, the user's self-evaluation value (score), the user's gender, and the user's skill evaluation value) to be used for personalization, and temporarily stores the information in the user information DB 93. Specifically, similar to the second embodiment, the user information acquisition unit 31 acquires the user's recorded voice from the user's recorded voice DB 91, acquires information on the self-evaluation value and the user's gender from a questionnaire, and also acquires information on the score (user's skill evaluation value) obtained by objectively evaluating the user's practice voice by a teacher or the like.

[0051] The voice personalization unit 32 determines (judges) self-esteem based on a difference value obtained by subtracting the user's skill evaluation value from the user's self-evaluation value among the user information (the user's recorded voice, the user's self-evaluation value, the user's gender, and the user's skill evaluation value) acquired from the user information DB 93. If the difference value is equal to or greater than a threshold value T2, the voice personalization unit 32 determines that the self-esteem is high, and if the difference value is less than T2, the self-esteem is determined to be low. Note that the threshold value T2 is an example of a second threshold value.

[0052] If the difference value is less than the threshold value T2, the voice personalization unit 32 outputs voice data of a teacher of the same gender from the teacher's recorded voice DB as a model voice based on the user's gender information.

[0053] On the other hand, if the difference value is equal to or greater than the threshold value T2, the voice synthesis unit 11 uses the user's recorded voice and the recorded voice of a teacher of the same sex as the user, obtained from the teacher's recorded voice DB 92, and outputs a voice having both the user's voice quality and the teacher's speaking style as a model voice. In this case, the voice synthesis unit 11 may also use the recorded voice of a teacher of the opposite sex to the user.

[0054] [Processing of the third embodiment] S31: The user information acquisition unit 31 acquires the user's recorded voice from the user's recorded voice DB 91 as user information, acquires the user's self-evaluation value and the user's gender from a questionnaire, and acquires the user's skill evaluation from an instructor, etc., and temporarily stores the same in the user information DB 93.

[0055] S32: The voice personalization unit 32 determines whether the difference value obtained by subtracting the user's skill evaluation value from the user's self-evaluation value among the user information (user's recorded voice, user's self-evaluation value, user's gender, user's skill evaluation value) obtained from the user information DB 93 is greater than or equal to threshold value T2.

[0056] S33: If the difference value is equal to or greater than the threshold value T2 (S32; YES), the voice synthesis unit 11 calculates the user's voice quality v based on the user's recorded voice acquired from the user's recorded voice DB 91 and the voice of a teacher of the same or opposite sex as the user acquired from the teacher's recorded voice DB 92. u and the teacher's speaking style t A voice that combines both voices is a model voice v e Let's say.

[0057] S34: On the other hand, if the self-evaluation value is less than (not greater than) the threshold value T2 (S32; NO), the voice personalization unit 32 selects the voice of a teacher of the same gender as the user as the model voice v e Let's say.

[0058] S35: The voice synthesis unit 11 or the voice personalization unit 32 generates a sample voice v e Output.

[0059] Effect of the Third Embodiment As described above, the voice providing device 30 according to the third embodiment switches between outputting a voice that combines the user's voice quality and the teacher's speaking style as a model voice, or outputting a voice of a teacher of the same gender as the user as a model voice, depending on the user's self-esteem determined based on the difference between the user's "self-evaluation value," which indicates the user's speaking skill level, and the user's "skill evaluation value," which is an objective evaluation of the user's practice voice by a third party such as a teacher. This provides an effect of obtaining an even greater practice effect than that of the second embodiment by taking into account whether the user's original skill level or self-evaluation level is high and addressing individual differences in reactions to the user's own voice.

[0060] Fourth Embodiment Next, a fourth embodiment will be described with reference to FIGS.

[0061] In the second and third embodiments, the voice providing devices 20 and 30 output model voices tailored to the user by determining the user's self-esteem. However, by taking into consideration the influence of personality traits, habits, and the like other than self-esteem, it may be possible to provide a model voice that is more suited to the user. Therefore, the voice providing device 40 according to the fourth embodiment calculates the number of times the user records and listens to their own voice during speaking (utterance) practice, and selects a model voice according to the number of times.

[0062] [Functional Configuration of Fourth Embodiment] The functional configuration of the fourth embodiment will be described with reference to Fig. 8. Fig. 8 is a diagram showing the functional configuration of a voice providing device according to the fourth embodiment. Note that in Fig. 8, the same functions as those in Fig. 2 according to the first embodiment, Fig. 4 according to the second embodiment, and Fig. 6 according to the third embodiment are denoted by the same reference numerals, and description thereof will be omitted.

[0063] As shown in Fig. 8, the voice providing device 40 according to the fourth embodiment has a voice synthesis unit 11, a user information acquisition unit 41, a voice personalization unit 42, and a listening frequency calculation unit 43. These are functions or means realized by instructions from the CPU 1004 shown in Fig. 1. In addition, a user's recorded voice DB 91, a teacher's recorded voice DB 92, and a user information DB 93 are constructed in the recording medium 1001, the auxiliary storage device 1002, or the memory device 1003.

[0064] The user's recorded voice DB 91, the teacher's recorded voice DB 92, and the user information DB 93 may be constructed in a database server or the like outside the voice providing device 40. In this case, the voice providing device 40 accesses the database server or the like to transmit and receive data to and from the recorded voice DB 91, the teacher's recorded voice DB 92, or the user information DB 93. In addition, the teacher's recorded voice DB 92 in the fourth embodiment stores data on recorded voices by male teachers and recorded voices by female teachers.

[0065] The listening frequency calculation unit 43 acquires the user's practice history, counts the number of times recording or playback was performed within a certain period (including a certain time) while practicing speaking (utterance) from this practice history, and calculates the "user's listening frequency" which indicates how often the user listened to their own recorded voice within the certain period.

[0066] The user information acquisition unit 41 acquires user information (the user's recorded voice, the user's gender, and the user's listening frequency) to be used for personalization and temporarily stores the information in the user information DB 93. Specifically, similar to the second and third embodiments, the user information acquisition unit 41 acquires the user's recorded voice from the user's recorded voice DB 91 and acquires information on the user's gender from a questionnaire, and also acquires information on the user's listening frequency from the listening frequency calculation unit 43. Note that the voice providing device 40 may not have the listening frequency calculation unit 43, and may acquire the user's listening frequency calculated outside the voice providing device 40.

[0067] The voice personalization unit 42 determines (judges) the listening habits (including personality traits that encourage listening to become a habit) based on the user's listening frequency from among the user information (the user's recorded voice, the user's gender, and the user's listening frequency) acquired from the user information DB 93. If the user's listening frequency is equal to or greater than a threshold F, the voice personalization unit 42 determines that the user has a listening habit of at least a certain level, and if the user's listening frequency is less than the threshold F, the voice personalization unit 42 determines that the user does not have a listening habit of at least a certain level. The threshold F is an example of a third threshold.

[0068] If the user's listening frequency is less than the threshold F, the voice personalization unit 42 outputs voice data of a teacher of the same gender from the teacher's recorded voice DB as a model voice based on the user's gender information.

[0069] On the other hand, if the user's listening frequency is equal to or greater than the threshold value F, the voice synthesis unit 11 uses the user's recorded voice and the recorded voice of a teacher of the same gender as the user, obtained from the teacher's recorded voice DB 92, and outputs a voice having both the user's voice quality and the teacher's speaking style as a model voice. In this case, the voice synthesis unit 11 may also use the recorded voice of a teacher of the opposite gender to the user.

[0070] [Processing of the fourth embodiment] S40: The listening frequency calculation unit 43 acquires the user's practice history, and from this practice history, counts the number of times recording or playback was performed within a certain period of time while practicing speaking (utterance), and calculates the "user's listening frequency" which indicates how often the user listened to their own recorded voice within the certain period of time.

[0071] S41: The user information acquisition unit 41 acquires the user's recorded voice from the user's recorded voice DB 91 as user information, acquires the user's gender from the questionnaire, and acquires the user's listening frequency from the listening frequency calculation unit 43, and temporarily stores it in the user information DB 93.

[0072] S42: The voice personalizing unit 42 determines whether the user's listening frequency is equal to or greater than a threshold value F from the user information (user's recorded voice, user's gender, user's listening frequency) acquired from the user information DB 93.

[0073] S43: If the user's listening frequency is equal to or greater than the threshold value F (S42; YES), the voice synthesis unit 11 determines the user's voice quality v based on the user's recorded voice acquired from the user's recorded voice DB 91 and the voice of a teacher of the same or opposite sex as the user acquired from the teacher's recorded voice DB 92. u and the teacher's speaking style t A voice that combines both voices is a model voice v e Let's say.

[0074] S44: On the other hand, if the user's listening frequency is less than (not greater than) the threshold value F (S43; NO), the voice personalization unit 42 selects the voice of a teacher of the same gender as the user as the model voice v e Let's say.

[0075] S45: The voice synthesis unit 11 or the voice personalization unit 42 generates a model voice v e Output.

[0076] Effect of the Fourth Embodiment As described above, the voice providing device 30 according to the fourth embodiment switches between outputting a voice that combines the user's voice quality and the teacher's speaking style as a model voice, or outputting a voice of a teacher of the same gender as the user as a model voice, depending on the user's habits of listening to their own voice (including the personality trait of trying to make listening a habit). This provides an advantage of obtaining a high practice effect in a different aspect from the second and third embodiments by addressing individual differences in reactions to the user's own voice, taking into account the influence of personality traits and habits other than self-esteem.

[0077] Supplementary Information The voice providing devices 10, 20, 30, and 40 can be realized by a computer and a program, but this program can also be recorded on a (non-transitory) recording medium or provided via a communication network such as the Internet.

[0078] Furthermore, the CPU 1004 as a processor may be a single processor or may be a multiple processor.

[0079] 10 Voice providing device 11 Voice synthesis unit 20 Voice providing device 21 User information acquisition unit 22 Voice personalization unit 30 Voice providing device 31 User information acquisition unit 32 Voice personalization unit 40 Voice providing device 41 User information acquisition unit 42 Voice personalization unit 43 Listening frequency calculation unit 91 User's recorded voice DB (an example of a user's recorded voice management unit) 92 Teacher's recorded voice DB (an example of a teacher's recorded voice management unit) 93 User information DB (an example of a user information management unit)

Claims

1. A voice providing device that provides a model voice, comprising: a voice personalization unit that determines whether a self-evaluation value indicating how skilled a user is in their own speaking style is equal to or greater than a first threshold value, and outputs, as the model voice, the voice of a teacher of the same gender as the user when the self-evaluation value is less than the first threshold value; and a voice synthesis unit that outputs a voice having a teacher's speaking style with the user's voice quality when the self-evaluation value is equal to or greater than the first threshold value.

2. A voice providing device that provides a model voice, comprising: a voice personalization unit that determines whether a difference value between a self-evaluation value indicating how skilled a user is in their own speaking style and a user skill evaluation value objectively evaluating the user's practice voice is equal to or greater than a second threshold value, and outputs, as the model voice, the voice of a teacher of the same gender as the user when the difference value is less than the second threshold value; and a voice synthesis unit that outputs a voice having a teacher's speaking style with the user's voice quality when the difference value is equal to or greater than the second threshold value.

3. A voice providing device that provides a model voice, comprising: a voice personalization unit that determines whether a user listening frequency indicating the frequency with which a user has listened to their own recorded voice within a certain period is equal to or greater than a third threshold value, and outputs, as the model voice, the voice of a teacher of the same gender as the user when the user listening frequency is less than the third threshold value; and a voice synthesis unit that outputs a voice having a teacher's speaking style with the user's voice quality when the user listening frequency is equal to or greater than the third threshold value.

4. The voice providing device according to claim 3, further comprising a listening frequency calculation unit that obtains the user's practice history and calculates the user's listening frequency by counting the number of recordings or reproductions made during the certain period while the user is practicing their speaking style from the practice history.

5. A voice providing method executed by a voice providing apparatus that provides a model voice, the method comprising: determining whether a self-evaluation value indicating to what extent a user has skills regarding the user's own speaking style is greater than or equal to a first threshold value; and when the self-evaluation value is less than the first threshold value, performing a voice personalization process of outputting the voice of a teacher of the same sex as the user as the model voice; and when the self-evaluation value is greater than or equal to the first threshold value, performing a voice synthesis process of outputting a voice having the speaking style of a teacher with the user's voice quality. A voice providing method that executes these processes.

6. A voice providing method executed by a voice providing apparatus that provides a model voice, the method comprising: determining whether a difference value between a self-evaluation value indicating to what extent a user has skills regarding the user's own speaking style and a user skill evaluation value objectively evaluating the user's practice voice is greater than or equal to a second threshold value; and when the difference value is less than the second threshold value, performing a voice personalization process of outputting the voice of a teacher of the same sex as the user as the model voice; and when the difference value is greater than or equal to the second threshold value, performing a voice synthesis process of outputting a voice having the speaking style of a teacher with the user's voice quality. A voice providing method that executes these processes.

7. A voice providing method executed by a voice providing apparatus that provides a model voice, the method comprising: determining whether a listening frequency of a user indicating the frequency with which the user has listened to the user's own recorded voice within a certain period is greater than or equal to a third threshold value; and when the user's listening frequency is less than the third threshold value, performing a voice personalization process of outputting the voice of a teacher of the same sex as the user as the model voice; and when the user's listening frequency is greater than or equal to the third threshold value, performing a voice synthesis process of outputting a voice having the speaking style of a teacher with the user's voice quality. A voice providing method that executes these processes.

8. A program for causing a computer to execute the method according to any one of claims 5 to 7.

Citation Information

Patent Citations

  • Auxiliary reading method and device, storage medium and electronic equipment

    CN111443890A

  • Method and device for leading interactive language learning

    JP2001159865A

  • Computer program for utterance leaning system and server device collaborating with the program

    JP2002244547A

  • Evaluation method and device for utterance and computer program for evaluating utterance

    JP2015068897A