Voice selection method, voice selection device and program

The voice selection device enhances navigation voice audibility by selecting or synthesizing a voice similar to the user's voice, adjusting based on noise and discomfort, addressing the challenge of clear navigation in noisy conditions.

WO2026004132A1PCT designated stage Publication Date: 2026-01-02NT T INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/023625
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing technologies lack effective methods for supporting clear voice navigation in environments with multiple competing sounds, particularly focusing on enhancing the audibility of navigation voices amidst various stimuli.

Method used

A voice selection device and method that utilizes a similarity calculation unit to select a navigation voice closest to the user's voice characteristics, optionally combined with voice synthesis to generate a personalized navigation voice, and adjusts output based on ambient noise, discomfort, and navigation performance.

Benefits of technology

Enables clear and comfortable navigation voice output tailored to the user's voice, reducing discomfort and improving audibility in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024023625_02012026_PF_FP_ABST
    Figure JP2024023625_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A computer, on the basis of a vocalization containing speech of a person, executes a voice synthesis procedure which generates a voice having text for guiding the behavior of the person as speech content. Thus, the output of a voice which the person can easily hear is made possible.
Need to check novelty before this filing date? Find Prior Art

Description

Audio selection method, audio selection device, and program

[0001] The present invention relates to a voice selection method, a voice selection device, and a program.

[0002] Voice navigation (guidance) is used in situations where it is difficult to see the screen, such as while walking or driving. In such situations, it is necessary to hear the navigation content while multiple stimuli are entering the ears. To support listening comprehension, for example, measures have been taken to reduce the volume of non-navigation content, such as music, when the navigation voice is played while driving.

[0003] On the other hand, there are voices that are easier to hear in situations where there are multiple speakers, as in Non-Patent Document 1. In this way, it is expected that the ease of hearing will increase depending on the type of voice (especially the voices of familiar people such as family members, as shown in Non-Patent Document 1).

[0004] Johnsrude, Ingrid S., et al., "Swinging at a cocktail party: Voice familiarity aids speech perception in the presence of a competing voice," Psychological science 24.10 (2013): 1995-2004

[0005] However, in the past, there has been insufficient research into methods for supporting listening in environments where multiple sounds are present, focusing on voices that are easy to hear.

[0006] The present invention has been made in view of the above points, and has as its object to make it possible to output a voice that is easy for a certain person to hear.

[0007] In order to solve the above problem, a computer executes a speech synthesis procedure for generating speech, the speech content of which is text for guiding the actions of a certain person, based on a voice including the speech of the certain person.

[0008] It is possible to output a voice that is easy for some people to hear.

[0009] FIG. 1 is a diagram illustrating an example of a hardware configuration of the voice selection device 10 in the first embodiment. FIG. 2 is a diagram illustrating an example of a functional configuration of the voice selection device 10 in the first embodiment. FIG. 3 is a flowchart illustrating an example of a processing procedure for voice selection processing in the first embodiment. FIG. 4 is a diagram illustrating an example of a functional configuration of the voice selection device 10 in the second embodiment. FIG. 5 is a diagram illustrating an example of a functional configuration of the voice selection device 10 in the third embodiment. FIG. 6 is a flowchart illustrating an example of a processing procedure for voice selection processing in the third embodiment. FIG. 7 is a diagram illustrating an example of a functional configuration of the voice selection device 10 in the fourth embodiment. FIG. 8 is a flowchart illustrating an example of a processing procedure for voice selection processing in the fourth embodiment. FIG. 9 is a diagram illustrating an example of a functional configuration of the voice selection device 10 in the fifth embodiment. FIG. 10 is a flowchart illustrating an example of a processing procedure for voice selection processing in the fifth embodiment.

[0010] The inventors conducted an experiment to assess the ease of listening by recording the voices of approximately 100 participants and having them listen to multiple voices, including their own recorded voices. The results showed that participants' own recorded voices were easier to hear than other people's voices.

[0011] Considering the results of the above experiments, it is thought that using one's own voice as the navigation voice used while walking or driving will enable navigation that is easier to hear even in environments where multiple voices are present.

[0012] Also, as in Reference 1, some people find their own voices unpleasant, so it is possible to switch between listening to their own voices depending on the situation, rather than always having them listen to it.

[0013] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0014] Fig. 1 is a diagram showing an example of the hardware configuration of an audio selection device 10 according to the first embodiment. The audio selection device 10 in Fig. 1 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, a communication device 105, a display device 106, an input device 107, a speaker 108, a microphone 109, and the like, all of which are interconnected via a bus B.

[0015] The program that realizes the processing in the audio selecting device 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the program does not necessarily have to be installed from the recording medium 101, but may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program as well as necessary files, data, etc.

[0016] The memory device 103 reads and stores the program from the auxiliary storage device 102 when an instruction to start the program is received. The processor 104 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to the audio selecting device 10 according to the program stored in the memory device 103. The communication device 105 is a module (antenna, circuit, etc.) for communication via a network. The display device 106 displays a GUI (Graphical User Interface) or the like according to a program. The input device 107 is composed of a touch panel or buttons, etc., and is used to input various operation instructions. The speaker 108 outputs audio. The microphone 109 inputs (collects) surrounding audio.

[0017] FIG. 2 is a diagram showing an example of the functional configuration of the voice selection device 10 according to the first embodiment. In FIG. 2, the voice selection device 10 includes a similarity calculation unit 11, a selection unit 12, and an output unit 13. These units are realized by processing executed by a processor 104 in accordance with one or more programs installed in the voice selection device 10. The voice selection device 10 also uses storage units such as a user voice storage unit 121 and a navigation voice DB 122. The navigation voice DB 122 can be realized using, for example, the auxiliary storage device 102 or a storage device connectable to the voice selection device 10 via a network.

[0018] The user voice storage unit 121 pre-stores voice data (hereinafter referred to as "user recorded voice") that is a recording of a speech made by a certain person (a user of the voice selection device 10). The content and length of the speech are not limited to a specific one. For example, the user recorded voice may be generated by requesting the user to read a prepared sentence of about several seconds and recording the content.

[0019] The navigation voice DB 122 is a database that stores voice data used for navigation (hereinafter referred to as "navigation voice"). Multiple types of voice data with different speaker characteristics for the same navigation (route guidance) are prepared in advance and stored in the navigation voice DB 122. It is desirable that the number of types of voice data is on the order of tens to hundreds. In this embodiment, it is assumed that N types of navigation voices are prepared. Note that speaker characteristics refer to the characteristics of the speaker, such as voice quality, gender, age, speaking style, etc.

[0020] In this embodiment, it is assumed that the voice selection device 10 is a personal terminal such as a smartphone or an in-vehicle device, and that there is a one-to-one relationship between the user and the voice selection device 10. However, if the voice selection device 10 is a server in a cloud system or the like, and provides functions to a terminal owned by the user via a network, the user may be identified by the terminal or based on the user's login from the terminal or the like.

[0021] The following describes the processing procedure executed by the voice selection device 10. Fig. 3 is a flowchart for explaining an example of the processing procedure for voice selection processing in the first embodiment.

[0022] In step S110, the similarity calculation unit 11 calculates a speaker vector V u A speaker vector is a vector that represents the characteristics of a speaker. A well-known neural network can be used to calculate the speaker vector from the voice vector. For example, an x-vector (Reference 2) may be used as the speaker vector.

[0023] Next, the similarity calculation unit 11 calculates the N types of navigation voice Un For each [i] (i=1, . . . , N), the navigation voice U n The speaker vector V of [i] n [i] (i=1, . . . , N) is calculated, and the speaker vector V n The similarity (S i (i=1, ..., N) is calculated (S130). The cosine similarity between speaker vectors can be used as the similarity. Since the similarity is the similarity between speaker vectors that indicates speaker characteristics, it can be said to be the similarity of speaker characteristics.

[0024] Next, the selection unit 12 selects the speaker vector V u The speaker vector V n The navigation voice related to [i] is the output target (output voice O n ) is selected as the target (S140).

[0025] Next, the output unit 13 outputs the output audio O n is output from the speaker 108 (S150).

[0026] Note that steps S110 to S140 may be performed only once before the start of navigation. n Step S150 may be executed every time a navigation voice is output.

[0027] As described above, according to the first embodiment, a navigation voice that is closest in speaker characteristics to (similar to) the user is output from among the prepared navigation voices. As a result, it is possible to output a voice that is easy for the user (a certain person) to hear. Therefore, it is expected that navigation that is easy to hear will be realized even in an environment where multiple voices are present, for example.

[0028] In the above example, the navigation voice with the highest similarity is selected as the output target, but any of the top M similarities (i.e., any of the navigation voices with relatively high similarities) may be selected as the output target.

[0029] Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be described. Points not specifically mentioned in the second embodiment may be the same as those in the first embodiment.

[0030] In the first embodiment, a voice similar to the user's voice was selected. However, the ease of listening may vary depending on the degree of similarity. Therefore, in the second embodiment, a voice synthesis technology is used to select a voice that is more similar to the user's own voice.

[0031] FIG. 4 is a diagram showing an example of the functional configuration of the voice selection device 10 according to the second embodiment. In FIG. 4, the same components as those in FIG. 2 are assigned the same reference numerals, and their description will be omitted. In FIG. 4, the voice selection device 10 has a voice synthesis unit 14 instead of the similarity calculation unit 11 and the selection unit 12. The voice synthesis unit 14 is realized by processing executed by the processor 104 of one or more programs installed in the voice selection device 10. The voice selection device 10 also uses a navigation text DB 123 instead of the navigation voice DB 122. The navigation text DB 123 can be realized using, for example, the auxiliary storage device 102 or a storage device connectable to the voice selection device 10 via a network.

[0032] The navigation text DB 123 stores the content spoken in the navigation voice as text (that is, text for guiding the route; hereinafter referred to as "navigation text").

[0033] The speech synthesis unit 14 generates (synthesizes) a navigation voice in which each navigation text is spoken, based on the user's recorded speech stored in the user speech storage unit 121. The speech synthesis unit 14 generates a navigation voice that sounds like the user by inputting the user's recorded speech. For example, a speech synthesis technology such as that described in Reference 3 can be used as the speech synthesis unit 14. Specifically, the speech synthesis unit 14 calculates a speaker vector from the user's recorded speech and inputs the speaker vector and the navigation text into a trained speech synthesis model, thereby generating a navigation voice that sounds like the user.

[0034] The output unit 13 outputs the navigation voice generated by the voice synthesis unit 14 .

[0035] The generation (synthesis) of a voice using each navigation text as the spoken content may be performed in advance, or may be performed at the timing when the spoken content of the navigation text is output.

[0036] The second embodiment may be combined with the first embodiment. For example, the second embodiment may be implemented when the highest similarity in the first embodiment is less than a threshold value.

[0037] As described above, according to the second embodiment, a navigation voice that closely matches the speaking characteristics of the user is output, making it possible to output a voice that is easy for the user (a certain person) to hear.

[0038] Next, a third embodiment will be described. In the third embodiment, differences from the above-described embodiments will be described. Points not specifically mentioned in the third embodiment may be the same as those in the above-described embodiments.

[0039] In the second embodiment, we considered using a synthesized voice of the user. However, constantly listening to the user's synthesized voice may be unpleasant. As mentioned above, some people find their own voice unpleasant (Reference 1). Therefore, in the third embodiment, we consider switching the voice to be output depending on the ambient noise level.

[0040] 5 is a diagram showing an example of the functional configuration of the audio selecting device 10 according to the third embodiment. In FIG. 5, the same or corresponding parts as those in FIG. 2 or FIG. 4 are designated by the same reference numerals, and the description thereof will be omitted as appropriate.

[0041] 5, the voice selection device 10 includes a noise level measurement unit 15 in addition to a voice synthesis unit 14, a selection unit 12, and an output unit 13. The noise level measurement unit 15 is realized by processing that is executed by the processor 104 of one or more programs installed in the voice selection device 10.

[0042] The noise level measuring unit 15 measures the noise level (for example, sound pressure level) of the surrounding sound input from the microphone 109, for example, periodically (for example, at one-minute intervals).

[0043] If the noise level is above a threshold, the selection unit 12 selects the user's synthesized voice generated by the voice synthesis unit 14 as the output target, and if the noise level is below the threshold, the selection unit 12 selects one of the navigation voices in the navigation voice DB 122 as the output target.

[0044] 6 is a flowchart for explaining an example of a processing procedure for selecting a voice in the third embodiment. The processing procedure in FIG. 6 is executed each time a voice is output during navigation (route guidance).

[0045] In step S210, the selection unit 12 determines whether the noise level last measured by the noise level measurement unit 15 (i.e., the latest noise level) is equal to or greater than the threshold value T.

[0046] If the noise level is equal to or greater than the threshold T (Yes in S210), the voice synthesis unit 14 calculates a speaker vector from the user's recorded voice, and inputs the speaker vector and the navigation text into a trained voice synthesis model to generate a user-like synthesized voice O. u The navigation text used at this time is a text corresponding to the navigation (route guidance) to be output. Next, the selection unit 12 generates the synthesized voice O u The output audio O n (S230).

[0047] On the other hand, if the noise level is less than the threshold value T (No in S210), the selection unit 12 selects one of the navigation voices in the navigation voice DB 122 as the output voice O. n For example, a navigation voice having a relatively low degree of speaker similarity to the user (for example, a navigation voice by a speaker of a different gender from the user) may be selected.

[0048] Following steps S230 and S240, the output unit 13 outputs the output audio O nis output from the speaker 108 (S250).

[0049] As described above, according to the third embodiment, when the ambient noise level is high (i.e., when the voice is difficult to hear), a voice that is close to the voice of the user is output. As a result, it is possible to output a voice that is easy for the user (a certain person) to hear while reducing the user's discomfort.

[0050] Next, a fourth embodiment will be described. In the fourth embodiment, differences from the third embodiment will be described. Points not specifically mentioned in the fourth embodiment may be the same as those in the third embodiment.

[0051] In the third embodiment, we considered controlling the output voice by taking into account the surrounding noise. However, it is possible that some users may feel uncomfortable listening to navigation using their own voice even in a noisy environment. Therefore, in the fourth embodiment, we take into account the results of measuring the degree of discomfort.

[0052] Fig. 7 is a diagram showing an example of the functional configuration of the voice selection device 10 according to the fourth embodiment. In Fig. 7, the same components as those in Fig. 5 are designated by the same reference numerals, and their description will be omitted. In Fig. 7, the voice selection device 10 further includes an discomfort level measurement unit 16. The discomfort level measurement unit 16 is realized by processing executed by the processor 104 of one or more programs installed in the voice selection device 10.

[0053] The discomfort level measurement unit 16 measures the user's discomfort level. For example, the discomfort level measurement unit 16 measures the discomfort level by inputting an image captured using a camera and performing emotion recognition based on facial expressions using publicly known technology. A trained machine learning model may be used for emotion recognition. The discomfort level is a value indicating the degree of discomfort, and here, the larger the value, the greater the degree of discomfort.

[0054] The selection unit 12 selects the user's synthesized voice as the output target when the noise level is equal to or higher than a threshold and the discomfort level is lower than the threshold, and otherwise selects the navigation voice in the navigation voice DB 122 as the output target.

[0055] 8 is a flowchart for explaining an example of the processing procedure for selecting audio in the fourth embodiment. In FIG. 8, the same steps as those in FIG. 6 are assigned the same step numbers, and their explanations will be omitted. In FIG. 8, if the ambient noise level is equal to or greater than the threshold T (Yes in S210), step S211 is further executed, and step S250 is followed by step S251.

[0056] In step S211, the selection unit 12 determines whether the discomfort level measured by the discomfort level measurement unit 16 after the output unit 13 has output the final sound is less than the threshold value F. The initial value of the discomfort level may be set to a minimum value (for example, 0). If the discomfort level is less than the threshold value F, the process proceeds to step S220, and if the discomfort level is equal to or greater than the threshold value F, the process proceeds to step S240.

[0057] In step S251, the discomfort level measurement unit 16 measures the discomfort level. n Synthetic voice O u In this case, the output audio n This may be done within a certain period after the output of the above.

[0058] As described above, according to the fourth embodiment, when the ambient noise level is high (i.e., when the voice is difficult to hear), if the user does not find their own voice unpleasant, a voice that is close to the speaker's characteristics is output. As a result, it is possible to output a voice that is easy for the user (a certain person) to hear while reducing the user's unpleasantness.

[0059] Next, a fifth embodiment will be described. In the fifth embodiment, differences from the fourth embodiment will be described. Points not specifically mentioned in the fifth embodiment may be the same as those in the fourth embodiment.

[0060] In the fourth embodiment, the navigation voice is output taking into consideration the surrounding noise and the user's emotions. In the fifth embodiment, in order to further improve the quality of navigation, the voice to be output is selected depending on the success rate of navigation.

[0061] Fig. 9 is a diagram showing an example of the functional configuration of the voice selection device 10 according to the fifth embodiment. In Fig. 9, the same components as those in Fig. 7 are designated by the same reference numerals, and their description will be omitted. In Fig. 9, the voice selection device 10 further includes a performance calculation unit 17. The performance calculation unit 17 is realized by a process in which one or more programs installed in the voice selection device 10 are executed by the processor 104.

[0062] The performance calculation unit 17 calculates performance related to past navigation. Performance is an index indicating the degree (degree of success) of the user's ability to move (travel) according to the navigation (i.e., according to the audio selected by the selection unit 12). Performance may be, for example, the percentage (L / K) of the number of times (L) the user was able to reach the destination according to the navigation out of a certain number (K) of past trips, i.e., the success rate. Hereinafter, performance based on such a history of past trips will be referred to as "past trip performance." According to the navigation means following the route intended by the navigation. Alternatively, performance may be the percentage of the current trip where the user was able to move according to the navigation at branch points such as intersections. In this case, if the number of branch points is K and the number of times the user was able to move according to the navigation is L, the performance (success rate) is L / K. Hereinafter, performance on the current trip will be referred to as "current trip performance." Whether it is the performance of a past trip or the performance of a current trip, the performance calculation unit 17 may record whether the user followed the navigation each time the navigation audio is output, and calculate performance based on the record. A trip refers to a single movement or movement section from a departure point to a destination.

[0063] The selection unit 12 selects the user's synthesized voice as the output target when the noise level is above a threshold, the discomfort level is below a threshold, and the performance is above a threshold, and otherwise selects the navigation voice in the navigation voice DB 122 as the output target.

[0064] 10 is a flowchart for explaining an example of the processing procedure for selecting audio in the fifth embodiment. In FIG. 10, the same steps as those in FIG. 8 are assigned the same step numbers, and their explanations will be omitted. In FIG. 10, if the discomfort level is less than the threshold F (Yes in S211), step S212 is further executed.

[0065] In step S212, the selection unit 12 determines whether the performance calculated by the performance calculation unit 17 is equal to or greater than the threshold value P. In the case of the performance of a past trip, the performance calculation unit 17 may calculate the performance at the start or end of each trip. In the case of the performance of the current trip, the performance calculation unit 17 may calculate the performance immediately before step S212 or a certain time after step S250.

[0066] If the performance is equal to or greater than the threshold P, the process proceeds to step S220, and if the performance is less than the threshold P, the process proceeds to step S240.

[0067] Note that performance basically refers to the performance of a trip in which navigation was performed based on the user's synthetic voice. Therefore, the performance calculation unit 17 may calculate performance only for trips in which the discomfort level did not exceed the threshold after the user's synthetic voice was output. However, the execution of step S212 likely indicates that the user does not find their synthetic voice unpleasant, and that this feeling is likely to have been the same in the past. Therefore, the performance value is likely to be similar whether or not the trips to be calculated as performance are narrowed down.

[0068] As described above, according to the fifth embodiment, when the ambient noise level is high (i.e., when the voice is difficult to hear) and the user does not find their own voice unpleasant, a navigation voice that closely resembles the speaker's characteristics is output if the performance is equal to or greater than the threshold value P. As a result, it is possible to output a voice that is easy for the user (a certain person) to hear while reducing the user's unpleasantness, thereby improving the success rate of navigation.

[0069] In the above embodiments, the selection of the voice to be output during navigation (route guidance) while walking or driving has been described, but the above embodiments may also be applied to guidance regarding human actions for operating various machines or other human actions.

[0070] [References] [Reference 1] Yanagida, Hikaru, Yusuke Ijima, and Naohiro Tawara, "Influence of Personal Traits on Impressions of One's Own Voice", INTERSPEECH 2023, 20-24 August 2023 [Reference 2] David Snyder et al. , "X-vectors: Robust DNN Embeddings for Speaker Recognition", 2018 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), April 2018 [Reference 3] Fujita, Kenichi, et al., "Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model.", arXiv preprint arXiv:2304.11976 (2023) The above describes in detail the embodiments of the present invention, but the present invention is not limited to such specific embodiments, and various modifications and variations are possible within the scope of the gist of the present invention described in the claims.

[0071] REFERENCE SIGNS LIST 10 Voice selection device 11 Similarity calculation unit 12 Selection unit 13 Output unit 14 Voice synthesis unit 15 Noise level measurement unit 16 Discomfort measurement unit 17 Performance calculation unit 100 Drive device 101 Recording medium 102 Auxiliary storage device 103 Memory device 104 Processor 105 Communication device 106 Display device 107 Input device 108 Speaker 109 Microphone 121 User voice storage unit 122 Navigation voice DB 123 Navigation text DB B Bus

Claims

1. A voice selection method characterized by the computer executing the following steps:

1. A voice synthesis procedure for generating, based on a voice including an utterance by a certain person, a voice having text as the spoken content for guiding the actions of the certain person.

2. The voice selection method according to claim 1, characterized in that a computer executes the following steps: a noise level measurement step for measuring a noise level; and a selection step for selecting, based on the noise level, either a voice generated by the voice synthesis step or a pre-prepared voice as an output target for guiding the person's actions.

3. The voice selection method according to claim 2, characterized in that the computer executes an discomfort measurement procedure that measures the user's discomfort level in accordance with the output of a voice when the voice generated by the voice synthesis procedure is selected as the output target by the selection procedure, and the selection procedure further selects, based on the discomfort level, either the voice generated by the voice synthesis procedure or a voice prepared in advance as the output target for guiding the person's actions.

4. The voice selection method according to claim 3, characterized in that the selection procedure further selects either a voice generated by the voice synthesis procedure or a voice prepared in advance as an output target for guiding the behavior of the certain person, based on the degree to which the certain person's behavior has followed the guidance of a voice previously selected as an output target.

5. A voice selection method characterized by being executed by a computer, which comprises: a similarity calculation step of calculating the degree of speaker similarity between each of a plurality of types of first voices having mutually different speaker characteristics and a second voice containing the speech of a certain person; and a selection step of selecting the first voices having a relatively high degree of similarity as output targets for guiding the actions of the certain person.

6. A voice selection device comprising: a voice synthesis unit configured to generate, based on a voice including an utterance of a certain person, a voice having text as the spoken content for guiding the actions of the certain person.

7. A program for causing a computer to execute a speech synthesis procedure for generating speech containing text for guiding the actions of a certain person, based on a speech including the speech of the certain person.

Citation Information

Patent Citations

  • Navigation apparatus and program therefor

    JP2006292918A

  • Voice system and voice output method of vehicle

    JP2020042163A

  • Information presentation method and information presentation device for vehicle

    JP2024067341A