System and program

The system addresses boredom in in-vehicle devices by allowing users to interact with a character through voice recognition, enhancing engagement and motivation through dialogue-based functionality.

JP2025100636APending Publication Date: 2025-07-03YUPITERU CORP
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2025063428
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Conventional in-vehicle electronic devices lack engagement and interaction, leading to boredom and reduced motivation for users, particularly drivers, as they operate the devices manually or through voice recognition without meaningful interaction.

Method used

A system that recognizes the voice of a vehicle occupant and interacts through a predetermined character, allowing users to engage in dialogue to execute functions, enhancing convenience and enjoyment by providing a sense of communication and interaction.

Benefits of technology

The system increases user engagement and motivation to actively use in-vehicle electronic devices by providing a fun and interactive experience, reducing boredom during long drives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100636000001_ABST
    Figure 2025100636000001_ABST
Patent Text Reader

Abstract

To provide a system and a program enabling a user to obtain fun in operation and driving, and providing him / her with an incentive to utilize a function positively without becoming bored.SOLUTION: A function required for a user is executed in accordance with a dialog between recognition information based on a result obtained by recognizing a voice of the user who is a passenger of a vehicle and an output of a voice of a prescribed character 100 indicating that the user's voice has been recognized. A voice uttered by the user is used as information recognized as a word phrase for establishing the dialog with the character 100, as the recognition information based on the result obtained by recognizing the voice of the user.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system and a program that recognize the voice of a vehicle occupant and execute predetermined processing.

Background Art

[0002] Currently, in-vehicle electronic devices equipped with functions such as navigation devices and radar detectors are widely spread. Such in-vehicle electronic devices have a function of outputting various types of information required by a user who is a vehicle occupant while the vehicle is running.

[0003] For example, when the in-vehicle electronic device has a function as a navigation device, a guidance route to a searched destination and various searched facilities can be displayed on a display device according to a user's operation. Further, for example, when the in-vehicle electronic device has a function as a radar detector, targets such as orbits, control / interrogation areas, and traffic monitoring devices can be appropriately displayed on the display device.

[0004] The operation of such in-vehicle electronic devices is basically performed by a user manually operating a switch provided in the device, a touch panel configured on the screen of the display device, or a remote controller capable of remote operation by infrared rays. Further, in-vehicle electronic devices having a voice recognition function of recognizing the voice spoken by a user and performing corresponding processing are also generally spread.

[0005] By using this voice recognition function, even a user who has difficulty in manual operation because he / she is driving can execute various functions such as route search and facility search by uttering a recognition phrase corresponding to a pre-registered command or the like.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0007] However, the voice recognition function in conventional in-vehicle electronic devices is no different from manual operation in that the device performs predetermined processing in response to an instruction unilaterally issued by the user. For this reason, the user only feels that they are operating a dull device, whether by manual operation or voice input.

[0008] In addition, general in-vehicle electronic devices also have a function of providing route guidance, target notification, etc. by means of pre-set voices. However, this function is also no different from information provision by screen display in that the device unilaterally provides information. For this reason, the user only feels that they are obtaining information from a dull device.

[0009] As described above, in the conventional technology, only the feeling of operating the device can be obtained, so there is no interest in it at all, and it is very boring for people in the vehicle, especially for a driver who is driving alone for a long time. For this reason, as time passes after purchase, the user loses the motivation to actively use the in-vehicle electronic device, and various functions often remain unused.

[0010] The present invention has been proposed to solve the problems of the prior art as described above, and its object is to provide a system and a program that can provide the user with fun in operation and driving, prevent boredom, and stimulate the motivation to actively use the functions.

Means for Solving the Problems

[0011] (1) The system of the present invention is characterized in that it executes a function required by a passenger of the vehicle in response to a dialogue based on recognition information based on the result of recognizing the voice of a user who is a passenger of the vehicle and the output of the voice of a predetermined character indicating that the voice of the user has been recognized. In this way, for example, the system does not execute functions in response to unilateral instructions from the user who is the occupant of the vehicle, but the user can execute functions required by the occupant of the vehicle by interacting with the character, which increases the convenience and enjoyment of operation and driving, and also provides a sense of communication with the character, preventing boredom. Also, for example, the user is motivated to operate the system in order to interact with the character, so the functions of the system are effectively utilized. The recognition information based on the result of recognizing the user's voice may be, for example, information that recognizes the voice uttered by the user as a phrase that establishes a dialogue with the character. In this way, for example, by outputting the voice of the character according to the recognized phrase, it is possible to indicate that the user's voice has been recognized and to establish a dialogue with the user. As for the recognition of the phrase, for example, in order to execute a function required by the passenger of the vehicle, a specific recognition phrase and a voice of a character that establishes a dialogue in relation to the specific recognition phrase are registered in advance, and the voice of the character is output according to whether or not it matches the result of recognizing the user's voice. In this way, for example, by preparing various combinations of recognition phrases for dialogue and voices of the character, it is possible to correspond to various processes that are executed according to the combinations. The predetermined character may be, for example, a speaker with a specific personality that is recognized by the voice output from the system. In this way, for example, the user can feel as if he or she is talking to a real person, which is favorable because it gives the user a sense of familiarity. As for the dialogue, for example, at least one round of interaction such as the start of a speech, a response to this, and a reply to the response may be made. By doing so, for example, instead of simply having the system process according to the user's command, a feeling can be obtained that the character intervenes between the user and the system and provides a service. Also, as for the dialogue, for example, it is preferable that the user's speech and the character's voice are established as the speech and the reply to this in terms of their semantic content. By doing so, a feeling that mutual communication of intentions is possible can be obtained. Also, as for the dialogue, for example, the reply to the speech may be off-topic in terms of semantic content. By doing so, an interesting feeling or surprise can be felt by the user. Furthermore, as for the dialogue, for example, it is preferable to mix the case where the speech and the reply to this are established as semantic content and the case where the reply to the speech is off-topic in terms of semantic content. By doing so, an incompleteness close to a real conversation can be produced, and a human touch can be made to be felt by the user. As a function required by a vehicle occupant, for example, in changing settings, outputting information, etc., it is preferable to be a function that enhances the convenience of the vehicle occupant. In particular, for example, it is preferable to be a predetermined process for outputting information that the user wants to know. As the information that the user wants to know, for example, by using voice recognition, it is preferable to be information that can be output to the system more easily or quickly than by manual operation. By doing so, for example, the user will come to actively use voice recognition, so the system can give the user more opportunities to interact with the character. In response to the dialogue, as for executing the functions required by the vehicle passengers, for example, as a result of the dialogue, it may be appropriate to execute the functions required by the vehicle passengers. By doing so, it is good because the desire to interact with the character arises in order to execute the functions required by the vehicle passengers. Also, as for executing the functions required by the vehicle passengers in response to the dialogue, for example, it may be appropriate to execute the functions required by the vehicle passengers during the dialogue. By doing so, it is good because the desire to speak and advance the dialogue increases in order to execute the functions required by the vehicle passengers. Furthermore, as for executing the functions required by the vehicle passengers in response to the dialogue, for example, it may be the execution of the process of outputting the voice of the character constituting the dialogue. By doing so, it is good because the user can enjoy the dialogue itself because they want to hear the voice of the character and thus engage in the dialogue.

[0012] As a predetermined process for outputting information that the user wants to know in response to the dialogue, it is advisable to use a process for outputting information required by the driver during vehicle operation. By doing so, for example, even when it is difficult for the driver to obtain desired information during driving due to difficult manual operations, restricted manual operations, etc., by using voice recognition, the user can obtain the desired information while enjoying the dialogue with the character.

[0013] As a predetermined process for outputting information that the user wants to know in response to the dialogue, it is advisable to use a process for outputting information through a plurality of selections. By doing so, for example, when it is necessary to go through a plurality of selections to output information, compared to the case where manual operations are required for each selection, voice input is easier for the user to operate. The information output through a plurality of selections may be, for example, information with a hierarchical structure or information that cannot be output without performing multiple operations. Such information may reduce the burden on the user compared to performing multiple manual operations. In particular, as the information output through a plurality of selections, for example, as in facility search, search around the current location, etc., it may be information obtained by sequentially narrowing down from among a plurality of candidates. The process of outputting such information is different from operations for moving the vehicle itself, such as brakes and accelerators, or operations for notifying something to the outside for safety during driving, such as turn signals, and it is difficult to directly execute it manually, so it may be suitable for voice recognition operations.

[0014] As part of the dialogue, when there is a next selection, it may be the voice of a character that guides this. In this way, for example, the user can sequentially proceed with predetermined processing by being prompted by the voice of the character to make a selection. As the voice of the character that guides to the next selection, for example, when the user's voice of "around" in a search around is recognized, it may be a voice that asks about the destination, such as "Where to go?" In this way, for example, it is clear to the user what to say next.

[0015] As part of the dialogue, when there is no next selection, it may be the voice of a character that indicates this. In this way, for example, the user can know that there is no need to proceed with the processing further, so it is possible to make a judgment such as ending the instruction or causing another process to be performed. As the voice of the character that indicates that there is no next selection, for example, when the user's voice of "family restaurant" in a search around is recognized, it may be a voice that clearly indicates that there is no next selection, such as "Going to a family restaurant alone?" In this way, the user can know that there is no need to proceed further, and it is possible to ask the user to make a new judgment.

[0016] As a predetermined process for outputting the information that the user wants to know according to the dialogue, it may be a process for causing the display means to display information about facilities around the desired location. In this way, for example, while eliminating the cumbersome manual operations for obtaining information on surrounding facilities, the user can obtain information on surrounding facilities while enjoying the interaction with the character.

[0017] As a predetermined process for outputting information that the user wants to know in response to the dialogue, it is preferable to relatively emphasize and display an icon indicating the surrounding facility at the position of the surrounding facility on the map displayed on the display means compared to the map. In this way, for example, on the map, the information on the facilities necessary for the user will stand out, so it becomes easier for the user to grasp the positions of the facilities and serves as a guide when selecting a desired facility by voice. As the relative emphasized display compared to the map, for example, it is preferable to use a display that makes the position of the facility clear. For example, by changing at least one of the brightness, saturation, and hue of the icon and the map, the contrast between the two can be enhanced. In particular, for example, if the icon is displayed in color and the map is displayed in grayscale, the icon can be seen very clearly.

[0018] As a predetermined process for outputting information that the user wants to know in response to the dialogue, display an icon indicating the surrounding facility at the position of the surrounding facility on the map displayed on the display means, and display the icon in a way that can distinguish the rank according to a predetermined standard. In this way, for example, by the display that can distinguish the ranks in the icon, it can be used as a guide for the user to select a facility. As the predetermined criterion, for example, it may be in the order of proximity to a specific location. By doing so, it is better because it is intuitive to understand which facility is the closest compared to the case where the facility name is displayed. Further, as the predetermined criterion, for example, it may be in the order of decreasing importance to the user. By doing so, it is better because the facilities that the user most wants to know are preferentially displayed. The importance to the user may be, for example, having been there in the past, having a high frequency of going there in the past, or matching the user's preference, etc. By doing so, it is better because the user can quickly know the facilities they want to visit. Also, for example, those with low importance to the user, such as having been there in the past but not wanting to go there again, may be excluded. By doing so, for example, it is possible to display only those that the user may visit. As a display that can distinguish the ranking, for example, it is better to use information that inherently indicates the ranking, such as numbers and alphabets. By doing so, for example, the meaning of the ranking can be easily grasped. Further, as a display that can distinguish the ranking, for example, it is better to use information that visually indicates the ranking, such as shapes like ◎, ○, △, ×, etc., or by making gradual changes according to brightness, saturation, and hue. By doing so, for example, the ranking can be intuitively grasped visually.

[0019] The voice of the character prompting the selection of the searched facility may be output with content matching the displayed icon. By doing so, for example, it is possible to prompt the selection of the facility with the voice of the content matching the displayed icon, so the user can also grasp the content confirmed visually by voice. As the content corresponding to the displayed icon, for example, if it is a number, it may be a voice asking which number is appropriate; if it is an alphabet, it may be a voice asking which letter of the alphabet is appropriate; if it is a shape, it may be a voice asking what the shape is; if it is a color, it may be a voice asking what the color is. By doing so, for example, the user can easily know what to say to select a facility. Also, for example, as the content corresponding to the displayed icon, it may be a voice indicating the range of candidates to be selected. Thereby, for example, if the voice is output as from what number to what number, the user can also grasp the range to be selected along with the content to be spoken.

[0020] When recognizing the voice of a user who selects any facility with the content corresponding to the displayed icon, it is preferable to display on the display means the searched route from a specific point to the selected facility. By doing so, for example, the user can easily select with the content corresponding to the displayed icon rather than selecting by facility name. As the content corresponding to the displayed icon, for example, if it is a number, it may be the number of the icon of the desired facility; if it is an alphabet, it may be the alphabet of the icon of the desired facility; if it is a shape, it may be the shape of the icon of the desired facility; if it is a color, it may be the color of the icon of the desired facility. By doing so, for example, the facility can be selected with words that are easier to remember than the facility name.

[0021] After the searched route is displayed on the display means, it is preferable to determine whether to start route guidance or re-search the route according to the dialogue. By doing so, for example, when the user does not want the searched route, the user can give an instruction to re-search while having a dialogue with the character, so that the trouble and boredom during re-search can be alleviated.

[0022] As a predetermined process for outputting information that the user wants to know in response to a dialogue, it is advisable to display, on the display means, a table of items for selecting which of the items corresponding to the recognition phrases registered in advance for voice recognition will function as the recognition phrase. In this way, for example, if only the items that the user often uses are selected from the displayed table of items, the recognition phrases are limited, so misrecognition can be reduced. The selected recognition phrases should be registered, for example, by rewriting them from the default state in a predetermined memory area, putting them in an empty area, rearranging those already in it, etc. In this way, the limited memory area can be effectively utilized.

[0023] Display, on the display means, a table of items for selecting which of the items corresponding to the recognition phrases registered in advance for recognizing the user's voice will function as the recognition phrase for the dialogue, and the items selected from the said table should be set for each page to be displayed on the display means. In this way, for example, even if the number of items that can be displayed on one page of the display screen of the display means is limited, the items that the user often uses can be brought to the top page, so the trouble of switching pages to display the items that the user wants to use can be saved.

[0024] It is advisable to set the items of the recognition phrases that may be recognized confusingly to be displayed on another page. By doing so, for example, phrases containing similar sounds, phrases that are likely to be confused when actually recognized, etc. will not be arranged on the same page and become candidates for recognition at the same time, so it is possible to reduce misrecognition. As phrases containing similar sounds, for example, phrases containing the same sounds in hiragana, the same Chinese characters, etc., such as "family restaurant" and "fast food", "hospital" and "beauty salon", "Western food" and "Japanese food", may be used. By doing so, since recognition phrases with confusing sounds are not displayed on the same page from the characters as well, the user will input either one of them, and misrecognition can be reduced. As phrases that are likely to be confused when actually recognized, for example, based on the record of past recognition results, different phrases recognized for the same utterance may be used. By doing so, for example, it is possible to eliminate the situation where the combination where misrecognition actually occurs becomes the same page.

[0025] The item corresponding to the recognition phrase has a major item corresponding to a recognition phrase including a plurality of concepts and a middle item corresponding to a recognition phrase corresponding to the plurality of concepts. It is advisable to display the middle item together with the major item on the same page. By doing so, for example, even when the user wants to search for a middle item, there is no need to change the hierarchy or page, so it is possible to directly select the desired item. As the major item, for example, a higher-level concept item including a large number of candidates, such as "hospital", may be used. By doing so, it will be convenient for users who want to search from as many candidates as possible. As the middle item, for example, a lower-level concept item including fewer candidates than the major item, such as "internal medicine" and "pediatrics" included in the concept of "hospital", may be used. By doing so, it will be convenient for users who want to immediately obtain information on the desired facility from a small number of candidates.

[0026] As a recognition phrase, it is advisable to register a phrase that is related to the phrase displayed as an item but is not displayed as an item. By doing so, for example, even if the user does not utter the recognition phrase itself displayed in the item, it suffices to utter a phrase related to the recognition phrase. Thus, even if the user's memory of the item name is vague or the user is unaccustomed to the operation, the search can be performed, which is acceptable. A phrase that is related to the phrase displayed as an item but not displayed as an item may be, for example, a middle item that corresponds to a large item being displayed on a specific page but is not displayed on that page. By doing so, for example, even if the middle item is not displayed, if the user utters the middle item, the middle item can be searched, and thus the information on the necessary facility can be obtained quickly, which is acceptable.

[0027] As a predetermined process for outputting information that the user wants to know in response to the dialogue, the character may be displayed on the display means. By doing so, for example, the user can become familiar with the appearance of the dialogue partner by seeing it on the display screen and may be motivated to actively use the system to view the appearance.

[0028] As a predetermined process for outputting information that the user wants to know in response to the dialogue, the character may be displayed at a position that does not obstruct the display of the searched route on the map. By doing so, for example, it is possible to maintain the visibility of the searched route while allowing the user to enjoy the display of the character, which is acceptable. As a position that does not obstruct the route display, for example, the character and the display of the searched route may be closer to opposite sides of the display screen. By doing so, the user can be made to feel as if the character is taking care not to interfere with the route search service, and thus a better impression can be given.

[0029] The character may be moved while being displayed at a position that does not obstruct the display of the searched route on the map. By doing so, for example, by continuously displaying a moving character, it is possible to make the user feel bored and appeal for its existence, and for users who want to continuously watch the character, they can enjoy the actions of the moving character.

[0030] The display area of the item corresponding to the recognized speech of the user has a part where the background image can be seen. The character displayed by the display means is larger than each item and does not have a part where the background image can be seen, and it is preferable that the display area of the item is closer to the position on the opposite side of the character on the display screen. By doing so, for example, when the character is closer to the side opposite to the searched route, the display area of the item may overlap with the searched route, but each item is smaller than the character, and since the display area of the item has a part where the background image can be seen, it is better that the background image is easier to see than the searched route overlapping with the character. As the part where the background image can be seen, for example, it may be a transparent part or a part with a gap. By doing so, for example, it is possible to see the background image while ensuring the visibility of the item display.

[0031] It is preferable to insert the voice of the pre-registered user's name into the voice output by the character in the dialogue. By doing so, for example, the user can feel a greater sense of familiarity with the character when being called by their own name in the conversation. It is preferable to insert the voice of the pre-registered user's name into the voice output by the character at the start of the dialogue. By doing so, for example, at the start of the dialogue, the user will be called by their name, so it can attract the user's attention.

[0032] When the voice including the name of the character is recognized, it is preferable to insert the voice of the pre-registered user's name into the voice output by the character. By doing so, for example, they will call each other by name, and the user can obtain a more intimate feeling.

[0033] As the conversation progresses, it is advisable to reduce the frequency of inserting the user's name into the voice output by the character. By doing so, for example, at first the name is called, but later the frequency of being called decreases, so it is possible to create a conversation between two people who are getting used to each other.

[0034] Voice data corresponding to the 50 - sound chart should be registered in advance, and the voice of the name selected by the user from this 50 - sound chart should be inserted during the conversation. By doing so, for example, compared with the case of registering a plurality of names in advance, the memory capacity can be saved.

[0035] As the voice of the character, it is advisable to insert the voice of the content indicating the situation of the character. By doing so, the user can hear the situation of the character. For example, if the character is in a negative situation, it can arouse the user's concern for the character, and if the character is in a positive situation, it can arouse the user's feeling of empathy for the character. As a statement indicating the situation of the character, for example, a statement such as "I'm busy" can be used to show whether the character is in a positive or negative state. By doing so, for example, the system does not unconditionally follow the user's instructions, but can arouse the user's concern and empathy, so that a feeling similar to a conversation with an actual person can be obtained.

[0036] For the voice of the character in the conversation, it is advisable to randomly select and output from among a plurality of different pre - set phrases, excluding the phrases that have already been output. By doing so, for example, it is possible to prevent only the same phrases from being output, so that the user will not get bored, which is good.

[0037] It is preferable that a list is set in which a plurality of different response phrases are registered for a common recognition phrase. By doing so, for example, even if the user makes a statement with common content, the character will respond in different expressions. Therefore, even when realizing the same process many times, the user can enjoy what kind of response will come from the character.

[0038] It is preferable that a plurality of the lists are set, and there is a list in which phrases different from other phrases are registered even within the same list. By doing so, for example, as a response to the user's speech with similar content, different phrases are output, which can give the user an unexpected impression, change the impression of the character, and prevent boredom.

[0039] Even if the character is the same person, there are a plurality of display modes, and as the output voice, different phrases are registered for each of the plurality of display modes. By doing so, for example, even for the same character, the speech content is different according to the plurality of display modes, so that while having a conversation with the same person, the user can enjoy the changes. As the plurality of display modes, for example, standard ones, ones that transform by changing clothes, hairstyles, etc., and ones with different body proportions may be used. By doing so, for example, it is possible to give the user various impressions such as adult-like, child-like, cute, and beautiful, and prevent boredom.

[0040] It is preferable that phrases to be output are registered when voice recognition is incorrect. By doing so, for example, even when the recognition is incorrect, it is possible to make the user recognize that the conversation with the character is continuing rather than returning some phrase.

[0041] It is preferable that a phrase serving as an escape route in case of an error in voice recognition be registered. In this way, for example, although there may be a possibility of an error in voice recognition, even if an error occurs, the frustration and anger of the user can be alleviated by the escape route of the character. As an escape route, for example, an expression indicating the will to try to recognize accurately, such as "I'll do my best in voice recognition. It may not be 100% recognizable, though", and an expression indicating that complete achievement is difficult may be used. In this way, for example, it is possible to create a psychology in the user to tolerate misrecognition rather than simply indicating that recognition is impossible.

[0042] When there is an input of the user's voice pointing out that the content the user expects as a response and the content of the character's voice output are different, it is preferable that a response phrase from the character to this be registered. In this way, for example, if the user points out the mistake of the character, there is a response from the character to this, so it is possible to enjoy the natural interaction between people in case of misunderstanding.

[0043] It is preferable that a phrase for advertising be registered. In this way, for example, it is possible to guide the user to a partnering facility.

[0044] As the voice of the character in the dialogue, it is preferable that a phrase about the reason for doing something at the facility the user is going to be registered. In this way, for example, by talking about the reason for going to a certain facility, it becomes easier for the user going to that facility to listen. For example, it is possible to attract the user's attention whether the reason for going to the destination is favorable or unfavorable for the user. As facilities where the reason for going to the destination is not favorable to the user, for example, facilities related to the user's illness such as "hospital", "dentist", and "pharmacy" may be considered. As facilities where the reason for going to the destination is favorable to the user, for example, facilities that are slightly more challenging than the facilities the user usually visits, such as "department store", "sushi", and "eel", may be considered. As phrases that make the character evoke various emotions in the user, for example, phrases that make the user feel attached to the character by showing concern for the user's state, such as "Is something wrong? Are you okay?" and "Don't overdo it", may be considered. Also, for example, phrases that make the user feel regret, such as "Are you brushing your teeth properly?", may be considered.

[0045] As the voice of the character in the dialogue, phrases that give positive or negative advice regarding what to do at the facility the user is going to may be registered. By doing so, for example, the user can determine what to do and what not to do at the destination based on the character's advice. As positive advice, for example, phrases that recommend something, such as "It's good to get your teeth cleaned once in a while", may be considered. As negative advice, for example, phrases that prohibit something, such as "Don't make that painful face. A real man endures with a cool face", may be considered.

[0046] As the voice of the character in the dialogue, phrases that are the one-sided impressions of the character regarding what to do at the facility the user is going to may be registered. By doing so, for example, the user can obtain the character's impressions and thus can use them as a reference for actions at the destination. As one-sided impressions of the character, for example, positive impressions such as "Black pork char siu that melts in your mouth. Wonderful~." would be good. By doing so, it would be good for the user to want to take similar actions at the facility. Also, for example, negative impressions such as "Going to a restaurant alone? Aren't you lonely?·· Leave me alone." would be good. By doing so, it would be good to make the user want to change the destination. Furthermore, for example, impressions that mix positive and negative impressions such as "Shopping malls seem fun just to stroll around. But I'm likely to get tired." would be good. By doing so, it would be good to be able to make the user aware of the situation at the destination in advance.

[0047] As the voice of the character in the dialogue, it would be good if phrases that make the user think the character actually exists are registered. By doing so, for example, the user can get the feeling that the character actually exists. As phrases that make the user think the character actually exists, for example, phrases that make the user think they are traveling together such as "I'm so happy that you're taking me on a business trip too~.", "Where are we going to stay?" would be good. By doing so, for example, the user can get the feeling of always being with the character.

[0048] (2) It would be good if recognition phrases registered in advance to recognize the user's voice are displayed as items. By doing so, for example, the user can see what to say in order to execute a predetermined process by looking at the displayed items.

[0049] (3) The recognition phrases registered in advance to recognize the user's voice may include recognition phrases that can be commonly recognized regardless of the display screen and recognition phrases whose recognition is restricted to what is displayed on the display screen. In this way, for example, by using recognition phrases that can be commonly recognized, it is possible to maintain the convenience of the functions required regardless of the display screen, and by using recognition phrases with restricted recognition, the possibility of misrecognition can be reduced.

[0050] As recognition phrases that can be commonly recognized, for example, they may be recognition phrases that are commonly used in multiple types of display screens. In this way, the possibility that the function is restricted according to the display screen is reduced, and the convenience can be maintained. For example, it may be a phrase for transitioning to another display screen, such as "return" or "cancel". Such phrases are, for example, phrases that need to be used on any screen. As recognition phrases with restricted recognition, for example, they may be phrases displayed on the display screen or phrases to be searched for. Such phrases are, for example, less likely to require recognition of other phrases, and it is advisable to reduce the number of recognition phrases to prevent misrecognition.

[0051] (4) It is advisable that the recognition phrases registered in advance for recognizing the user's voice include those displayed on the display screen and those not displayed on the display screen. In this way, for example, even when the user utters a phrase not displayed on the display screen, if it is a recognition phrase set in advance, some reaction can be obtained from the character, which can be a surprise for the user and give the user the joy of discovery.

[0052] It is advisable that voice data output by the character is registered when voice other than the phrases set in advance as recognition phrases is input. In this way, for example, even if it is other than the recognition phrases, there is always a reaction from the character, so that the user is not reminded that it is an interaction with the system by the conversation being interrupted or being notified as an error, and the feeling of continuing the conversation with the character can be maintained.

[0053] (5) The state of a function that changes quantitatively may be directly set to a desired quantitative level by voice recognition. In this way, for example, compared with the case of changing the quantitative level step by step between a small value and a large value, the desired value can be set immediately, so it does not take much effort for the user and the time restricted by the operation is shortened. As the function that changes quantitatively, for example, it may be volume, brightness, scale, etc. In this way, it is good because it is possible to save the trouble of selecting buttons such as volume, brightness, and scale to display +- or a scale and then increasing or decreasing step by step by selecting +- or sliding on the scale.

[0054] (6) As the recognition phrases registered in advance for recognizing the user's voice, an indicator corresponding to the function that changes quantitatively and a number indicating the quantitative level may be set. In this way, for example, when the user says an indicator corresponding to a desired function and a number, the quantitative level of the function can be directly increased or decreased to the level of that number. As the indicator corresponding to the function that changes quantitatively, for example, it may be "volume", "brightness", "scale", etc., and the numbers may be "one", "two", "two zeros", etc. In this way, for example, it is good because the input can be made using function names and numbers that are intuitively easy for the user to understand. In particular, for example, like "volume three", "brightness two", "scale two zeros", the name and the quantitative level may be grouped into one recognition phrase. In this way, the speech can be completed in one go.

[0055] On the display screen of the display means, a button for selecting a function and having an indicator corresponding to the function displayed may be displayed. In this way, for example, the user only needs to say the indicator displayed on the button, which is easier for the user to understand.

[0056] On the display screen of the display means, it is preferable to display a button for selecting a function, on which the value of the current quantitative level of the function is displayed. In this way, for example, the user can check the current quantitative level by looking at the value displayed on the button and determine whether to increase or decrease it to the desired value. The value of the current quantitative level may be, for example, the scale of a map. In this way, for example, when a small-scale value is displayed on the button and the user requests a larger value, the scale value displayed on the button directly changes to that value, and the user can immediately view a wide-range map. Also, for example, when a large-scale value is displayed on the button and the user requests a smaller value, the scale value displayed on the button directly changes to that value, and the user can immediately view a detailed map.

[0057] The button on which the instruction corresponding to the function is displayed is preferably characterized in that there is a button for executing the function without transitioning to another display screen. In this way, for example, the function can be executed interactively without transitioning to another display screen to execute the function, so the current screen display may not be disturbed.

[0058] (7) Whether voice recognition is being accepted or not may be displayed on the display means. In this way, for example, the user can visually recognize whether voice operation is possible.

[0059] As to whether voice recognition is being accepted or not, for example, the color of the display screen may be changed for display. In this way, for example, the state of voice recognition can be immediately perceived through the user's vision. In particular, for example, when voice recognition is being accepted, it may be set to the normal color, and when it is not being accepted, it may be set to grayscale. In this way, for example, the user can intuitively understand whether voice recognition is being accepted or not.

[0060] (8) Whether voice recognition is being accepted or not may be displayed on the display means together with the character. By doing so, for example, by looking at the character, it is possible to tell whether an operation by voice is possible or not, which is good because it makes it easier for the user to notice. As a mode of display on the character, for example, it is preferable that the state of the character clearly changes. By doing so, for example, the state of voice recognition can be immediately perceived. As a mode in which the state of the character clearly changes, for example, an item may be added to the character, the color or form of the character may be changed, a specific pose or gesture may be made, or a line may be displayed in a speech bubble. By doing so, for example, a user who knows the normal state of the character can immediately notice the change by comparing it with the normal state. As a mode of adding an item to the character, for example, equipment related to sound such as a headset, headphones, earphones, or microphone may be worn by the character. By doing so, for example, the user can be made to associate voice recognition, which is a function related to sound. As a mode of changing the color or form of the character, for example, the hairstyle or hair color may be changed, or the character may change into another costume. By doing so, for example, the difference from the normal state can be easily understood. In particular, for example, if the ears of the character are made larger, it may indicate a posture of trying to hear the voice well. Furthermore, as a pose or gesture, for example, a gesture of putting a hand behind the ear may also indicate a posture of trying to hear the voice well.

[0061] (9) While the voice of the character is being output, voice recognition may be stopped until the voice output is finished. By doing so, for example, when there is an input of the user's voice during the character's voice, it is possible to prevent other processes from being executed by voice recognition. As other processes, for example, processes with low urgency such as volume, brightness, scale, and peripheral search may be used. By doing so, it is not necessary for processes that do not necessarily need to be executed immediately to be executed by voice recognition.

[0062] It is advisable to accept voice recognition even while the character's voice is being output. By doing so, for example, when there is an input of the user's voice during the character's voice, other processes do not need to be interrupted by voice recognition. As other processes, for example, a process of transmitting position information regarding the place of supervision to an information collection server may be used as a process with high urgency. By doing so, it is not necessary for the process that the user wants to execute immediately to be interrupted.

[0063] When there is an input of the user's voice while the character's voice is being output, it is advisable to stop the output of the character's voice. By doing so, for example, an impression can be given as if the user's voice interrupted the character's speech. The output of the stopped character's voice does not need to be repeated from the beginning, for example, without outputting it again, so as not to give an unnatural impression. Also, if the output of the stopped character's voice is started from where it was stopped, for example, an impression can be given that the character was waiting during the user's speech.

[0064] When a predetermined input operation is performed while the character's voice is being output, it is advisable to stop the character's voice. In this way, for example, when one does not want a third party to hear the conversation with the character, the voice of the character can be stopped by a predetermined input operation. As the predetermined input operation, for example, it may be an operation of a physical key. In this way, since one only needs to operate a location that always exists fixedly, an emergency operation can be easily performed.

[0065] (10) In response to the input of the user's voice corresponding to the recognition phrase registered in advance for recognizing the user's voice, it is advisable to make a post that transmits the location information regarding the place of inspection to the information collection server. In this way, for example, a post notifying the place of inspection can be easily and quickly made by voice instead of by manual operation. As the place of inspection, for example, it may be a check location where the implementation position and time change, such as an inspection and interrogation area. In this way, information about a place of inspection for which it is difficult to register the position in advance can be collected by a post. As the location information, for example, it may be the current location or the center position coordinates. In this way, for example, the place of inspection can be specified with information that can be obtained surely and easily, and since voice recognition is used, the deviation from the actual inspection position can be small.

[0066] (11) As the predetermined recognition phrase, it is advisable that a phrase for activating voice recognition and a phrase for starting to accept a post are registered. In this way, for example, by the user uttering a phrase for activating voice recognition and a phrase for starting to accept a post, multi-stage operations such as voice recognition and accepting a post can be easily executed.

[0067] As the phrase for activating voice recognition and the phrase for starting to accept a post, for example, they may be registered as separately divided phrases. By doing so, for example, it is possible to prevent the acceptance of submissions from starting simultaneously with the start of voice recognition and the occurrence of misrecognition. Also, as the phrases for activating voice recognition and the phrases for starting the acceptance of submissions, for example, they may be set as a series of phrases. By doing so, for example, since the acceptance of submissions can be started all at once simultaneously with the start of voice recognition, it is possible to immediately enter a state where submissions are possible.

[0068] (12) As the predetermined recognition phrases, it is preferable that the phrases for activating voice recognition and the phrases for making submissions are registered. By doing so, for example, the user can easily perform multi-step operations such as voice recognition and submission by uttering the phrases for activating voice recognition and the phrases for making submissions.

[0069] As the phrases for activating voice recognition and the phrases for making submissions, for example, they may be registered as separate phrases. By doing so, for example, it is possible to prevent submissions from being made simultaneously with the start of voice recognition and from being erroneously submitted. Also, as the phrases for starting voice recognition and the phrases for making submissions, for example, they may be set as a series of phrases. By doing so, for example, since submissions can be made all at once simultaneously with the start of voice recognition, it is possible to immediately submit and minimize the deviation between the place of regulation and the place of submission as much as possible.

[0070] The predetermined recognition phrases for submissions are preferably shorter than other recognition phrases. By doing so, for example, for submissions with high urgency, it is possible to give instructions in a short time.

[0071] The predetermined recognition phrases for submissions are preferably more redundant than other recognition phrases. By doing so, for example, when using normal voice recognition, the possibility of uttering other recognition phrases is reduced, and it is possible to prevent erroneous submissions. By being redundant, for example, it may be considered that the number of characters is large or it is divided into multiple parts. By doing so, for example, it is possible to reduce the possibility that a user who does not intend to post will accidentally speak.

[0072] The predetermined recognition phrase for posting may be a phrase that the user is unlikely to say normally. By doing so, for example, it is possible to prevent frequent misrecognition from occurring due to phrases that the user frequently uses in the car normally, and to prevent unintended posts by the user.

[0073] The voice of the character in the dialogue for posting may be shorter than the voice in other dialogues. By doing so, for example, it is possible to shorten the dialogue time and enable quick posting.

[0074] (13) The dialogue may start with an input of voice from the user corresponding to a predetermined recognition phrase. By doing so, for example, since a predetermined process starts with an input of voice from the user, it is not necessary for the system to start automatically and execute the process.

[0075] (14) The dialogue may start with an input of voice from the user corresponding to a predetermined recognition phrase in response to an output of voice of a predetermined character. By doing so, for example, since the output of the voice of the character becomes an opportunity to determine whether the user starts a predetermined process, it is possible to increase the opportunity to use the system. The output of the voice of the character may, for example, inquire about the presence or absence of the start of a specific process. By doing so, for example, the user can be prompted to use the process. If the timing of the output of the voice of the character is, for example, regular, it is possible to increase the frequency of using the process. Also, if the timing is set to be when a predetermined condition is satisfied, it is possible to prompt the user to use it at a timing suitable for using the process.

[0076] When the execution of the function required by the vehicle occupant starts, if there is no voice input from the user and a predetermined time has elapsed, the process may end. By doing so, for example, it is possible to prevent malfunction caused by continuous voice recognition for a long time.

[0077] (16) From the time when the execution of the function required by the vehicle occupant starts, until a predetermined time elapses, a character may output a voice prompting the user for voice input. By doing so, for example, when the user forgets to start the process, it is possible to prompt voice input.

[0078] (17) The functions as the system of (1) to (16) may be configured as a program for realizing them by a computer.

Advantages of the Invention

[0079] According to the present invention, it is possible to provide a system and a program that can give the user a sense of fun in operation and driving, prevent boredom, and stimulate the motivation to actively use the functions.

Brief Description of the Drawings

[0080]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Embodiment for Carrying Out the Invention

[0081] [Basic Configuration] FIG. 1 shows an external view of an embodiment of a navigation device suitable as an electronic device constituting the system of the present invention, FIG. 2 shows its functional block diagram, and FIGS. 3 and 4 show an example of a display form. As shown in FIG. 1, the navigation device includes a portable device main body 2 that can be carried and a cradle 3 that is an attachment member for holding it. By attaching the device main body 2 to the cradle 3, it functions as an in-vehicle navigation device, and by detaching it from the cradle 3, it functions as a portable navigation device (PND). In this embodiment, the device main body 2 is detachable from the cradle 3, and it can be installed and used on the dashboard of the vehicle together with the cradle 3, or the device main body 2 can be removed from the cradle 3 and used as a portable (PND). However, it may not be easily detachable from the cradle 3, and it may be an in-vehicle fixed type or a dedicated portable type. Furthermore, for example, it may be realized by installing a predetermined application program (including pre-installation) on a portable terminal such as a mobile phone. The application program may be the system itself for performing navigation, or it may be a program for accessing a server on which the system for performing navigation is installed and using the system. The present invention can be applied to various types of such navigation systems.

[0082] The device main body 2 is detachably attached to the cradle 3. The device main body 2 includes a flat rectangular case main body 4. On the front surface of the case main body 4, a display unit 5 is arranged, and on the display unit 5, a touch panel 8 for detecting which part of the display unit 5 has been touched is provided, and warning lamps 9 are provided on both sides of the front surface. The cradle 3 includes a cradle main body 6 for holding the device main body 2, and a pedestal portion 7 for instructing the cradle main body 6 in an arbitrary posture at a predetermined position (such as a dashboard) in the vehicle interior. The pedestal portion 7 is adsorbed and fixed on the dashboard or the like by a suction cup provided on the bottom surface. The pedestal portion 7 and the cradle main body 6 are connected via a connection mechanism such as a ball joint so as to be rotatable within a predetermined angle range. Since it is a ball joint, the pedestal portion 7 and the cradle main body 6 can rotate relative to each other within an arbitrary angle range in the three-dimensional direction and maintain their positions at an arbitrary angular position due to the frictional resistance at the joint portion. Therefore, the device main body 2 attached to the cradle main body 6 can also be arranged in an arbitrary posture on the dashboard.

[0083] Furthermore, on one side surface of the case main body 4, an SD memory card slot portion 21 is provided, and an SD memory card 22 in which map data or the like is recorded can be inserted into the SD memory card slot portion 21. This SD memory card slot portion 21 includes a memory card reader / writer for reading and writing information to and from the inserted SD memory card 22. Also, on the side surface of the case main body 4 where the SD memory card slot portion 21 is provided, a DC jack 10 is provided. The DC jack 10 is for connecting a cigar plug cord (not shown), and can receive power supply by connecting to the cigar socket of the vehicle via the cigar plug cord.

[0084] On the other hand, on the side surface opposite to the SD memory card slot portion 21, a power switch and a USB terminal 23 are provided. By connecting to a personal computer via this USB terminal 23, it is possible to perform software application version upgrades and the like.

[0085] Inside the case body 4, the following devices and components are arranged. That is, a microwave receiver 11 is arranged inside the back side of the case body 4. The microwave receiver 11 receives microwaves in a predetermined frequency band, and when it receives microwaves in the set frequency band, it detects the signal level of the received microwaves. Specifically, it uses the RSSI voltage corresponding to the signal level and the electric field strength. The above-mentioned predetermined frequency band includes, for example, the frequency band in which the frequency of microwaves emitted from a vehicle speed measuring device is included.

[0086] Also, a GPS receiver 12 that receives GPS signals and determines the current position is arranged inside the upper surface side of the case body 4. An infrared communication device 14 is arranged inside the front side of the case body 4. The infrared communication device 14 performs data transmission and reception with a communication device incorporating an infrared communication device such as a mobile phone 15. Further, a speaker 20 is also incorporated in the case body 4.

[0087] Furthermore, the device of this embodiment includes a wireless receiver 13 and a remote control receiver 16. The wireless receiver 13 receives wireless signals of a predetermined frequency flying in. The remote control receiver 16 performs data communication with a remote control (portable device: slave unit) 17 and performs various settings for the device. The wireless receiver 13 receives wireless signals in a predetermined frequency band flying in. This predetermined frequency can be, for example, the frequency band of the wireless used when an emergency vehicle notifies its own vehicle position to a base station.

[0088] In addition to the navigation function, the navigation device of this embodiment also has a target notification function as a target notification device that notifies targets such as vehicle speed measuring devices and other traffic monitoring points existing around. These navigation functions and target notification functions are stored on the EEPROM of the control unit 18 as programs executed by a computer included in the control unit 18, and are realized by the computer included in the control unit 18 executing them.

[0089] That is, the control unit 18 executes predetermined processing based on information input from the various input devices (GPS receiver 12, microwave receiver 11, wireless receiver 13, touch panel 8, infrared communication device 14, remote control receiver 16, SD memory card slot unit 21, USB terminal 23, etc.) described above, and outputs predetermined information, warnings, and messages using output devices (display unit 5, warning lamp 9, speaker 20, infrared communication device 14, SD memory card slot unit 21, USB terminal 23, etc.). This predetermined processing is for executing the above-described various functions, and accesses the database 19 and the SD memory card 22 as necessary.

[0090] Here, the database 19 can be realized by a non-volatile memory (e.g., EEPROM) built into the microcomputer of the control unit 18 or externally attached to the microcomputer. The database 19 stores information necessary for implementing certain target objects, maps, and other functions at the time of shipment. Data about subsequently added target objects can be updated through predetermined processing. The processing for this update can be performed, for example, by mounting the SD memory card 22 storing additional data in the SD memory card slot unit 21 and transferring the data from the SD memory card 22 to the database 19. Also, this data update can be performed by using the infrared communication device 7 or by using a personal computer or other external device connected via the USB terminal 23.

[0091] The control unit 18 for realizing the target object notification function operates as follows. That is, when the microwave receiver 11 receives a desired microwave, the control unit 18 outputs a predetermined warning. Such warnings include outputting a buzzer using the speaker 20, outputting a voice to notify of the reception of a microwave (detection by a vehicle speed measuring device, etc.), and outputting a message in characters or images using the display unit 5.

[0092] In addition, when the wireless receiver 13 receives a desired wireless signal, the control unit 18 outputs a predetermined warning. Such warnings include the output of a buzzer using the speaker 20, the output of voice to notify the reception of microwaves (such as the approach of an emergency vehicle), and the output of a message in the form of characters or images using the display unit 5. It is preferable that the form of the warning associated with the reception of this wireless signal is different from the form of the warning associated with the reception of the above-mentioned microwaves.

[0093] Furthermore, when the current position detected by the GPS receiver 12 and the position of a target such as a traffic monitoring point stored in the database 19 are in a predetermined positional relationship, the control unit 18 outputs a predetermined warning. Therefore, the database 19 contains information about the target to be detected (position information of the target including longitude and latitude, type information of the target, etc.), traffic safety information for driving safely with more caution such as accident-prone areas and traffic control information, landmarks, and various kinds of information useful for driving. Each piece of information is registered by associating the specific type of information (type of target, type of traffic control, accident-prone area, name of landmark, etc.) with the position information.

[0094] The distance to the target when issuing a warning can be changed for each type of target. As warning modes, similar to the above, there are warnings by voice using the speaker 20 and warnings using the display unit 5. FIG. 3 shows an example of a warning by the display unit 5. In this embodiment, since the device is a navigation device and has map data as described later, as a basic screen, the control unit 18 has a function of reading road network information around the current position and displaying a map around the current position on the display unit 5. Then, a warning screen 70 is displayed on the display unit 5 over the currently displayed screen (here, a map around the current position). FIG. 3 shows an example of the display of the warning screen 70 when the distance between the current position and an LH system, which is a type of speed measurement device and one of the traffic monitoring points, reaches 500 m while the map 50 is being displayed. Furthermore, a process is performed to output a warning voice indicating the warning type and distance, such as "There is an LH system 500 m ahead", from the speaker 20.

[0095] On the one hand, the control unit 18 for realizing the navigation function operates as follows. First, the database 19 stores road network information for navigation. The information for navigation stored in this database 19 may store all information about the whole country at the time of shipment, or map data, etc. may be provided by storing them in the SD memory card 22 for each region. The user may prepare the SD memory card 22 storing the necessary map data, etc., and mount it on the SD memory card slot unit 21 for use. Note that the map data, etc. stored in the SD memory card 22 may be transferred to and stored in the database 19, or the control unit 18 may access the SD memory card 22 and read and use the data therefrom.

[0096] The control unit 18 has a function of reading the road network information around the current position from the database 19 and displaying the map around the current position on the display unit 5. The control unit 18 can search for a route from one position to another position using this road network information. In addition, the database 19 includes a telephone number database that stores telephone numbers in association with the location information and names of residences, companies, facilities, etc. of the telephone numbers, and an address database that stores addresses in association with the location information of the addresses. Further, in the database 19, the location information of traffic monitoring points such as speed measurement devices is stored together with their types.

[0097] In addition, the control unit 18 has a function of performing the processing of a general navigation device. That is, it displays a map around the current position on the display unit 5 at any time and displays a destination setting button. When the control unit 18 detects a touch at a position corresponding to the display position of the destination setting button on the touch panel 8, it performs destination setting processing. In the destination setting processing, a destination setting menu is displayed on the display unit 5 to prompt the user to select a method for setting the destination. The destination setting menu has a phone number search button and an address search button for prompting the user to select a method for setting the destination. When it is detected that the phone number search button has been touched, a phone number input screen is displayed, and position information corresponding to the input phone number is acquired from the database 19. When it is detected that the address search button has been touched, an address selection input screen is displayed, and position information corresponding to the input address is acquired from the database 19. Then, the acquired position information is set as the position information of the destination, and a recommended route from the current position to the destination is obtained based on the road network information stored in the database 19. As a method for calculating this recommended route, a known method such as Dijkstra's algorithm can be used, for example.

[0098] Then, the control unit 18 displays the calculated recommended route together with the surrounding map. For example, as shown in FIG. 4, a route 53 from the current position 51 to the destination 52 is displayed as a recommended route in a predetermined color (for example, red). This is the same as a general navigation system. In this embodiment, the control unit 18 compares the position information on the route 53 with the position information of targets such as traffic monitoring points stored in the database 19, and displays the positions of traffic monitoring points and the like located on the route 53. Note that, as information provided for driving safely on the route 53, it may be limited to traffic monitoring points among various targets stored in the database 19.

[0099] As a result, for example, as shown in FIG. 4, a balloon is displayed from a position on the road, and the type of traffic monitoring point at that position stored in the database 19 is displayed in the balloon. For example, FIG. 4(a) is a display example of a warning point search screen, showing an example in which traffic monitoring points 55a to 55f are displayed. In the balloon of traffic monitoring point 55a, the character "N" indicating the N system is displayed. In the balloons of traffic monitoring points 55b, 55c, and 55e, the characters "LH" indicating the LH system are displayed. In the balloon of traffic monitoring point 55d, the character "loop" indicating the loop coil is displayed. In the balloon of traffic monitoring point 55f, the character "H" indicating the H system is displayed. Thus, when the touch panel 8 detects a touch on the balloon position in the simplified display state where a character is displayed in the balloon, as shown in FIG. 4(b), the simplified display is switched to the detailed display. The detailed display is to display the warning screen shown in FIG. 3 in the balloon.

[0100] Also, as shown in FIG. 4(a), a traffic monitoring point type designation section 56 is displayed on the lower right side within the display section 5. The traffic monitoring point type designation section 56 is a part that performs button display listing the types of traffic monitoring points existing on the route 53. In the example of FIG. 4(a), a type button 57a with the character "N" indicating the N system of traffic monitoring point 55a, type buttons 57b with the characters "LH" indicating the LH system of traffic monitoring points 55b, 55c, and 55e, a type button 57c with the character "loop" indicating the loop coil of traffic monitoring point 55d, and a type button 57d with the character "H" indicating the H system of traffic monitoring point 55f are displayed. When the control unit 18 detects a touch on the display section of this type button from the touch panel 8, the touched type button is reversely displayed, and the display mode of the traffic monitoring point corresponding to the touched type button is switched to the detailed display if it is in the simplified display, or switched to the simplified display if it is in the detailed display.

[0101] The car navigation device according to this embodiment has a surrounding search function as a method for setting a destination (hereinafter including a waypoint). When this surrounding search function is selected and activated, the control unit 18 extracts facilities that meet the specified conditions and are close to the current position, and draws the extraction result on the display unit 5. Then, the user can select any one from the candidates drawn on the display unit 5 and specify it as the destination (including the waypoint) by confirming it.

[0102] That is, in order to realize such a function, the database 19 of this embodiment stores information about each facility in association with information regarding the classification of the facility together with its location information. The information regarding the classification of the facility is a plurality of types of items grouped by the names of the genres that each facility matches.

[0103] In this embodiment, when any one of the items is specified, the control unit 18 extracts facilities that meet the specified item and are close to the current position, and draws the extraction result on the display unit 5 as candidates for the destination. The user can specify it as the destination by selecting any one from the candidates drawn on the display unit 5.

[0104] When any one of the items is selected, the control unit 18 accesses the database 19 and extracts facilities that exist within a reference distance (for example, 10 km) centered on the current position and match the classification of the specified item. Then, icons (marks) indicating the facilities are superimposed and drawn at corresponding positions on the map drawn on the display unit 5 from the closest ones for a predetermined number (for example, 10) of facilities. As information about the facilities stored in the database 19, for each facility, the respective icons are registered in association, and at the time of drawing, the associated icons are read out and drawn at predetermined positions on the display unit 5.

[0105] In addition, when the current position exists on the map drawn on the display unit 5, a bicycle icon indicating the bicycle is drawn overlaid on the position corresponding to the current position. Further, when drawing such facility icons, numbers are added and displayed in ascending order from the bicycle. The numbers are updated as the bicycle moves.

[0106] When the user selects a desired facility, the control unit 18 that recognizes this obtains a recommended route with such a facility as the destination and draws the result overlaid on the map of the display unit 5.

[0107] Furthermore, the control unit 18 has a function of performing route guidance to the set destination. That is, in response to the user's selection to start guidance, the control unit 18 sequentially detects the position of the vehicle by GPS or autonomous navigation, draws map information representing roads and the like on the display unit 5, and performs road guidance to the destination using images and sounds.

[0108] The above instructions for input such as designation and selection by the user can be performed when the control unit 18 detects that a button with the name of an item or function is displayed on the display unit 5, the facility icon is touched, the selection is made by remote control operation, or a predetermined voice is input.

[0109] [Function of interacting with characters] 《Basics of the interaction function》 Hereinafter, the function of interacting with characters, which is a feature of the present embodiment, will be described. This interaction function is a function realized by the control unit 18 executing a program, and is a function realized by recognition information based on the result of recognizing the voice of the user who is a passenger in the vehicle and the output of the voice of a predetermined character indicating that the voice of the user has been recognized.

[0110] To execute this dialogue function, the database 19 stores recognition words (recognized phrases) and response phrases which are the voice data of corresponding characters. The control unit 18 recognizes the voice of the user input via the microphone 20 based on the recognition words. That is, the control unit 18 recognizes the voice uttered by the user as phrases based on the voice signal of the user, and determines whether the recognized phrases match the recognition words. Then, the control unit 18 causes the response phrase corresponding to the recognition word determined to match the recognized phrase to be output to the speaker 20. By preparing various combinations of recognition words for dialogue and response phrases which are the voices of characters, as will be described later, it is possible to support various processes executed according to the dialogue.

[0111] Thereby, the user can have a simple dialogue with the character using the electronic device which is the system. The character of this embodiment is set as a girl named "Ray", for example. Various functions such as the above-mentioned peripheral search can also be performed in the dialogue with the character. The database 19 stores "Rei-tan" as the name of the character and as a recognition word.

[0112] In addition, the database 19 stores image data (including still images and moving images) of the character. Based on this image data, the control unit 18 can display a still image or a moving image of the character on the display unit 5. Note that for some response phrases, the control unit 18 can make the user feel the situation where the character is speaking realistically by using lip sync to move the mouth and body of the character displayed on the display unit 5 so as to match the response phrase to be output.

[0113] Furthermore, for the voices and images of the characters stored in the database 19, multiple modes are set for the same person. For example, the database 19 divides a single person named "Ray" into three modes: "Standard Ray", "Transformed Ray", and "Chibi Ray", and stores the corresponding response phrases and image data for each mode. "Standard Ray" is the basic mode of the character. "Transformed Ray" is a mode in which the clothes and hairstyle are different from those of "Standard Ray". "Chibi Ray" is a mode in which the body proportion is changed to make the character look like a child. The control unit 18 outputs the voice and image of the character according to the selection of any mode by the user.

[0114] 《Recognition Words》 Figures 5 and 6 are examples of the list of recognition words stored in the database 19. Each column of this list of recognition words is assigned a list number. As will be described later, each column contains recognition words that can be recognized on each display screen M1, V1, V1a, S1, S1a, S2, S2a, S2_list, S2a_list, R1, R1a, R1b, A1, A2, A3 displayed on the display unit 5.

[0115] The activation word "Let's go, voice control" for activating voice recognition is registered in list number 1. When the control unit 18 recognizes the activation word, it activates the voice recognition function for executing various functions such as surrounding search. Therefore, in the initial state, the control unit 19 waits in a state where only the activation word can be recognized by voice recognition. This list number 1 corresponds to the display screen M1 described later.

[0116] In list number 2, after activating voice recognition, recognition words for starting any one of peripheral search, scale change, volume change, and brightness change are registered. For example, "しゅうへん" (peripheral) is registered as the recognition word for starting peripheral search. "すけーる" (scale) is registered as the recognition word for starting scale change. "おんりょう" (volume) is registered as the recognition word for starting volume change. "きど" (brightness) is registered as the recognition word for starting brightness change. This list number 2 corresponds to the display screen V1 described later. Furthermore, in list number 2, the recognition word "じたく" (home) for searching for the route to home in one utterance when returning home is registered. In this way, since the recognition words corresponding to the functions frequently used by the user are registered as the recognition words when voice recognition is activated, voice recognition can be started and executed immediately.

[0117] In list number 3, after starting route guidance, recognition words corresponding to the functions to be executed are registered. For example, in addition to the recognition words in list number 2, recognition words for canceling route guidance such as "あんないちゅうし" (cancel guidance), "あんないていし" (stop guidance), and "あんないやめる" (end guidance) are registered. By preparing multiple types of recognition words for executing the same function of canceling route guidance in this way, the possibility of not being recognized can be reduced even if the user's utterance content is somewhat different. This list number 3 corresponds to the display screen V1a described later. In particular, "しゅうへん" (peripheral), "おんりょう" (volume), and "きど" (brightness) will be the instruction words displayed on the buttons of the display screens V1 and V1a as described later. Therefore, if the user utters the displayed instruction word, instruction input becomes possible.

[0118] For list numbers 4 and 5, recognition words for specifying a genre are registered in the peripheral search. These genre recognition words are displayed in hiragana characters on the buttons of the items for specifying the genre. On the screen display where the buttons of the recognition words belonging to list number 4 are shown, the recognizable genres are limited to the recognition words belonging to list number 4. This list number 4 corresponds to the display screen S1 described later. Note that for family restaurants, multiple types of recognition words for specifying the same genre, such as "ふぁみれす" (famiresu) and "ふぁみりーれすとらん" (family restaurant), and for electronics stores, "かでん" (kaden) and "でんきや" (denkiya), are prepared. By doing so, even if the user's utterance content is somewhat different, the possibility of unrecognizability can be reduced.

[0119] In the screen display where the buttons of the recognition words belonging to list number 5 are shown, the recognizable genres are limited to the recognition words belonging to list number 5. This list number 5 corresponds to the display screen S1a described later. In this way, since the recognizable genres on each display screen are limited to the recognition words corresponding to the buttons of the items shown on each display screen, the possibility of misrecognition is reduced.

[0120] Furthermore, in list number 4, recognition words "にぺーじ" (nipeji) and "つぎ" (tsugi) for transitioning to the display screen S1a of list number 5 are registered. Also, in list number 5, recognition words "いちぺーじ" (icheji) and "まえ" (mae) for transitioning to the display screen S1 of list number 4 are registered. By preparing multiple types of recognition words for executing the same function of transitioning the genre display screen, even if the user's utterance content is somewhat different, the possibility of unrecognizability can be reduced.

[0121] For list numbers 6, 7, 8, and 9, recognition words for selecting facilities searched by the surrounding search are registered. On the display screens corresponding to list numbers 6 and 7, the facilities searched by the surrounding search are displayed on the map at their locations with icons numbered in ascending order of proximity to the current position. As recognition words, the names of the numbers for selecting facilities are registered. For example, recognition words for indicating numbers such as "ichi", "ichiban", "ni", and "niban" are registered. By preparing multiple types of recognition words for selecting the same target as the destination in this way, it is possible to reduce the possibility of failure to recognize even if the user's utterance content is somewhat different. Note that on the display screens corresponding to list numbers 6 and 7, all the searched facilities (in this example, 10 locations) can be selected by voice regardless of whether they are displayed on the screen or not.

[0122] In addition, for list numbers 6 and 7, recognition words for changing the scale of the displayed map are registered. That is, for list number 6, the recognition word "suke-ru" for accepting the start of scale change is registered. This list number 6 corresponds to the display screen S2 described later. For list number 7, in addition to the recognition word "scale" for accepting the start of scale change, recognition words for changing the scale are registered. This list number 7 corresponds to the display screen S2a described later.

[0123] Recognition words for changing the scale include recognition words for enlarging the scale of the map for detailed display, recognition words for reducing the scale of the map for wide-area display, and recognition words for indicating the scale number and unit. Recognition words for enlarging the scale of the map include "expansion", "detail", and recognition words for reducing the scale of the map include "contraction", "wide area", etc., and include recognition words for changing the scale of the map step by step. By "step by step" it means that multiple levels of scale are set in advance and the levels are raised or lowered one by one. In this way, by preparing multiple types of recognition words for executing the same function of changing the scale step by step, even if the user's speech content is somewhat different, the possibility of not being recognized can be reduced. In particular, when raising or lowering the scale, since words used in daily life are used as recognition words, the user does not have to be confused about the words for instructing volume change.

[0124] In addition, recognition words for changing the scale include recognition words that can specify the scale number and unit, such as "ten meters", "twenty-five meters", "fifty meters". As a result, when the user utters and specifies the recognition word for the desired scale, the control unit 18 that recognizes this directly changes it to the specified scale of the map. Furthermore, in list numbers 6 and 7, recognition words "list", "list display" for transitioning to the display screen of the list display of list number 8 described later are registered. In this way, by preparing multiple types of recognition words for executing the same function of transitioning to the list display screen, even if the user's speech content is somewhat different, the possibility of not being recognized can be reduced.

[0125] On the display screens corresponding to list numbers 8 and 9, facilities found in the surrounding area search are displayed in a list using buttons with numbers displayed on them. The names of the numbers for selecting the facilities displayed in this list are registered as recognition words. List number 8 corresponds to the display screen S2_list described later, and list number 9 corresponds to the display screen S2a_list described later. The names of the registered numbers are the same as those of list numbers 8 and 9 above. However, on list numbers 8 and 9, the facilities that can be selected are limited to those displayed on the display screen. In this example, on the display screen S2_list corresponding to list number 8, the first to fifth facilities can be selected, and on list number S2_list, the sixth to tenth facilities can be selected.

[0126] Furthermore, in list numbers 8 and 9, the recognition words "list" and "cancel list" for canceling the list display are registered. In this way, by preparing multiple types of recognition words for executing the same function of canceling the list display, the possibility that the user's speech will not be recognized even if it differs slightly is reduced. Furthermore, like list numbers 4 and 5, in list number 8, the recognition words "page 2" and "next" for transitioning to the display screen S1a of list number 9 are registered. In addition, in list number 9, the recognition words "page 1" and "previous" for transitioning to the display screen S1 of list number 8 are registered.

[0127] In list numbers 10, 11, and 12, the recognition words "Guide me", "Please guide me", and "Please guide me" that instruct the start of the searched route guidance are registered. In this way, by preparing multiple types of recognition words that instruct the same function of route guidance, the possibility that the user's speech will not be recognized even if it differs slightly can be reduced. In particular, by preparing the command tone of "Guide me", as well as the friendly request tone of "Please guide me" and "Please guide me", users can enjoy changing the tone of speech depending on the level of intimacy with the character.

[0128] List number 10 corresponds to a display screen that displays the searched route after route search. This display screen corresponds to the display screen R1 described later. List number 11 corresponds to a display screen that changes the scale of the map while displaying the searched route. This display screen corresponds to the display screen R1a described later. List number 12 corresponds to a display screen that changes the route conditions while displaying the searched route. This display screen corresponds to the display screen R1b described later.

[0129] Also, in list numbers 10, 11, and 12, recognition words "suke-ru" for starting scale change, "shu-hen" for re-searching surrounding facilities, "jouken" and "ru-to jouken" for starting route condition change are registered. In this way, by preparing multiple types of recognition words for instructing the same function of route condition change, even if the user's utterance content is somewhat different, the possibility of not being recognized can be reduced.

[0130] Furthermore, in list number 11, similar to list number 7, a recognition word for changing the scale is registered. Also, in list number 12, a recognition word for changing the route condition is registered. The route condition is a search condition for what to prioritize when searching for a route. As the route conditions in this embodiment, highway priority, general road priority, and recommended route are set. The recommended route is set with default conditions, for example, time-to-destination priority, distance-to-destination priority, etc. Correspondingly, in list number 12, "sui-shou", "ippan-dou", and "kou-soku-dou" are registered.

[0131] In list numbers 13, 14, and 15, recognition words for changing various functions of this embodiment are registered. In list number 13, a recognition word for changing the map scale in the same way as above is registered. List number 13 corresponds to the display screen A1 described later.

[0132] In list number 14, recognition words for changing the volume are registered. List number 14 corresponds to display screen A2, which will be described later. The recognition words for changing the volume include recognition words for increasing the volume, recognition words for decreasing the volume, and recognition words for indicating the volume level.

[0133] The recognition words for increasing the volume include "appu", "ookiku", and the recognition words for decreasing the volume include "daun", "chiisaku", etc., and include recognition words for changing the volume step by step. By "step by step" it means that a plurality of volume levels are set in advance, and the levels are increased or decreased one by one. By preparing a plurality of types of recognition words for executing the same function of changing the volume step by step in this way, even if the user's speech content is somewhat different, the possibility of not being recognized can be reduced. In particular, when increasing or decreasing the volume, since words used in daily life are used as recognition words, the user does not have to be confused about the words for instructing volume change.

[0134] In addition, the recognition words for changing the volume include recognition words that can specify numbers indicating the volume level, such as "zero", "ichi", "ni". Thus, when the user speaks and specifies the desired volume level, the control unit 18 that recognizes this directly changes the volume to the specified level. Note that "zero" is a recognition word that realizes the function of setting the volume to zero, that is, muting the sound output from the speaker 20. Thus, when the user does not want the character's voice to be heard by others, the user can immediately make the character's voice inaudible just by speaking the recognition word for setting the volume to zero.

[0135] In list number 15, recognition words for changing the brightness are registered. List number 15 corresponds to display screen A3, which will be described later. The recognition words for changing the brightness include recognition words for increasing the brightness, recognition words for decreasing the brightness, and recognition words for indicating the brightness level.

[0136] Recognition words for increasing brightness include "appu", "akaruku", and recognition words for decreasing brightness include "daun", "kuraku". Recognition words that change brightness step by step are included, such as these. By "step by step", it means that multiple levels of brightness are set in advance, and the levels are increased or decreased one by one. In this way, by preparing multiple types of recognition words each for executing the same function of changing brightness step by step, even if the user's speech content is somewhat different, the possibility of not being recognized can be reduced. In particular, when increasing or decreasing brightness, since words used in daily life are used as recognition words, the user can easily indicate the change in brightness without confusion.

[0137] In addition, recognition words for changing brightness include recognition words that can specify numbers indicating the levels of brightness, such as "zero", "ichi", "ni". Thus, when the user speaks the level of the desired brightness, the control unit 18 that recognizes this directly changes the brightness to the specified level. Note that "zero" is a recognition word that realizes the function of setting the brightness to zero and the function of darkening the screen of the display unit 5 so that the display content cannot be seen. Thereby, when the user does not want the character image to be seen by others, the user can immediately make the character image invisible just by speaking the recognition word for setting the brightness to zero.

[0138] In the above list numbers 13, 14, 15, recognition words for executing various functions are registered in common. That is, "sukeru" related to scale change, "shuhen" related to peripheral search, "onryo" related to volume change, and "kido" related to brightness change are registered.

[0139] On the display screen of list number 13, when the user speaks "sukeru", the control unit 18 that recognizes this completes the setting of the changed scale. On the other hand, on the display screens of list numbers 14, 15, when the user speaks "sukeru", the control unit 18 that recognizes this transitions to the display screen for scale change, that is, the display screen corresponding to list number 13.

[0140] On the display screen of list number 14, when the user speaks "onryou", the control unit 18 that recognized this completes the setting of the changed volume. On the other hand, on the display screens of list numbers 13 and 15, when the user speaks "onryou", the control unit 18 that recognized this causes a transition to the display screen for volume change, that is, the display screen corresponding to list number 14.

[0141] On the display screen of list number 15, when the user speaks "kido", the control unit 18 that recognized this completes the setting of the changed brightness. On the other hand, on the display screens of list numbers 13 and 14, when the user speaks "kido", the control unit 18 that recognized this causes a transition to the display screen for brightness change, that is, the display screen corresponding to list number 15.

[0142] In this way, even for the same recognition word, by changing the function according to the display screen, the user does not need to remember different recognition words in order to execute related but different functions.

[0143] Furthermore, in list numbers 13, 14, and 15, the recognition word "jitaku", which causes the route to home to be searched with a single utterance when returning home, is registered. In this way, since the recognition word corresponding to the function frequently used by the user is registered as the recognition word when voice recognition is activated, voice recognition can be started and immediately executed.

[0144] Note that for each list number from 1 to 15, recognition words that can be commonly recognized on multiple display screens and are used to execute various functions are registered. First, as phrases to return to the previous process, recognition words such as "modoru" (return), "torikeshi" (cancel), and "yarinaoshi" (redo) are registered. Also, as cancellation phrases to cancel functions and processes, recognition words such as "onseikaijo" (voice recognition cancellation), "shuuryou" (end), and "genzaichi" (transition to current position display) are registered. "Onseikaijo" is a recognition word for canceling voice recognition. "Shuuryou" is a recognition word for ending the current process. "Genzaichi" is a recognition word for transitioning to the current position display. These are recognition words that can be commonly recognized on all display screens after voice recognition is activated. By uttering these common recognition words, the user can redo the process or return to the initial state at any time.

[0145] Furthermore, for list number 16, recognition words for recognizing the user's Yes (affirmative) and No (negative) utterances are registered. This recognition word is used to recognize the user's simple Yes or No utterance in response to a question from the character. Such recognition words can also be commonly recognized on all screens. As the recognition words for Yes, "okkee" (okay), "yoroshiku" (please), "hai" (yes), "iesu" (yes), "un" (yes) are registered. As the recognition words for No, "iranai" (don't need), "sutoppu" (stop), "iie" (no), "noo" (no), "dame" (no good), "kiniiranai" (don't like), "urusai" (annoying) are registered. By preparing multiple types of recognition words for recognizing the user's Yes and No in this way, the possibility of not being recognized can be reduced even if the user's utterance content is somewhat different. In particular, the expressions of Yes and No vary depending on the person speaking, but the user can give Yes and No responses that match the character's tone.

[0146] As described above, the recognition words are, as will be described later, some that are displayed in characters on the display screen and some that are not displayed on the display screen. Those displayed in characters are good because the user can visually confirm what to say in order to make the control unit 18 execute its function.

[0147] 《Response Phrase》 Figures 7 to 21 are extracts of a part of a list of response phrases that the control unit 18 outputs as the voice of a character according to the recognition result of the above recognition words. The database 19 stores such a list of response phrases and voice data. Among the columns filled in gray in the list, the upper column is the number of the response phrase. The lower column is the content of the process executed by the control unit 18 according to the recognition word spoken by the user. "れーたん、ぼいすこんとろーる" outside the column is the activation word registered in the recognition word list number 1 of the above.

[0148] One or more phrases are registered for each process executed by the control unit 18 as response phrases. Figure 7 is an extract of a list of response phrases for "Standard Ray". The first column from the left in Figure 7 is a plurality of response phrases when the user speaks the activation word "れーたん、ぼいすこんとろーる" and the control unit 18 starts voice recognition. The second column is the response phrase when the control unit 18 turns down the volume. The third column is the response phrase when the user selects the facility No. 8 and the control unit 18 changes the destination to No. 8. The fourth column is the response phrase when the control unit 18 cancels the guidance. The fifth column is the response phrase when the control unit 18 changes the volume or the brightness level to 3.

[0149] Such a list of response phrases is set separately for each of the three aspects of the character, namely, the "standard Ray", the "transformed Ray", and the "miniature Ray". Depending on these three aspects, the response phrases in the list are different. Fig. 8(a) extracts the response phrases at the start of speech recognition for the "transformed Ray" from the list. The response phrases for the "transformed Ray", such as "Do you have any business?", "Speech recognition will start. Please speak slowly and clearly.", are set in a more polite tone compared to the response phrases for the "standard Ray". Also, Fig. 8(b) extracts the response phrases at the start of speech recognition for the "miniature Ray" from the list. The response phrases for the "miniature Ray", such as "Raytan, voice control start.", "I've been waiting.", are set in a more casual tone compared to the response phrases for the "standard Ray". Also, Figs. 9(a) to (i) extract the response phrases for the user's utterances. These specific response phrases will be explained in the processing examples described later. Note that as for the user's dialogue, for example, when the user's utterance and the character's voice are such that their semantic content is established as an utterance and a response to it, a feeling of mutual communication can be obtained. Even if the response to the utterance has a semantic content that is off-topic, the user can feel an element of interest or surprise. By mixing the case where the utterance and the response to it are semantically valid and the case where the response to the utterance has an off-topic semantic content, an imperfection similar to that of a real conversation can be produced, and the user can be made to feel a human touch.

[0150] In such a list, one or more response phrases are registered in the column for the same process. The control unit 18 randomly outputs, from among those registered in each column, a response phrase corresponding to the same process. However, the control unit 18 randomly selects from among those excluding the response phrases output previously, so that the same phrase is not repeatedly output. In particular, for processes that are frequently used, such as the start of speech recognition, as shown in Fig. 7, a large number of response phrases are prepared, and the control unit 18 randomly outputs from among them.

[0151] The operation example of this control unit 18 is as follows. 1. When there are five phrases A, B, C, D, and E, first randomly select one from the five. 2. If B is selected first, then select one from the four phrases A, C, D, and E next time. 3. If E is selected next, then select one from the three phrases A, C, and D next time. 4. After all selections are completed, return to the first step next time and randomly select one from the five phrases A, B, C, D, and E. In this embodiment, in response to the dialogue, as a predetermined process for outputting information that the user wants to know, a process of executing various functions listed below is performed. If the process according to the dialogue is a process of outputting information as a result of the dialogue, in order to execute the function, the desire to interact with the character arises. If the process according to the dialogue is a process of outputting information during the dialogue, in order to execute the function, the desire to proceed with the dialogue by speaking sequentially increases. If the process according to the dialogue is a process of outputting a response phrase that constitutes the dialogue, the user can enjoy the dialogue itself because they want to hear the character's voice and interact.

[0152] 《Surrounding Search Process》 Based on the recognition word as described above, the control unit 18 performs voice recognition to recognize the voice spoken by the user, and performs various processes according to the dialogue that outputs the voice of the character based on the response phrase. First, the surrounding search process according to such a dialogue will be described.

[0153] Surrounding search is a process of obtaining the target information by sequentially narrowing down from a plurality of candidates through a plurality of selections, and there are phases of genre selection, facility selection, and route selection. Such information is the information that the vehicle passengers want to know, especially the information that the vehicle driver wants to know.

[0154] Hereinafter, an example of the surrounding search process will be described with reference to FIGS. 10 to 21. FIGS. 10 to 21 are explanatory diagrams showing examples of display screens in each phase on the left and examples of corresponding recognition words and response phrases on the right. Each display screen M1 (including M1a), V1, V1a, S1, S1a, S2, S2a, S2_list, S2a_list, R1, R1a, R1b, A1 (including A1a), A2 (including A2a), A3 (including A3a) corresponds to the list numbers 1 to 15 of recognition words (FIGS. 5 and 6) as described above.

[0155] The descriptions shown on the left of the display screen examples in FIGS. 10 to 21 indicate "recognition word recognized as the user's voice → character response phrase output accordingly → display screen transitioned thereby".

[0156] The control unit 18 causes the display unit 5 to display these display screens by drawing on the memory the screen configuration data such as buttons, icons, texts, maps, etc. stored in the database 19. Also, as described above, the control unit 18 can perform surrounding search, route search, and route guidance.

[0157] (Waiting for voice recognition to start) First, when waiting for voice recognition to start for a predetermined process such as surrounding search, the control unit 18 is in a state where only the activation word can be recognized. The display screen M1 in FIG. 10 is an example of a screen waiting for voice recognition to start. On this display screen M1, an icon indicating the current position 51 is displayed together with a map of the surroundings of the current position 51. At the upper left of the display screen M1, the current time (10:00) is displayed, and below it, a scale button B1 indicating the current quantitative level of the map scale and the unit (100 m), and a voice recognition button B2 with an instruction (voice recognition) to activate voice recognition are displayed. At the lower right of the display screen M1, the upper body of the character 100 is displayed. The character 100 is in the standard pose, and its display position is such that it does not obstruct the display of various buttons, the current position, the current time, etc.

[0158] In this state of waiting for voice recognition to start, predetermined processes such as surrounding search cannot be performed by voice recognition. Therefore, a user who wants to change the map scale selects by touching the scale button B1. When the scale button B1 is selected in this way, as shown in the display screen M1a of FIG. 10, the control unit 18 displays a + button B1a and a - button B1b to the right of the scale button B1. When the user wants to increase the scale, the user selects by touching the + button B1a, and when the user wants to decrease the scale, the user selects by touching the - button B1b. In response to this, the control unit 18 gradually changes the scale of the displayed map.

[0159] In such a state of waiting for voice recognition to start, when the user speaks "re - tan, boi su kon to ro - ru", the control unit 18 outputs a corresponding response phrase and activates voice recognition. As a result, for example, the speaker 20 outputs a voice of "Do you have something to do?". Then, the control unit 18 changes the display screen M1 to the display screen V1 shown in FIG. 11.

[0160] In addition, even when the user touches and selects the voice recognition button B2, the control unit 18 outputs the above - mentioned response phrase and activates voice recognition. The control unit 18 displays a button corresponding to a function that can be selected either by manual operation or by voice input in a different color from other buttons on the display unit 5. For example, in the display screen of the present embodiment, buttons that can be selected by voice input are displayed in blue, and other buttons are displayed in black. Thus, the user can visually recognize functions that can be input by voice. Note that even for the same function, whether it can be selected by voice input depends on the display screen. For example, in the display screens M1 and M1a of FIG. 10, since the voice recognition function has not started, the scale button B1 is in a state where it cannot be selected by voice input and is displayed in black. And after the voice recognition function starts as described later, the scale button B1 becomes in a state where it can be selected by voice input and is displayed in blue.

[0161] (Voice Recognition Start) The display screen V1 in Fig. 11 is an example of a screen for starting voice recognition while waiting for navigation. Note that the display screen V1a in Fig. 12 is an example of a screen for starting voice recognition after the start of navigation guidance, which will be described later. When the control unit 18 activates voice recognition, it causes the map scale button to be displayed in blue on the display screen V1 to indicate that voice input is possible. In addition, the control unit 18 causes buttons for other functions that allow voice input to be displayed in blue on the display screen V1 to indicate that voice input is possible. Among these buttons, the voice cancellation button B2x is a button with the instruction "Voice Cancellation" for ending voice recognition displayed on it. The surrounding button B3 is a button with the instruction "Surroundings" for starting a surrounding search displayed on it. Since the surrounding search cannot function during navigation guidance, the control unit 18 does not display the surrounding button B3 on the display screen V1a in Fig. 12. The control unit 18 displays a guidance cancellation button B4x with the instruction "Cancel Guidance" for canceling the guidance displayed at the same position as the surrounding button B3 of V1 on the display screen V1a.

[0162] Also, on the display screens V1 and V1a, a volume button B5 and a brightness button B6 are displayed in the same column as the scale button B1, the voice cancellation button B2x, and the guidance cancellation button B4x. The volume button B5 is a button with the instruction "Volume" for starting volume change and the number "5" indicating the current sound volume level displayed on it. The brightness button B6 is a button with the instruction "Brightness" for starting brightness change and the number "5" indicating the current brightness level displayed on it. Furthermore, a home button B7 is displayed in the upper right corner of the display screen V1. The home button B7 is a button with the instruction "Home" for performing route guidance to the home as the destination displayed on it. These instructions and the numbers indicating the quantitative levels are phrases registered as recognition words as shown in Figs. 5 and 6 above.

[0163] As described above, the display area of an item composed of a plurality of buttons has a visible portion of the map which is the background image. That is, there are gaps between the buttons, and the background map is visible. On the other hand, the display area of the character 100 is larger than that of a single button, and in the overlapping portion with the map, there is no visible portion of the background. Note that in the gaps that occur between the hair and the body, the hands and feet and the torso, etc., there may be a visible portion of the background map, but such portions are excluded from the display area of the character 100.

[0164] The position of the character 100 to be displayed on the display unit 5 by the control unit 18 is a position that does not obstruct the visibility of the current position and the guidance route. In the display screen V1 of FIG. 11 and the display screen V1a of FIG. 12, the character 100 is displayed at a position closer to the lower right on the screen, and each of the above buttons is displayed at a position closer to the opposite side, the upper left or the upper right, in a left-right or up-down relationship with this.

[0165] (Genre Selection) In the above-described display screen V1 during navigation standby, when the user speaks "around", the control unit 18 transitions the display screen V1 to the display screen S1 of the vicinity 1 shown in FIG. 13 and outputs the voice of the response phrase "Hey, where are you going?". In this way, in the vicinity search that advances the process by making multiple selections, when the character 100 makes a speech prompting the next selection of genre selection, the user can immediately know what to do next. On the other hand, in the display screen V1a during navigation guidance, the vicinity search function does not work. For this reason, even when the user speaks "around", the control unit 18 that recognizes this outputs the voice of the character "What did you say? I don't understand." and does not transition to the screen of the vicinity 1. In this way, when a process that cannot be advanced is selected next, by the character 100 speaking a response phrase corresponding to this, the user can know that the process cannot be performed on the display screen.

[0166] As described above, the control unit 18 proceeds with the processing based on at least one round of dialogue, such as the input of voice to activate voice recognition from the user → the output of a response phrase of a character asking about the matter → the input of voice to start a function from the user. The output of the response phrase in this dialogue serves the function of indicating that the character has recognized the user's voice, along with the semantic content indicated by the response phrase.

[0167] As described above, on the display screen S1 in FIG. 13 transitioned by the user's utterance, the control unit 18 displays a genre button B8 on which the genre name as a recognized word is displayed. The genre names displayed on these genre buttons B8 correspond to the recognized words in list number 4 in FIG. 5. The user can also touch and select this genre button B8. Since the genre names are in hiragana or katakana without including Chinese characters, it is good because it is easy for the driver or the like to visually understand what to say.

[0168] On this display screen S1, when the user utters the recognized word "famiresu" displayed on the button, the control unit 18 causes the character to output a voice of the response phrase "Going to a family restaurant alone?". Then, the control unit 18 shifts to the process of searching for candidates for family restaurants. For this reason, as the phase of genre selection, there is no further need for the user to make any selection. That is, "Going to a family restaurant alone?" can be a voice that makes the user understand that no further genre selection is required. Also, this response phrase can make the user aware that they are alone other than the character 100 who is the conversation partner, and can enhance the intimacy with the character 100.

[0169] Also, for the case where the user wants to select from a genre other than the genre displayed on the display screen S1, the control unit 18 causes the display screen S1 to display a page button B9 on which a recognition word "2 pages" for transitioning to the next page is displayed. On this display screen S1, when the user speaks "nipeiji" or "tsugi", the control unit 18 causes the screen to be displayed on the display unit 5 to transition to the display screen S1a of the periphery 2 in FIG. 14.

[0170] On this display screen S1a, a genre button B8 on which a genre name corresponding to the recognition word of list number 5 in FIG. 5 is displayed is displayed. When the user speaks the recognition word of the genre name displayed on the genre button B8, the control unit 18 causes a response phrase for this to be output. Also, the control unit 18 causes the display screen S1a to display a page button B9 on which a recognition word "1 page" for transitioning to the previous page is displayed. On this display screen S1a, when the user speaks "ichipeiji" or "mae", the control unit 18 causes the screen to be displayed on the display unit 5 to transition to the display screen S1 of the periphery 1 in FIG. 13.

[0171] (Facility Selection) In this way, when a genre is selected, the control unit 18 performs a search for facilities around the current location belonging to the genre, and causes the screen to be displayed on the display unit 5 to transition to the display screen S2 of Facility Selection 1 as shown in FIG. 15. This display screen S2 displays the results of the facility search. That is, on the display screen S2, the facilities hit by the search are displayed by icons C1 numbered in order of proximity to the current location at their locations on the map. In FIG. 15, icons C1 with numbers 1 to 5 displayed inside a circle are displayed at positions on the map. In this way, the results of the surrounding search only display the facilities of the target (genre) necessary for the user. At this time, by setting the map display to grayscale and displaying the icon C1 in color, the location and number of the facility become clearer. Also, when transitioning from another screen to the screen of S2, the control unit 18 causes a response phrase prompting the selection of a destination such as "Where shall we go?" to be output.

[0172] The user selects a facility on the display screen S2 according to the content corresponding to the display of the icon C1. That is, when the user speaks the number displayed on the icon C1 of the desired facility, the control unit 18 outputs a corresponding response phrase and performs a route search to the facility. For example, when the user speaks "ichi" or "ichiban", the character 100 speaks "Number 1", and a route search to the facility displayed as number 1 is executed. In this way, by the character 100 speaking the number spoken by the user, the user can confirm the content of their own instruction.

[0173] Note that on the display screen S2, a scale button B1 and a list button B10 are displayed. As described above, the scale button B1 is a button for changing the map scale. On this display screen S2, when a user who wants to change the scale of the map speaks "suke-ru", the control unit 18 outputs a response phrase prompting which scale to change to, and transitions to the display screen S2a of facility selection 2 shown in FIG. 16. For example, when the user speaks "suke-ru", the control unit 18 outputs a response phrase "Please give an instruction", and as shown in FIG. 16, a + button B1a and a - button B1b for changing the scale are displayed. The scale change process by voice input will be described later.

[0174] The list button B10 is a button that changes the screen for selecting a facility from icon display to list display. On the display screen S2, when the user speaks "list" or "list display", the control unit 18 changes the screen to be displayed on the display unit 5 to the display screen S2_list of facility selection 1 shown in FIG. 17. On this S2_list display screen, the searched facilities are displayed not as icons but as a list arranged in the order of proximity to the current location by facility selection buttons B11. Each facility selection button B11 displays a number assigned in the order of proximity to the current location, the facility name, the distance from the current location to the facility, the location, etc. The user can select a facility in the same manner as above by touching the facility selection button B11 or speaking the number displayed on the button. Since the facility selection button B11 displays, in addition to the number, the facility name, the distance from the current location, the location, etc., the user can obtain more detailed information about the searched facilities.

[0175] When the number of searched candidates is more than the number that can be displayed on one page (for example, 5), the control unit 18 displays a page button B9 that displays "Page 2" indicating the existence of the next page. When the number of candidates is less than the number that can be displayed on one page, the control unit 18 hides the page button B9. When the user speaks "next page" or "next", the control unit 18 changes to the display screen S2a_list of facility selection 2 shown in FIG. 18. The selection process of the facilities in the list displayed on this display screen S2a_list is the same as above. Also, the control unit 18 displays a page button B9 that displays the recognition word "Page 1" for changing to the previous page on the display screen S2_list. On this display screen S1a, when the user speaks "page 1" or "previous", the control unit 18 changes the screen to be displayed on the display unit 5 to the display screen S1 of facility selection 1 in FIG. 16.

[0176] (Route Selection) As described above, when a facility is selected, the control unit 18 searches for a route 53 from the current position 51 to the selected facility as the destination 52, and displays the result on the route selection 1 display screen R1 shown in FIG. 19. The searched route 53 is color-coded to indicate the route 53 that should be given the highest priority based on the pre-set route conditions. At this time, the character 100 moves to a position on the map where it does not interfere, according to the route guidance direction. That is, the character 100 is displayed at a position that does not block the displayed searched route including the current position on the map. However, as described above, the character 100 is already displayed in the lower right of the screen, and in the example of the display screen R1, it does not interfere with the searched route, so there is no movement in this example.

[0177] On the other hand, the control unit 18 displays a plurality of buttons on the display screen R1 on the side opposite to the character, that is, on the left side. If the number of the plurality of buttons is small, there will be no part overlapping the searched route 53. Even if the number of buttons increases, there are gaps between the buttons, and the map and the route 53 can be seen. These buttons include a scale button B1. When the user wants to change the scale, the operation is performed on the route selection 2 display screen R1a shown in FIG. 20, which will be described later.

[0178] Furthermore, on the display screen R1, there is a route condition button B12 on which the instruction word "Route Condition" for changing route conditions is displayed. When the user speaks "conditions" or "route conditions", the control unit 18 causes a transition to the display screen R1b of route selection 3 shown in FIG. 21. On this display screen R1b, condition selection buttons B12a, B12b, and B12c for selecting search conditions to be prioritized, namely "highway", "ordinary road", and "recommended", are displayed beside the "Route Condition" button. On this display screen R1b, when the user speaks "ordinary road", "high-speed road", or "recommended", the control unit 18 outputs a response phrase corresponding to each, changes the search conditions to the conditions spoken, and displays the most prioritized route 53 under the changed search conditions. For example, if it is an ordinary road, the control unit 18 outputs a response phrase indicating the advantages of each search condition, such as "Yes, let's go slowly on the ordinary road.", and if it is a highway, "Time will be shortened on the highway." Also, if it is a recommended route, a response phrase indicating the trust relationship with the character, such as "Leave it to me.", is output.

[0179] And on the display screen R1, there is a guidance start button B13 on which the instruction word "Guidance Start" is displayed when the user wishes to receive guidance based on the displayed route 53. When the user speaks "start guidance", "please guide me", or "guide me" on the screen where the desired route 53 is displayed, the control unit 18 outputs a response phrase "Guidance start. Drive as I say." and starts route guidance.

[0180] Furthermore, when the user speaks "surroundings" to perform a surrounding search or a return phrase such as "return", "cancel", or "redo", the control unit 18 outputs a response phrase "Re-search" and causes a transition to the display screen S1 shown in FIG. 13. The subsequent surrounding search process is as described above. For the utterance of the return phrase, in the display screens S2, S2a, S2_list, S2a_list, R1, R1a, and R1b shown in FIGS. 15 to 21, the same process is performed.

[0181] In addition, for the utterance of the return phrase on the display screens V1 and V1a for starting voice recognition shown in FIGS. 11 and 12, the control unit 18 outputs a response phrase "Please call again." and causes a transition to the display screen M1 shown in FIG. 10 to return to waiting for voice recognition activation. For the utterance of the return phrase on the display screens S1 and S1a for peripheral search shown in FIGS. 13 and 14, the control unit 18 outputs a response phrase "From the beginning." and causes a transition to the display screen V1 for starting voice recognition shown in FIG. 11. Thus, even for the same recognition word, depending on the current display screen and the corresponding phase, the display screen to be transitioned is different. That is, by changing the display screen to be transitioned for the return phrase, it is possible to prevent over-returning and suppress the trouble for the user to do it again.

[0182] And in this embodiment, in addition to the return phrase, a cancellation phrase is prepared. When the user utters "voice cancellation", "termination", or "current location", it always returns to the display screen M1 waiting for voice recognition activation. At this time, the control unit 18 outputs a response phrase "Please call again." to encourage the next use. In this way, by registering a recognition word with a different return position according to the display screen and phase and a recognition word that returns to a certain position, the user can quickly transition to the desired screen according to the situation.

[0183] Furthermore, when a recognition word that cannot be processed is uttered according to the screen and phase, the control unit 18 can prompt the user to utter another word by outputting a response phrase indicating that it cannot be processed with that recognition word. For example, on the display screen V1a in FIG. 12 and the display screens S2, S2a, S2_list, and S2a_list in FIGS. 15 to 18, a response phrase "What did you say? I don't understand." is output.

[0184] 《Map Scale Change》 In response to the dialogue, as a process of changing functions, the control unit 18 performs processes of changing the map scale, volume, and brightness. First, the procedure for changing the map scale will be described with reference to FIG. 22. Note that when transitioning from the display screen S2 in FIG. 15 to the display screen S2a in FIG. 16, and when transitioning from the display screen R1 in FIG. 19 to the display screen R1a in FIG. 20, the change in the map scale is the same.

[0185] When a user who wants to change the map scale speaks "suke-ru" on the display screens V1 and V1a in FIGS. 11 and 12 where voice recognition has started, the control unit 18 causes a transition to the display screens A1 and A1a in FIG. 22. The display screen A1 is a screen that has transitioned from the display screen V1, and the display screen A1a is a screen that has transitioned from the display screen V1a. On the display screen A1, only when transitioning from another screen (scene), the control unit 18 outputs a response phrase "Please give an instruction." In this way, even when the screen is switched, by having the character 100 make a speech prompting an instruction, the user can be made to be aware that it is always beside them, which is good. Also, the display screen A1a is a screen for changing the map scale during navigation guidance.

[0186] On these display screens A1 and A1a, a + button B1a and a - button B1b are displayed to the right of the scale button B1. As described above, when the user touches the + button B1a, the control unit 18 gradually increases the map scale. Also, when the user touches the - button B1b, the control unit 18 gradually decreases the map scale. Further, in this embodiment, when the user speaks "kakudai" or "shousai", the control unit 18 that recognizes this gradually increases the map scale. When the user speaks "shukushou" or "kouiki", the control unit 18 that recognizes this gradually decreases the map scale.

[0187] Then, when scaling up, the control unit 18 outputs a response phrase that allows the user to confirm what changes have been made, such as "It's enlarged." When scaling down, it outputs a response phrase like "I made it a bit wider." This enables the user to confirm audibly what processing has been done. Also, when changing the scale would cause it to exceed the upper or lower limit, the control unit 18 outputs a response phrase indicating that the change cannot be made, such as "Boo~, it's at the limit." This prevents the user from repeatedly giving futile change instructions without being able to confirm whether the upper or lower limit has been reached.

[0188] Also, when the user speaks a desired scale value, such as "jyuumeetoru" or "ichikiro", the control unit 18 that recognizes this directly changes the map scale to the spoken value. At this time, as a response phrase corresponding to the user's speech, the control unit 18 outputs a response phrase corresponding to the input value, such as "10 meters" or "1 kilometer". This allows the user to confirm the scale level they have instructed.

[0189] Then, when the user speaks "sukeeru" or a return phrase, the control unit 18 outputs a response phrase indicating that the scale change process is complete, such as "The setting is complete." The processing for the release phrase is the same as above.

[0190] Also, in the case of the display screen A1, when the home button is displayed and the user touches this button or says "jitaku" (home in Japanese), the control unit 18 processes as follows. First, if home registration has not been done in advance, the control unit 18 outputs a response phrase prompting home registration, such as "You said home, but you haven't told me yet, right?" This prompts the necessary processing as a prerequisite for executing a specific function through voice, so that the user can understand what is required to use the function. If home registration has been done, the control unit 18 outputs a response phrase "Yes, let's go home." and searches for a route home. This way, the user can get the impression that the character will go home with them.

[0191] Note that in the case of the display screen A1a, since navigation guidance is in progress, specific functions are restricted. That is, around search and route guidance to home cannot be executed on the display screen A1a. Therefore, even if the user says "shuuhen" (around in Japanese), the control unit 18 outputs a response phrase "What did you say? I don't understand." and does not perform an around search. Also, even if the user says "jitaku", the control unit 18 outputs a response phrase "What did you say? I don't understand." and does not search for a route home. In this way, the user can know through the response phrase that the process cannot be performed.

[0192] 《Volume Change》 The procedure for changing the volume will be described with reference to FIG. 23. When a user who wants to change the volume speaks "volume" on the display screens V1 and V1a of FIGS. 11 and 12 where voice recognition has started, the control unit 18 causes a transition to the display screens A2 and A2a of FIG. 23. The display screen A2 is a screen that has transitioned from the display screen V1, and the display screen A2a is a screen that has transitioned from the display screen V1a. On the display screen A2, only when transitioning from another screen (scene), the control unit 18 outputs a response phrase "Please give an instruction." In this way, even when the screen is switched, by having the character 100 make a speech prompting an instruction, the user can be made to be aware that it is always by their side. Also, the display screen A2a is a volume change screen during navigation guidance.

[0193] On these display screens A2 and A2a, a + button B5a and a - button B5b are displayed to the right of the volume button B5. When the user touches the + button B5a, the control unit 18 gradually increases the volume. When the user touches the - button B5b, the control unit 18 gradually decreases the volume. Also, in this embodiment, when the user speaks "up" or "bigger", the control unit 18 that recognizes this gradually increases the volume. When the user speaks "down" or "smaller", the control unit 18 that recognizes this gradually decreases the volume.

[0194] Then, when increasing the volume, the control unit 18 outputs a response phrase that allows the user to confirm what change has been made, such as "Volume up~". When decreasing the volume, it outputs "Volume down~". Thus, the user can confirm by voice what process has been performed. Also, when the volume change causes the volume to exceed the upper limit or lower limit, the control unit 18 outputs a response phrase informing that the change cannot be made, such as "Boo, it's the limit though." This can prevent the user from repeatedly making useless change instructions without being able to confirm whether the upper limit or lower limit has been reached.

[0195] Also, when the user speaks a desired volume value such as "zero", "one", or "two", the control unit 18 that recognizes this directly changes the volume to the spoken value. At this time, the control unit 18 outputs a response phrase corresponding to the input value, such as "It's zero", "It's one", or "It's two", as a response phrase corresponding to the user's speech. Thereby, the user can confirm the volume level they have instructed.

[0196] Then, when the user speaks "volume" or a return phrase, the control unit 18 outputs a response phrase notifying that the volume change process has been completed, such as "The setting is complete." The process for the cancel phrase is the same as above. Other processes are the same as the processes described for the map scale above.

[0197] 《Brightness Change》 The procedure for brightness change will be described with reference to FIG. 24. On the display screens V1 and V1a in FIGS. 11 and 12 where voice recognition has started, when the user who wants to change the brightness speaks "brightness", the control unit 18 causes a transition to the display screens A3 and A3a in FIG. 24. The display screen A3 is the screen that has transitioned from the display screen V1, and the display screen A3a is the screen that has transitioned from the display screen V1a. On the display screen A3, only when transitioning from another screen (scene), the control unit 18 outputs a response phrase "Please give an instruction." In this way, even when the screen is switched, by having the character 100 make a speech prompting an instruction, the user can be made to be aware that it is always beside them, which is good. Also, the display screen A3a is the brightness change screen during navigation guidance.

[0198] On these display screens A3 and A3a, a + button B6a and a - button B6b are displayed beside the brightness button B6. When the user touches the + button B6a, the control unit 18 gradually increases the brightness. When the user touches the - button B6b, the control unit 18 gradually decreases the brightness. Also, in this embodiment, when the user says "appu" or "akaruku", the control unit 18 that recognizes this gradually increases the brightness. When the user says "daun" or "kuraku", the control unit 18 that recognizes this gradually decreases the brightness.

[0199] Then, when increasing the brightness, the control unit 18 outputs a response phrase that allows the user to confirm what change has been made, such as "Brightness up~" when increasing and "Brightness down~" when decreasing. This enables the user to confirm audibly what process has been performed. Also, when changing the brightness would cause it to exceed the upper or lower limit, the control unit 18 outputs a response phrase informing the user that the change cannot be made, such as "Boo, it's the limit though." This prevents the user from repeatedly giving unnecessary change instructions without being able to confirm whether the upper or lower limit has been reached.

[0200] Also, when the user says a desired brightness value such as "zero", "one", "two", the control unit 18 that recognizes this directly changes the brightness to the spoken value. At this time, as a response phrase corresponding to the user's speech, a response phrase corresponding to the input value, such as "It's zero", "It's 1", "It's 2", is output. This enables the user to confirm the brightness level they have instructed.

[0201] Then, when the user says "kido" or utters a return phrase, the control unit 18 outputs a response phrase informing the user that the brightness change process is complete, such as "The setting is complete." The processing for the release phrase is the same as above. Other processing is the same as the processing described for the map scale above.

[0202] In the display screens A1 and A1a, the control unit 18 causes a transition to the display screens A2, A2a or the display screens A3, A3a by speaking "onryou" or "kido". In the display screens A2 and A2a, the control unit 18 causes a transition to the display screens A3, A3a or the display screens A1, A1a by speaking "kido" or "sukeru". In the display screens A3 and A3a, the control unit 18 causes a transition to the display screens A1, A1a or the display screens A2, A2a by speaking "sukeru" or "onryou". In this way, by making the screens for the quantity level change processing mutually transitionable, it can be convenient for users who want to change the settings all at once.

[0203] 《Timeout Processing》 Next, the timeout processing will be described. The timeout processing is a process in which the control unit 18 ends the process when a predetermined time has elapsed without voice input from the user since the start of a predetermined process. And, until a predetermined time elapses from the start of the predetermined process, the control unit 18 outputs a voice prompting voice input as a response phrase of the character.

[0204] A more specific process will be described with reference to FIG. 25. As the timeout processing, timeouts V1, S1, S2, R1, A1, and A1a are set. First, the process of timeout V1 will be described. When there is no voice recognition for 15 seconds after the control unit 18 activates voice recognition on the above display screen V1 or V1a, as the timeout status 1, a response phrase prompting the user's voice input is output. Here, since it is the screen at the start of voice recognition, a response phrase asking for a voice input such as "Hey, say something" is output. And when there is no voice recognition for 15 seconds from the output of this response phrase, the control unit 18 outputs a response phrase notifying that the voice recognition has ended as the timeout status 2 and ends the voice recognition. Here, a response phrase that seems a little angry about the absence of voice input, such as "Humph, already." is output.

[0205] The processing of timeout S1 is as follows. After the control unit 18 transitions to the display screens S1 and S1a, if there is no voice recognition for 15 seconds, it outputs a response phrase prompting voice input as the timeout status 1. Here, since it is a screen for surrounding search, it outputs a response phrase asking about the destination, such as "Hey, hey, where are you going?". Then, if there is no voice recognition for 15 seconds from the output of this response phrase, the control unit 18 outputs a response phrase notifying that the voice recognition has ended as the timeout status 2, and ends the voice recognition. Here, it outputs a response phrase that seems a bit angry about the lack of voice input, such as "Humph, already.".

[0206] The processing of timeout S2 is as follows. After the control unit 18 transitions to the display screens S2, S2a, S2_list, and S2a_list, if there is no voice recognition for 15 seconds, it outputs a response phrase prompting voice input as the timeout status 1. Here, since it is a screen for facility selection, it outputs a response phrase prompting facility selection, such as "Still not decided?". Then, if there is no voice recognition for 15 seconds from the output of this response phrase, the control unit 18 outputs a response phrase notifying that the voice recognition has ended as the timeout status 2, and ends the voice recognition. Here, it outputs a response phrase that seems a bit angry about the lack of voice input, such as "Humph, already.".

[0207] The processing of timeout R1 is as follows. After the control unit 18 transitions to the display screens R1, R1a, and R1b, if there is no voice recognition for 15 seconds, it outputs a response phrase prompting voice input as the timeout status 1. Here, since it is a screen for route selection, it outputs a response phrase notifying that the next process will proceed even without voice input, such as "I'll guide you if you're silent." Then, if there is no voice recognition for 15 seconds from the output of this response phrase, the control unit 18 outputs a response phrase notifying the end of voice recognition as the timeout status 2 and terminates the voice recognition. Here, it outputs a response phrase notifying the start of the next process, such as "Guiding start. Can you drive as I say?"

[0208] The processing of timeout A1 is as follows. After the control unit 18 transitions to the display screens A1, A2, A3, A1a, A2a, and A3a, if there is no voice recognition for 15 seconds, it outputs a response phrase prompting voice input as the timeout status 1. Here, since it is a screen for setting change, it outputs a response phrase prompting setting change, such as "Please give instructions." Then, if there is no voice recognition for 15 seconds from the output of this response phrase, the control unit 18 outputs a response phrase notifying the end of voice recognition as the timeout status 2 and terminates the voice recognition without changing the settings. Here, it outputs a response phrase that seems slightly angry about the lack of voice input, such as "Well, come on."

[0209] The processing of timeout A1a is as follows. After the control unit 18 transitions to the display screens A1, A2, A3, A1a, A2a, A3a and no input for instructing a setting change is made by button selection or voice, if there is no voice recognition for 15 seconds, as the timeout status 1, a response phrase prompting voice input is output. Here, a response phrase that prompts a determination instruction for setting change by voice input, such as "Instruction, please", is output. Then, if there is no voice recognition for 15 seconds, the control unit 18 outputs, as the timeout status 2, a response phrase notifying that the voice recognition has ended, completes the setting change, and ends the voice recognition. Here, a response phrase notifying the completion of the setting change, such as "The setting is completed.", is output.

[0210] Although timeout processing may occur frequently, as described above, by changing the content of the response phrases for timeout status 1 and timeout status 2 according to the display screen and phase, it is possible to prevent the user from getting bored.

[0211] [Advantages and Effects] According to the above-described embodiment, instead of the system providing information in response to a one-sided instruction from the user who is a vehicle occupant, the user can obtain the information they want by interacting with the character. Therefore, the fun of operation and driving increases, and a feeling of communicating with the character is obtained, preventing boredom. In addition, since the user is motivated to operate the system in order to interact with the character, the functions of the system are effectively utilized.

[0212] Since the response phrases are registered so that a predetermined character is recognized as a speaker with a specific personality by the voice output from the system, the user can feel as if they are interacting with an actual person and feel familiar.

[0213] As for the dialogue, for example, it is set so that at least one round of interaction, such as the start of a speech, a response to it, and a reply to the response, is made. Therefore, rather than simply having the system process according to the user's command, the user can get the feeling that the character intervenes between himself and the system and provides the service.

[0214] By using voice recognition, the user can output more easily or quickly than by manual operation. Therefore, the user will actively use voice recognition, and more opportunities can be given to the user to interact with the character.

[0215] In particular, even when it is difficult for the driver to obtain desired information during driving due to difficult manual operation, restricted manual operation, etc., by using voice recognition, the user can obtain the desired information while enjoying the interaction with the character.

[0216] When it is necessary to go through multiple selections to output information, compared with the case where manual operation is required for each selection, voice input is easier for the user to operate. As information output through multiple selections, for example, information with a hierarchical structure or information that can only be output after multiple operations places less burden on the user than performing multiple manual operations.

[0217] In particular, a process of outputting information obtained by sequentially narrowing down from a plurality of candidates, such as facility search, search around the current location, etc., is different from operations for moving the vehicle itself, such as brakes and accelerators, or operations for notifying something externally for safety during driving, such as turn signals. It is difficult to directly execute by manual operation, so it is suitable for voice recognition operations.

[0218] By being prompted by the voice of the character to make a selection, the user can sequentially proceed with a predetermined process. For example, as a response phrase to guide the user to the next selection, when the user's voice "around" is recognized in a nearby search, a voice asking about the destination, such as "Hey, where are we going?", is output, so that it becomes clear to the user what to say next.

[0219] When there is no next selection, the voice of the character indicating this is output, so that the user can understand that there is no need to further proceed with the process and can make a judgment such as ending the instruction or causing another process to be performed. For example, as the voice of the character indicating that there is no next selection, when the user's voice "family restaurant" is recognized in a nearby search, by saying something like "Going to a family restaurant alone?", a new judgment can be requested from the user.

[0220] The user can obtain information on facilities around a desired location such as the current location while enjoying the interaction with the character without the need for troublesome manual operations.

[0221] The icon indicating the searched nearby facilities is relatively highlighted compared to the map, so that the information on the necessary facilities stands out for the user on the map, making it easier for the user to grasp the location of the facilities and serving as a guide when selecting a desired facility by voice. In particular, since the icon is displayed in color and the map is displayed in grayscale, the icon can be seen very clearly.

[0222] The icon of the facility is displayed with a distinguishable indication of the ranking of the proximity to the current location, so that it can serve as a guide for the user to select the facility. In particular, since it uses information that inherently indicates a ranking, such as numbers and alphabets, the meaning of the ranking can be easily grasped.

[0223] By recognizing the voice of a user who selects any facility according to the displayed icon and searching for and displaying the route to the selected facility, the user can easily select according to the content corresponding to the displayed icon rather than selecting by facility name. In particular, since the user only needs to speak the displayed number, the user can select a facility with a word that is easier to remember than the facility name.

[0224] When the user does not want the searched route, the user can instruct a re-search while interacting with the character, so that the trouble and boredom during re-search can be alleviated.

[0225] When the user sees the appearance of the character who is the interaction partner on the display screen, the user will feel familiar and will want to actively use the system to see the appearance.

[0226] And the character is displayed at a position that does not block the displayed searched route on the map, so that while enjoying the display of the character, the visibility of the searched route can be maintained. When the character moves closer to the side opposite to the searched route, the display area of the item may overlap with the searched route. However, each item is smaller than the character, and the display area of the item has a part where the background image can be seen. Therefore, the background image is more visible than the searched route overlapping with the character.

[0227] The voice of the character is randomly selected and output from among a plurality of different preset response phrases, excluding the response phrases that have already been output, so that not only the same phrases are not output, and the user does not get bored.

[0228] Since a list in which a plurality of different response phrases are registered for a common recognition word is set, even if the user makes a speech with common content, the character can reply with different expressions, and even when the same process is realized many times, the user can enjoy what kind of reply will come from the character.

[0229] The character has multiple display modes while being the same person in the standard form, transformed form, and chibi form. As for the output voice, different phrases are registered for each display mode. Therefore, even for the same character, the utterance content varies according to the multiple display modes, and while it is a conversation with the same person, the user can enjoy the changes. In particular, by making it the standard form, the transformed form with changed clothes, hairstyle, etc., and the form with changed body proportions, various impressions such as adult-like, child-like, cute, and beautiful can be given to the user, preventing boredom.

[0230] Since the recognition words are displayed as items on buttons or the like, the user can see the displayed items to know what to say in order to execute a predetermined process.

[0231] There are recognition phrases that can be commonly recognized regardless of the display screen and recognition words whose recognition is restricted to what is displayed on the display screen among the recognition words. Therefore, by using the recognition phrases that can be commonly recognized, the convenience of the functions required regardless of the display screen is maintained, and by using the recognition words whose recognition is restricted, the possibility of misrecognition can be reduced. For the recognition words commonly used on multiple types of display screens, the possibility that the function is restricted according to the display screen is low, and the convenience can be maintained. In particular, since they are phrases for transitioning to other display screens, such as the return phrase and the cancel phrase, they are phrases that need to be used on any screen. On the other hand, the recognition words displayed on the display screen and the recognition words to be searched have little need to recognize other phrases, and there is a high need to reduce the number of recognition words to prevent misrecognition.

[0232] The state of the function that changes quantitatively can be directly set to the desired quantitative level by voice recognition. Therefore, compared with the case of changing step by step between a small value and a large value, the desired value can be set immediately, saving the user's effort and shortening the time restricted by the operation. In particular, it is possible to save the effort of selecting buttons such as volume, brightness, and scale, displaying +- or a scale, and following the steps to increase or decrease by selecting +- or sliding on the scale.

[0233] As recognition words, an indicator corresponding to a function that changes quantitatively and a number indicating the quantitative level are set. So, when the user utters the indicator and the number corresponding to the desired function, the quantitative level of the function can be directly increased or decreased to the level of that number. In particular, using "onryou", "kido", "sukeeru" as indicators and "ichi", "ni", "nikiro" as numbers, etc., the user can input using function names and numbers that are intuitively easy to understand.

[0234] Since an indicator corresponding to the function is displayed on the button for selecting the function, the user can select the function by uttering the indicator displayed on the button, which is easy for the user to understand.

[0235] For the button for selecting the function, since the value of the current quantitative level of the function is displayed, the user can check the current quantitative level by looking at the value displayed on the button and determine whether to increase or decrease it to the desired value. For example, in the case of the scale of a map, when a small scale value is displayed on the button and the user utters a larger value, the scale value displayed on the button also directly changes to that value, and the user can immediately view a wide - range map. Also, for example, when a large scale value is displayed on the button and the user utters a smaller value, the scale value displayed on the button also directly changes to that value, and the user can immediately view a detailed map.

[0236] When no voice input is received from the user and a predetermined time has elapsed since the start of a predetermined process, a timeout process for ending the process is performed, so it is possible to prevent incorrect operations caused by continuous voice recognition for a long time.

[0237] From the time when a predetermined process is started until a predetermined time elapses, a character outputs a voice prompting the user for voice input, so that when the user forgets to start the process, the user can be prompted for voice input.

[0238] [Other Embodiments] The present invention is not limited to the above-described embodiments, and can be configured in various forms including the forms exemplified below.

[0239] 《Genre Registration》 A function of genre registration will be described in which a user causes only desired items to be displayed on the display screen of the display unit 5 for the items of the genre of the surrounding search. By preliminarily tabulating the genres that the user will frequently use in this way, an original surrounding search screen for the user can be created. More specifically, the control unit 18 causes the display unit 5 to display a table of items for selecting which of the items corresponding to the recognition words of the genre will function as the recognition word. When the user selects a desired item from this table, the control unit 18 sets the selected item to a page for button display on the display unit 5.

[0240] Such an example of the display screen will be described with reference to FIG. 26. First, in the display screen (1), in the same manner as in the above-described embodiment, the control unit 18 causes the character 100 to be displayed superimposed on the lower right of the map screen displayed on the display unit 5, and causes the surrounding button B3 for selecting and searching for surrounding facilities by genre to be displayed superimposed on the left of the map screen.

[0241] In this display screen (1), when the user touches the surrounding button B3 or utters the two words "rei" and "shuhen" registered as recognition words for activating voice recognition and instructing a surrounding search, the control unit 18 that has recognized this causes a transition to a one-shot genre selection screen shown in the display screen (2).

[0242] As described above, the display screen (2) displayed by activating two-word voice recognition displays a table composed of a plurality of items by arranging a plurality of buttons displaying genre names. When the user selects each button or utters the genre name from the table of items of this genre name, the control unit 18 searches for the facilities of that genre as shown in the display screen (3), numbers the searched facilities in ascending order from the current location 51, and displays them on the map.

[0243] When the user touches a desired facility icon displayed on the display screen (3) or speaks the number displayed on the desired facility icon, the control unit 18 executes a route search from the current location to the selected facility and, as shown on the display screen (4), displays the searched route 53 in color-coded form.

[0244] At this time, the character 100 moves to a position on the map where it does not get in the way, depending on the route guidance direction. That is, the character 100 is displayed at a position that does not block the displayed searched route including the current position on the map. Further, when the user speaks to start the guidance, as shown on the display screen (5), the control unit 18 starts the route guidance. Even during this route guidance, the character 100 is displayed at a position that does not block the display of the searched route 53 including the current position 51 on the map.

[0245] Three buttons, namely, "Return", "Change", and "Current Location", are displayed at the lower part of the table on the display screen (2). Among these, the "Return" button is a button that, when the user touches it or speaks "Return", the control unit 18 that recognizes this causes a transition to the previous screen. The "Current Location" button is a button that, when the user touches it or speaks "Current Location", the control unit 18 that recognizes this causes a transition to the display screen that displays the current location 51. Here, whether the "Return" button is selected or the "Current Location" button is selected, a transition is made to the display screen (1).

[0246] The "Change" button is a button that, when the user touches it or speaks "Change", the control unit 18 that recognizes this causes a transition to the one-shot genre change screen shown on the display screen (6). On this one-shot genre change screen, the genres currently registered as the one-shot genre selection screen are displayed in the same order. Here, when the user touches a button of any genre or speaks the genre name, the control unit 18 that recognizes this causes a transition to any one of the display screens (7) to (10) of the major item to which each genre name belongs.

[0247] For example, when the user selects "Station", the screen transitions to display screen (7) where the genre names belonging to the major item "Hotel, Public Facility, Others" to which "Station" belongs are displayed. On display screen (7), the genre names displayed on display screen (2) are shown in a specific dark color, and the genre names not displayed on display screen (2) are shown in a lighter color. For example, "Station" displayed on display screen (2) is shown in blue characters, and "Parking Lot" not displayed on display screen (2) is shown in gray characters.

[0248] Here, when the user touches or speaks the genre name displayed on display screen (7), the control unit 18 that recognizes this sets the selected genre name as the genre name not displayed on display screen (7). At this time, the color of the genre name is changed to a lighter display. For example, when the user selects "Station", the characters of "Station" become gray.

[0249] For example, when the user touches or speaks the genre name displayed on display screen (7), the control unit 18 that recognizes this sets the selected genre name as the genre name to be displayed on display screen (2). At this time, the color of the genre name is changed to a specific dark color display. For example, when the user selects "Parking Lot", the characters of "Parking Lot" become blue.

[0250] Also, when a genre name belonging to a hospital such as "General Hospital" on display screen (6) is selected, the screen transitions to display screen (8). When a genre name belonging to shopping such as "Convenience Store" is selected, the screen transitions to display screen (9). When a genre name belonging to a restaurant such as "Family Restaurant" is selected, the screen transitions to display screen (10). In each display screen, the selection of the genre names to be displayed on display screen (2) and the selection of the genre names not to be displayed are as described above.

[0251] Also, on the display screen (7), when the user selects the "▽" button to send a page, the screen transitions to the next display screen (8). On the display screen (8), when the user selects the "△" button to go back a page, the screen transitions to the display screen (7). In this way, "▽" and "△" function as buttons for switching pages among the display screens (7) to (10).

[0252] As described above, on each of the display screens (7) to (10), after the genre name is changed, when the user touches the "Back" button or says "Go back", the screen transitions to the display screen (6), and the selected genre name to be displayed is shown while the selected genre name not to be displayed is no longer shown. Then, when the user touches the "Back" button or says "Go back", the screen transitions to the display screen (2) that reflects the display of the changed display screen (6). In the above example, "Parking lot" is displayed at the position of "Station" and "Station" is no longer displayed.

[0253] Also, to return to the default display, on the display screen (6), when the user selects the "Return to Default" button or says "Return to default", the control unit 18 that recognizes this returns to the default genre name and its layout. Then, when the user touches the "Back" button or says "Go back", the screen transitions to the display screen (2) that has returned to the default settings.

[0254] As described above, in the display screen (6) and the display screen (2) which reflects it, the genre name can be selected for each display position. That is, in the display screen (6), when the "station" in the upper left is selected, the control unit 18 also recognizes which display position on the display screen it is together with the selected genre name. Then, the control unit 18 can display any genre name selected in the subsequent display screens (7) to (10) at the said display position, that is, the upper left position. Therefore, when it is desired to change the order of the genre names in the display screen (2), it can be realized by selecting the genre name at the desired display position in the display screen (6) and changing it to the genre name to be displayed at the said display position. In this case, it is also possible to select to display the same genre name at different positions. For example, the genre name "station" can be displayed in the upper left and "station" can also be arranged below it. In this way, for frequently used genre names, if the same genre name is displayed at multiple locations on the display screen (2), it will be easier to touch and select that genre name.

[0255] Note that if the control unit 18 allows the user to touch and drag a desired genre name and release it at a desired position in the display screen (6), so that the control unit 18 can change the display position of the said genre name to a desired position, the change of the display position will be easier.

[0256] Also, the number of genre names that can be displayed on the display screen (2) is limited to a predetermined number. For example, here it is 25. Therefore, when 25 genre names have already been registered as the genre names to be displayed, it is not possible to increase the number of genre names to be displayed any further.

[0257] As described above, by selecting only the items frequently used by the user from the displayed item list, the recognized words are limited, so misrecognition can be reduced. The selected recognized words are registered, for example, by rewriting them from the default state in a predetermined memory area, putting them in an empty area, rearranging the existing ones, etc. In this way, the limited memory area can be effectively utilized.

[0258] Such setting of items by the user may be set for each of the plurality of pages that display the items selected by the user on the display unit 5 as shown in FIGS. 13 and 14 of the above embodiment. For this reason, even if the number of items that can be displayed on one page of the display screen is limited, by bringing the items frequently used by the user to the top page, the trouble of switching pages to display the desired items can be saved.

[0259] Since the large items corresponding to the recognized words including a plurality of concepts and the middle items corresponding to the recognized words corresponding to the plurality of concepts can be displayed on the same page, even when the user wants to search for a middle item, there is no need to change the hierarchy or page, and the desired item can be directly selected.

[0260] For example, a high-level concept large item including a large number of candidates, such as "hospital", is a convenient item for a user who wants to search from as many candidates as possible. Middle items such as "Internal Medicine" and "Pediatrics" included in the concept of "hospital" are convenient for users who want to immediately obtain information on the desired facility from a small number of candidates.

[0261] Information regarding the classification of facilities is made such that efficient designation can be performed by classifying and separating them in stages and hierarchies like large items and middle items. That is, one or more middle items belong to a large item.

[0262] For example, examples of major items include "Hotel, Public Facilities, Others", "Hospital, Clinic", "Shopping", "Restaurant", etc. Examples of middle items belonging to "Hotel, Public Facilities, Others" include "Station", "IC (Interchange) / JCT (Junction)", "Post Office", "Government Office", "Toilet", "Parking Lot", "Hotel", "Bank / Credit Union", "School", "Gas Station", etc. Examples of middle items belonging to "Hospital, Clinic" include "General Hospital", "Internal Medicine", "Pediatrics", "Surgery", "Obstetrics and Gynecology", "Dentistry", "Ophthalmology", "Otolaryngology", "Orthopedics", "Dermatology", "Urology", "Diagnostic Internal Medicine", "Neurology", "Proctology", "Veterinary Hospital", etc. Examples of middle items belonging to "Shopping" include "Convenience Store", "Drugstore", "Department Store", "Shopping Mall", "Supermarket", "Automobile Supplies", "Home Center", "Household Appliances", etc. Examples of middle items belonging to "Restaurant" include "Family Restaurant", "Fast Food", "Japanese Cuisine", "Western Cuisine", "Chinese Cuisine", "Ramen", "Grilled Meat", "Sushi", "Pork Cutlet", "Udon", "Curry", "Beef Bowl", etc.

[0263] Note that each middle item may further belong to one or more minor items. By classifying and separating in a hierarchical manner such as major items, middle items, and minor items, efficient designation can be performed from a large number of items. When only one item belongs, in fact, the item at that level is equivalent to not existing. That is, when there is only one middle item belonging to a major item, it may be considered that there are no middle items and minor items up to the major item. Also, when there is only one minor item belonging to a middle item, it may be considered that there are no minor items up to the middle item. When a major item is designated, it means that all the middle items belonging to it are designated. When a certain middle item is designated, the control unit 18 processes it as if all the minor items belonging to that middle item are designated.

[0264] In addition, as recognition words, words related to the phrases displayed as items but not displayed as items may be registered. For example, even if the user does not utter the recognition word itself displayed in the item, it is sufficient to say a word related to the recognition word. Therefore, even if the user's memory of the item name is vague or the user is not familiar with the operation, the search can be performed.

[0265] For example, even on a page where the phrase displayed as an item is the large item "hospital" and the medium items "internal medicine" and "pediatrics" are not displayed, "internal medicine" and "pediatrics" may also be registered as voice-recognizable recognition words. In this way, even if the medium items are not displayed, if the more specific medium items "internal medicine" and "pediatrics" associated with the large item "hospital" are spoken, the medium items can be searched, and thus the user can quickly obtain the information on the facilities required.

[0266] Depending on the user's settings as described above or the default settings, the items of recognition words that may be misrecognized may be set to be displayed on different pages. For example, phrases containing similar sounds, phrases that are likely to be misrecognized when actually recognized, etc. will not be candidates for recognition simultaneously on the same page, so misrecognition can be reduced. Examples of phrases containing similar sounds are phrases containing the same sound in hiragana, the same Chinese characters, etc., such as "famiresu" and "fast food", "hospital" and "beauty salon", "western food" and "Japanese food".

[0267] In this way, since the recognition phrases with confusing sounds are not displayed on the same page, the user will input one of them, and misrecognition can be reduced. Examples of phrases that are likely to be misrecognized when actually recognized are different phrases recognized for the same utterance based on the record of past recognition results. In this way, it is possible to prevent the combination where misrecognition actually occurs from being on the same page.

[0268] 《Display of Characters》 Regarding the display of characters, various modes are possible. In the display screens (4) and (5) of FIG. 26 described above, the control unit 18 moves the character 100 left and right so that the character 100 displayed on the screen does not overlap with the current position and the searched route 53. For example, the display of the character 100 and the searched route 53 may be closer to the opposite sides on the display screen. By doing so, it is possible to make the user feel that the character 100 is being careful not to interfere with the route search service, so that the user can feel more favorably disposed.

[0269] The movement of the character 100 may be such that it disappears once and then appears at the destination, or it may move while the display is continued. If the movement is frequent, the user may feel annoyed, but for the user who can appeal the existence of the character 100 and who wants to continue watching the character 100, the operation of the moving character 100 can be enjoyed.

[0270] The control unit 18 may cause the display unit 5 to display whether or not voice recognition is activated and being received. By doing so, the user can visually recognize whether or not voice operation is possible. As for whether or not voice recognition is being received, the color of the display screen is changed for display. By doing so, for example, the state of voice recognition can be immediately perceived through the user's vision. In particular, when voice recognition is being received, the normal color is used, and when it is not being received, the grayscale is used. By doing so, it is possible to make the user intuitively understand whether or not voice recognition is being received.

[0271] The control unit 18 may display whether or not voice recognition is being received on the character. By doing so, it is possible to know whether or not voice operation is possible just by looking at the character, making it easier for the user to notice.

[0272] As a display mode for the character, for example, it is set to a mode in which the state of the character clearly changes. By doing so, for example, the state of voice recognition can be immediately perceived. As a mode in which the state of the character clearly changes, for example, an object is added to the character, the color or form of the character is changed, a specific pose or gesture is made, or a line is displayed in a speech bubble. By doing so, for example, a user who knows the normal state of the character can immediately notice the change by comparing it with the normal state.

[0273] As a mode of adding an object to the character, for example, equipment related to sound such as a headset, headphones, earphones, or microphone is worn on the character. By doing so, the user is made to associate voice recognition, which is a function related to sound. As a mode of changing the color or form of the character, for example, the hairstyle and its color are changed, or the character is changed into another costume. By doing so, the difference from the normal state becomes easy to understand. In particular, if the ears of the character are made larger, it becomes a gesture indicating that the character is trying to hear the voice well. Furthermore, as a pose or gesture, for example, a gesture of putting a hand behind the ear also becomes a gesture indicating that the character is trying to hear the voice well.

[0274] 《Use of Nickname》 The control unit 18 may insert the voice of the user's call name registered in the database 19 in advance into the response phrase, which is the voice output by the character. That is, the control unit 18 makes the character say the user's name in a nickname style during the conversation with "Kirishima Rei". For example, the character can be made to call it like "Ke-kun" or "Yu-kun".

[0275] For example, as an activation word for activating voice recognition and surrounding search, "Rei, vicinity" is registered. The following is an example of a conversation when a user whose nickname is registered as "Yu-kun" makes "Kirishima Rei" perform a surrounding search on the map display screen. What is inside the " " is the uttered phrase.

[0276] User: "Ray, around here" -> Voice recognition activated -> Transition to the surrounding search genre screen Ray: "What~ User, I'm busy though." User: "Internal medicine" -> Transition to the map screen, display the surrounding internal medicine icons Ray: "What's wrong? Are you okay? Which doctor are you going to?" User: "2" (the number two) Ray: "Got it, number two." -> Display the information of the second internal medicine hospital on the map User: "Search" -> Route search and display on the map Ray: "What~? Going here?" User: "Please." Ray: "Understood, leave it to me." -> Guidance starts

[0277] In this way, the user can feel a greater sense of familiarity with the character by being called by their own nickname within the conversation. Also, since the voice of the nickname is inserted into the voice output at the start of the conversation, the user will be called by their own name, which can attract the user's attention well. Furthermore, when the voice containing the character's name is recognized, the voice of the user's name is inserted into the output voice, so they will call each other by name, and the user can obtain a more intimate feeling.

[0278] Note that as the conversation progresses, the frequency of inserting the user's name into the voice output by the character may be reduced. In this way, for example, at first the name was called, but later the frequency of being called decreases, so it is possible to create a conversation between two people who are getting used to each other. Furthermore, voice data corresponding to the 50 - sound chart is registered in advance in the database 19, and the voice of the name selected by the user from this 50 - sound chart may be inserted during the conversation. In this way, for example, compared to the case of registering multiple names in advance, the memory capacity can be saved, which is good.

[0279] "Submission" When the user discovers a traffic control or inspection location, the control unit 18 can execute the submission function by transmitting the location information regarding the traffic control location to the information collection server in response to the input of the user's voice corresponding to the recognition word registered in the database 19 in advance. In addition, the control unit 18 can receive the distribution from the information collection server regarding the location information of the submitted traffic control location and display it together with the map on the display unit 5. Therefore, the navigation device is connected to the control unit 18 and includes a communication unit that transmits and receives information to and from the outside.

[0280] This communication unit is connected wirelessly or by wire to a communication device inside the vehicle having an Internet connection function. For example, the communication unit can be realized by connecting to a mobile phone such as a smartphone having a tethering function or a wireless LAN router through Wi-Fi (registered trademark) connection, Bluetooth (registered trademark) connection, or USB connection. Such a communication device may be incorporated in the navigation device.

[0281] First, the process of submission by touch operation is as follows. 1. The user discovers a traffic control or inspection location while driving 2. Touch the button that instructs the setting of the submission location on the display screen 3. Set the current position on the map when the button is touched as the submission location and display an icon at the corresponding position 4. Touch the submission button on the standby screen while parked or stopped 5. A list of multiple submission locations set in steps 1 to 3 is displayed 6. Touch any of the locations to be submitted 7. Touch the submission button for final confirmation 8. Transmit the information of the selected submission location to the server that collects the submission locations 9. The server that collects the submission locations registers the received information as a traffic control location 10. The server that collects the submission locations distributes the registered information 11. Receive and display the distributed information

[0282] Note that between 6 and 7, it is also possible to select whether the target is Orbith, the N system, inspection, interrogation, etc., select the target direction such as the driving lane, oncoming lane, right direction, left direction, etc., select the implementation period, etc.

[0283] In the case of posting by such button operations, during driving, since a plurality of hierarchical procedures are required after discovering the posting location, a plurality of button operations are necessary.

[0284] However, in this aspect, posting can be performed in response to the user's voice input. That is, when the user utters a recognition word registered in the database 19 in advance, the control unit 18 recognizes this and causes the communication unit to transmit the current position at the time of the utterance to the server as the posting location via the communication device. The server registers the received information as the enforcement location and distributes the registered information. The communication unit receives the distributed information via the communication device, and the control unit 18 causes the enforcement location to be displayed on the map based on the received information. For this reason, it is possible to perform it simply and quickly without the trouble of button operations.

[0285] As the enforcement location, as described above, it includes Orbith and the N system, but in particular, it is preferably a check location where the implementation position and time change, such as an enforcement and interrogation area. In this way, it is possible to collect information about enforcement locations where it is difficult to register the position in advance through posting.

[0286] As the position information, for example, it may be the current position or the central position coordinates of the map. In this way, the enforcement location can be specified by information that can be obtained reliably and simply, and since voice recognition is used, the deviation from the actual enforcement position can be small.

[0287] As recognition words, it is preferable that a phrase for activating voice recognition and a phrase for starting to accept a post are registered. By doing so, the user can easily execute multi-step operations such as voice recognition and accepting a post by uttering the phrase for activating voice recognition and the phrase for starting to accept a post.

[0288] As the phrase for activating voice recognition and the phrase for starting to accept a post, for example, they may be registered as separately divided phrases. By doing so, for example, it is possible to prevent the acceptance of a post from starting simultaneously with the start of voice recognition and causing misrecognition. Also, as the phrase for activating voice recognition and the phrase for starting to accept a post, for example, they may be set as a series of phrases. By doing so, for example, the acceptance of a post can be started all at once along with the start of voice recognition, so that it is possible to immediately enter a state where a post can be made.

[0289] As recognition words, if a phrase for activating voice recognition and a phrase for making a post are registered, the user can easily perform multi-step operations such as voice recognition and making a post by uttering the phrase for activating voice recognition and the phrase for making a post.

[0290] As the phrase for activating voice recognition and the phrase for making a post, for example, they may be registered as separately divided phrases. By doing so, for example, it is possible to prevent a post from being made simultaneously with the start of voice recognition and being accidentally posted. Also, as the phrase for starting voice recognition and the phrase for making a post, for example, they may be set as a series of phrases. By doing so, for example, a post can be made all at once along with the start of voice recognition, so that it is possible to immediately make a post and minimize the deviation between the place of management and the place of posting as much as possible.

[0291] Note that the recognition words for posting may be shorter than other recognition words. By doing so, for example, for a highly urgent post, it is possible to give an instruction in a short time.

[0292] Also, the recognition word for posting may be more redundant than other recognition words. By doing so, for example, when using normal speech recognition, the possibility of uttering other recognition phrases is reduced, and incorrect posting can be prevented. Being redundant means, for example, having a large number of characters or being divided into multiple parts. By doing so, for example, the possibility that a user who does not intend to post will accidentally speak can be reduced.

[0293] Also, the recognition word for posting may be a phrase that the user rarely utters in daily life. By doing so, for example, it is possible to prevent frequent misrecognition from occurring due to phrases that the user frequently uses in the car, and prevent posting that the user does not intend.

[0294] When the recognition word for posting is recognized, the control unit 18 may, in response thereto, output the voice of a character registered in advance, thereby realizing an interaction function as in the above embodiment. In this case, the voice of the character in the interaction for posting may be shorter than the voice in other interactions. By doing so, for example, the interaction time can be shortened and posting can be made quickly.

[0295] 《Variations of Response Phrases》 The response phrases of the characters pre-registered in the database 19 can be in various forms as follows. When outputting the voice of the character prompting the selection of the searched facility in the content matching the displayed icon, it may be the voice asking which number is good if it is a number, the voice asking which letter is good if it is a letter, the voice asking what the shape is if it is a shape, and the voice asking what the color is if it is a color. By doing so, for example, the user can easily know what to say to select the facility.

[0296] Also, for example, as the content corresponding to the displayed icon, it may be voice indicating the range of candidates to be selected. By doing so, for example, if voice is output in the form of "from number X to number Y", the user can grasp not only the content to be spoken but also the range to be selected.

[0297] As the voice of a pre-registered character, voice content indicating the situation of the character may be inserted. By doing so, the user can hear the situation of the character. For example, if the character is in a negative situation, it can arouse the user's concern for the character, and if the character is in a positive situation, it can arouse the user's feeling of empathy for the character.

[0298] As a statement indicating the situation of the character, for example, a statement such as "I'm busy" that allows the user to know whether the character is in a positive or negative state may be used. By doing so, the system does not unconditionally follow the user's instructions, but can arouse the user's feelings of concern and empathy, so that a feeling similar to a conversation with an actual person can be obtained.

[0299] A plurality of lists may be set, and there may be a list in which phrases different from other phrases are registered even within the same list in the response phrase. By doing so, for example, as a response to a user's utterance of similar content, different phrases can be output to give the user an unexpected impression, change the impression of the character, and prevent boredom. As different phrases, for example, as shown in the Ray response phrase 3 in FIG. 9(a), among the positive inquiry phrases such as "Give me some instructions.", "What should I do?", "What to do?", "How to change it?", an ambiguous phrase such as "Do as you like." may be inserted. Also, as shown in the Ray response phrase 24 in FIG. 9(b), among the phrases indicating retry such as "Re-search", "Search again", "Once more", "I'll do it again", a phrase that reads the user's mind such as "You've been trying me out in various ways, right?" may be inserted.

[0300] Even when voice recognition is incorrect, the output phrase may be registered. By doing so, for example, even when the recognition is incorrect, it is possible to make the user recognize that the conversation with the character is continuing rather than returning some phrase.

[0301] It is preferable to register a phrase that serves as an escape route in case of an error in voice recognition. By doing so, for example, although there is a possibility of an error in voice recognition, even if an error occurs, the escape route of the character can soothe the user's irritation and anger.

[0302] As an escape route, for example, like "I'll do my best with voice recognition. It may not be 100% recognizable" shown in the Ray response phrase 1 in FIG. 8(a), it is preferable to use an expression that shows the will to try to recognize accurately along with an expression that shows that complete achievement is difficult. By doing so, for example, it is possible to make the user have a forgiving mentality towards misrecognition rather than simply indicating that recognition is impossible.

[0303] When there is an input of the user's voice indicating that the content the user expects as a response is different from the content of the character's voice output, it is preferable that a response phrase from the character to this is registered. By doing so, for example, if the user points out a mistake of the character, there is a response from the character to this, so it is good that a natural conversation between people in case of misunderstanding can be enjoyed.

[0304] It is preferable that a phrase for advertising is registered. By doing so, for example, it is good that the user can be induced to a partnering facility. For example, it is preferable to use a phrase that invites to a specific facility, such as "How about ○○ shop?".

[0305] It is preferable that a phrase about the reason for doing something at the facility the user is going to is registered. By doing so, by talking about the reason for going to a certain facility, it becomes easier for the user going to that facility to listen. For example, it is possible to attract the user's attention whether the reason for going to the destination is unfavorable or favorable for the user.

[0306] As a facility where the reason for going to the destination is not favorable for the user, for example, it is preferable to use a facility related to the user's illness, such as "hospital", "dentist", "pharmacy". As a facility where the reason for going to the destination is favorable for the user, for example, it is preferable to use a facility with a slightly higher hurdle than the facilities frequently visited in daily life, such as "department store", "sushi", "eel".

[0307] As phrases for the character to evoke various emotions in the user, for example, as shown in response phrases 6-8 in FIG. 9(c), phrases that make the user feel attachment to the character by showing consideration for the user's state, such as "Hospital? Is something wrong? Are you okay?", "Don't overdo it." It is preferable to use. Also, for example, as shown in response phrase 6-29 in FIG. 9(d), it is preferable to use a phrase that causes the user to have a feeling of regret, such as "Are you brushing your teeth properly?".

[0308] It is preferable that phrases that give positive or negative advice on what to do at the facility the user is going to are registered. By doing so, for example, based on the advice of the character, the user can determine what to do or not to do at the destination. As positive advice, for example, as shown in response phrase 6-29 in FIG. 9(d), it is preferable to use a phrase that recommends something, such as "You should get your teeth cleaned once in a while." As negative advice, for example, as shown in response phrase 6-23 in FIG. 9(e), it is preferable to use a phrase that prohibits something, such as "Don't make a painful face. It's manly to endure with a cool face."

[0309] It is preferable that phrases that are one-sided impressions of the character on what to do at the facility the user is going to are registered. By doing so, for example, since the impression of the character can be obtained, the user can use it as a reference for actions at the destination.

[0310] As one-sided impressions of the character, for example, as shown in response phrase 6-14 in FIG. 9(f), it is preferable to have positive impressions, such as "Black pork char siu that melts in your mouth. Delicious~." By doing so, the user will want to take similar actions at the facility. Also, for example, it is preferable to have negative impressions, such as "Going to a restaurant alone? Aren't you lonely? ··· Just leave me alone?" By doing so, it can make the user want to change the destination. Furthermore, for example, as shown in response phrase 6-21 in FIG. 9(g), it is also acceptable to have a mixture of positive and negative impressions, such as "The shopping mall seems fun just to walk around. But it also seems tiring." By doing so, the user can be made aware of the situation at the destination in advance.

[0311] Phrases that make the character seem to exist in reality may be registered. By doing so, for example, the user can obtain a feeling that the character actually exists.

[0312] As phrases that give the impression of actually existing, for example, as shown in response phrase 6-24 of FIG. 9(h), "I'm so happy that you're taking me on a business trip too~.", or as shown in response phrase 6-25 of FIG. 9(i), "Where are we going to stay?" would be good phrases that give the impression of being together. By doing so, for example, the user can get the feeling of always being with the character.

[0313] When voice other than the phrases set in advance as recognized phrases is input, it is preferable that the voice data output by the character is registered. By doing so, for example, even if it is other than the recognized phrases, there is always a reaction from the character, so the user is not reminded that it is an interaction with the system due to the conversation being interrupted or being notified as an error, and the feeling of continuing the conversation with the character can be maintained. As a phrase when voice other than the set phrases is input, a phrase that cannot be recognized as shown in the ray response phrase 50 of FIG. 9(j) can be considered. Note that the response phrases are not limited to being output randomly from a plurality of them. If the order of output from a plurality of them is determined and they are output sequentially in that order, and when the response phrase of the last rank is output, it returns to the response phrase of the first rank and is output, it is possible to prevent the same response phrase from being output repeatedly. Furthermore, for each recognized phrase, the corresponding response phrase may be determined one-to-one. By doing so, the memory amount of information can be saved. Even in such a case, if the response phrases with the content as shown above are set, the same effect can be obtained.

[0314] 《Timing of Voice Output and Voice Recognition》 The timing of voice output and voice recognition can be set in various ways as follows. While the voice of the character is being output, the control unit 18 may stop voice recognition until the voice output is completed. By doing so, for example, when there is an input of the user's voice during the character's voice, it is possible to prevent other processes from being executed by voice recognition. As other processes, for example, processes with low urgency such as volume, brightness, scale, and peripheral search may be used. By doing so, it is not necessary for processes that do not necessarily need to be executed immediately to be executed by voice recognition.

[0315] The control unit 18 may also accept voice recognition while the voice of the character is being output. By doing so, for example, when there is an input of the user's voice during the character's voice, other processes may not be interrupted by voice recognition. As other processes, for example, processes with high urgency such as the process of transmitting position information regarding the place of inspection to the information collection server may be used. By doing so, it is not necessary for processes that the user desires to execute immediately to be interrupted.

[0316] When there is an input of the user's voice while the voice of the character is being output, the control unit 18 may stop the output of the voice of the character. By doing so, for example, it is possible to give an impression as if the character's speech is blocked by the user's voice. The output of the stopped voice of the character does not need to give an unnatural impression by repeating the same utterance from the beginning, for example, without outputting it again. Also, if the output of the stopped voice of the character is started from, for example, the point where it was stopped, it is possible to give an impression that the character was waiting during the user's speech.

[0317] While the voice of the character is being output, the control unit 18 may stop the voice of the character when a predetermined input operation is performed. In this way, for example, when it is not desired for a third party to hear the conversation with the character, the voice of the character can be stopped by a predetermined input operation. As the predetermined input operation, for example, an operation of a physical key may be used. In this way, since it is only necessary to operate a location that always exists fixedly, an immediate operation can be easily performed.

[0318] 《Timing of Starting Voice Recognition》 The timing of starting voice recognition can be variously configured as follows. The control unit 18 may start a predetermined process for outputting information that the user wants to know in response to the conversation, based on an input of voice from the user corresponding to a predetermined recognition word. In this way, for example, since the predetermined process starts based on an input of voice from the user, it is not necessary for the system to start and execute the process on its own.

[0319] The control unit 18 may start a predetermined process for outputting information that the user wants to know in response to the conversation, when there is an input of voice from the user corresponding to a predetermined recognition phrase in response to the output of the voice of a predetermined character. In this way, for example, since the output of the voice of the character serves as an opportunity for the user to determine whether to start a predetermined process, the opportunity to use the system can be increased.

[0320] The output of the voice of the character may be, for example, to inquire about the presence or absence of the start of a specific process. In this way, for example, the user can be prompted to use the process. If the timing of the output of the voice of the character is, for example, periodic, the frequency of using the process can be increased. Also, if the timing is when a predetermined condition is satisfied, the user can be prompted to use it at a timing suitable for using the process.

[0321] 《Others》 As a mode of relatively emphasizing the icons indicating surrounding facilities compared to the map, it is preferable to make the display such that the positions of the facilities are clear. For example, by changing at least one of the brightness, saturation, and hue of the icons and the map, the contrast between the two can be made prominent.

[0322] When making a display that can distinguish the ranks according to a predetermined standard for the icons indicating surrounding facilities, the predetermined standard may be, for example, the order of high importance to the user. In this way, the facilities that the user most wants to know can be preferentially displayed. The importance to the user may be, for example, having been there in the past, having a high frequency of going there in the past, or matching the user's preferences, etc. In this way, the user can quickly know the facilities they want to go to. Also, for example, those with low importance to the user, such as having been there in the past but not wanting to go there again, may be excluded. In this way, for example, only those that the user may go to can be displayed.

[0323] As a display that can distinguish the ranks, for example, it may be information that visually shows the ranks, such as changing step by step according to shapes such as ◎, ○, △, ×, etc., brightness, saturation, and hue. In this way, for example, the ranks can be intuitively grasped visually.

[0324] When recognizing the voice of a user who selects one of the facilities according to the content corresponding to the displayed icon, for example, if it is an alphabet, it may be the alphabet of the icon of the desired facility, if it is a shape, it may be the shape of the icon of the desired facility, and if it is a color, it may be the color of the icon of the desired facility. Even in this way, the facility can be selected with words that are easier to remember than the facility name.

[0325] In the display area of the button-displayed item, the part where the background image can be seen may be, for example, a transparent part. In this way, for example, while ensuring the visibility of the item display, the background image can also be seen better.

[0326] The recognition words registered in advance for recognizing the user's voice may include those displayed on the display screen and those not displayed on the display screen. By doing so, for example, even if the user utters a phrase not displayed on the display screen, if it is a recognition phrase set in advance, some reaction from the character can be obtained, which can be a surprise for the user and give the user the joy of discovery.

[0327] As for the indicator corresponding to the function that changes quantitatively as a recognition word and the number indicating the quantitative level, for example, the name and the quantitative level may be combined into one recognition phrase like "onryosan", "kidoni", "sukeru nikiro". In this way, the utterance can be completed in one time.

[0328] The button on which the indicator corresponding to the function is displayed may be a button that executes the function without transitioning to another display screen. By doing so, for example, the function can be executed through dialogue without transitioning to another display screen to execute the function, so the current screen display does not need to be disturbed.

[0329] In this embodiment, an example of a navigation device has been described, but it can be implemented as a function of various electronic devices. For example, it may be incorporated as a function of a radar detector, a drive recorder, or a car audio. Also, the value of the quantitative level, the screen size of the display unit 5, various time settings, etc. described in this embodiment can be arbitrary within the range where the effects of the present invention are achieved. Further, the control unit 18 may be provided with a function of setting the priority order of each function and alarm based on an instruction from the user from the remote control 17 or the like, and the control unit 18 may be configured to perform processing according to the set priority order.

[0330] Furthermore, in the above-described embodiments, the apparatus includes the database 19 that stores various types of information, and the control unit 18 accesses the database 19 to read necessary information and performs various processes. However, the present invention is not limited to this. For example, part or all of the information registered in the database 19 may be registered in a server. Then, the navigation device, the radar detector, and other electronic devices may be configured to have a function of communicating with the server, and the control unit 18 may access the server as appropriate to obtain necessary information and execute processes. Furthermore, at least part of the functions of the control unit 18 may be placed in the server, and the functions may be executed by the server, and the electronic device held by the user may be configured to obtain the execution results.

[0331] The functions as the control system in the above-described embodiments are configured as a program for causing a computer included in the control unit 18 to realize them. However, the program is not limited to this and may be distributed and arranged in a plurality of computers for distributed processing.

Explanation of Reference Numerals

[0332] 2 Apparatus main body 3 Cradle 4 Case main body 5 Display unit 6 Cradle main body 7 Pedestal part 8 Touch panel 10 Alarm lamp 11 Microwave receiver 12 GPS receiver 13 Wireless receiver 18 Control unit 19 Database 20 Speaker 21 SD memory card slot 22 SD memory card 51 Current position 52 Destination 53 Route

Claims

1. A system that executes a function required by a vehicle passenger based on the result of recognizing the voice of a user who is a passenger of the vehicle, comprising, as a recognition phrase registered in advance for recognizing the voice of the user, a combination of a name corresponding to a function that changes quantitatively and a number indicating a quantitative level in one recognition phrase. The system is characterized by this.

2. The display screen of the display means has a function of displaying a button for selecting a function, on which the value of the current quantitative level of the function is displayed. The system according to claim 1, characterized by this.

3. While the voice of the character is being output, it has a function of stopping the voice recognition until the voice output is finished, and when there is an input of the user's voice during the voice of the character, it has a function of preventing the execution of the process of changing at least any one of the volume, brightness, and scale by voice recognition. The system according to claim 1 or 2, characterized by this.

4. While the voice of the character is being output, it has a function of stopping the voice of the character when a physical key is operated. The system according to any one of claims 1 to 3, characterized by this.

5. A program for causing a computer to realize the functions of the system according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Unit and method for control by voice recognition and record medium where program for control by voice recognition is recorded

    JP1999242497A

  • Continuous picture display device

    JP1999344994A

  • Device and method for control and storage medium storing program for executing operation processing therefor

    JP2000099306A

  • Speech recognition device

    JP2000163091A

  • Input device and program

    JP2003122393A