Voice response system

The voice response system addresses the limitation of single-answer systems by allowing multiple tone responses and integrating with external devices for personalized interaction, enhancing user experience.

JP2026086864APending Publication Date: 2026-05-26CASE CHARTER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CASE CHARTER CO LTD
Filing Date
2026-02-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing voice response systems provide only a single answer to a question, lacking user-friendliness and flexibility in response generation.

Method used

A voice response system comprising a requesting device, providing device, and server that allows setting permissions for information provision, generates multiple responses in different tones, and integrates with external devices for diverse response generation, including personality and preference analysis.

Benefits of technology

Enhances user-friendliness by providing varied responses in different tones, reducing processing load on the device, and enabling personalized and flexible interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086864000001_ABST
    Figure 2026086864000001_ABST
Patent Text Reader

Abstract

In voice response systems, we provide a voice response system that is more user-friendly for the user. [Solution] A voice response device that provides voice responses to input character information comprises a response acquisition means for acquiring multiple different responses to the character information, and a voice output means for outputting each of the multiple different responses in a different tone of voice. With such a voice response device, since multiple responses can be output in different tones of voice, even when a single solution to a character piece of information cannot be uniquely identified, different solutions can be output in different tones of voice in an easily understandable way to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related applications

[0001] This international application claims priority based on Japanese Patent Application No. 2012 - 137065, Japanese Patent Application No. 2012 - 137066, and Japanese Patent Application No. 2012 - 137067, which were filed with the Japan Patent Office on June 18, 2012, and incorporates by reference the entire contents of Japanese Patent Application No. 2012 - 137065, Japanese Patent Application No. 2012 - 137066, and Japanese Patent Application No. 2012 - 137067 into this international application.

Technical Field

[0002] The present invention relates to a voice response system that enables responses to be made by voice.

Background Art

[0003] As the above - mentioned voice response device, there is known one that searches a dictionary for an answer to an input question and outputs the searched answer by voice (see, for example, Patent Document 1). Also, there is known a technique for generating an answer to a question based on the content of the dialogue with the user (see, for example, Patent Document 2).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] In the above - mentioned technology, it is simply set to give one answer specified by a dictionary to one question. In a voice response system, making it more user - friendly for the user is one aspect of the present invention.

Means for Solving the Problems

[0006] In the first phase of the invention, A voice response system comprising a requesting device that initiates a request, a providing device that provides information, and a server capable of communicating with the requesting device and the providing device, The aforementioned provider device is configured to allow the setting of whether or not to permit the provision of information. The server is configured to receive a request from the requesting device, and if the request includes a request specifying the providing device, to send the request to the specified providing device, and if the provision of information is permitted, to receive the provided information from the providing device. A providing unit is configured to generate an audio response based on the provided information as a response to a request from the requesting device, and to provide the response to the requesting device. It is equipped with. In another aspect of the invention, A voice response device that provides voice responses to input character information, A response acquisition means for acquiring multiple different responses to the aforementioned character information, A voice output means that outputs the aforementioned multiple different responses, each in a different tone of voice, It is characterized by having the following features.

[0007] Such a voice response device can output multiple responses in different voice tones. Therefore, even when a single character of information cannot be uniquely identified, different solutions can be output to the user in a clear and understandable way using different voice tones. This makes it more user-friendly.

[0008] The voice response device of the present invention may be configured, for example, as a terminal device owned by the user, or as a server that communicates with this terminal device. Furthermore, character information may be input using an input means such as a keyboard, or it may be input by converting speech into character information.

[0009] By the way, in the above-mentioned voice response device, as in the invention of the second phase, A voice input method for the user to input voice, and a system that converts the input voice into text information. An external device that generates multiple different responses to the character information and transmits them to the voice response device, and a voice transmission means that transmits to the external device, Equipped with, The response acquisition means acquires the response from the external device. You may do so.

[0010] With such a voice response device, since it can receive voice input, it can be configured to input text information via voice. Furthermore, since the response can be generated by an external device, the processing load on the voice response device can be reduced.

[0011] Furthermore, in the voice transmission means, the operation of "converting the input voice into text information" may be performed by the voice response device or by an external device. Furthermore, in the above-mentioned voice response device, as in the invention of the third phase, The voice response device or the external device includes a response recording means that records a plurality of different responses for each of a plurality of character information, including a positive response and a negative response for each character information. The response acquisition means acquires the positive response and the negative response as a plurality of different responses, The voice output means may be configured to reproduce the positive response and the negative response in different tones.

[0012] Such a voice response device can reproduce responses from different perspectives, such as positive and negative responses, using different tones of voice, making it possible to reproduce the sound as if different people were speaking. Therefore, it is less likely that the user listening to the voice will feel any sense of unease.

[0013] Note that the tone of voice may be changed according to the type of response and the choice of words in the response. For example, when responding in a gentle tone, it may be played back with a calm female voice, and when responding in an intense tone, it may be responded with a courageous male voice, etc. That is, the response content and character may be associated with each other, and the tone of voice may be set according to the character.

[0014] Also, in the above voice response device, as in the invention of the fourth aspect, it can be configured to be used at a workplace or a company reception, or to be configured to convey something that the user is reluctant to say directly to someone else.

[0015] When using the voice response device at the reception, the name and company name of the person coming for sales are recorded in advance in the voice response device or an external device, and when the person coming to the reception announces this name or company name, a response may be generated so that the voice of a rejection phrase is played back. That's all.

[0016] Also, when configured to convey something that is difficult to say instead, for example, before a date, if you talk to this device saying that you want to say such and such today, at an appropriate timing (for example, a preset time, or when a certain time has elapsed after the conversation has stopped), the voice response device may talk (play back the voice) on your behalf.

[0017] Alternatively, it may be configured to say words that trigger something difficult to say, for example, words like "Come to think of it, didn't I say I had something to tell her?" That is, instead of immediately outputting a response, the response may be output when a playback condition is satisfied, such as after a certain time has elapsed.

[0018] Furthermore, in the above voice response device, as in the invention of the fifth aspect, the external device or the voice response device may obtain information for generating a response to the character information from another voice response device. Also, in the above voice response device, as in the invention of the sixth aspect, when information for generating a response to the character information is requested from another voice response device, the device may return information corresponding to this request.

[0019] In this case, the voice response device may be equipped with sensors for detecting position information, temperature, humidity, illuminance, noise level, etc., and databases such as dictionary information, and extract necessary information according to the request.

[0020] According to such a voice response device (external device), information for generating a response can be obtained from another voice response device. In this case, information unique to another voice response device, such as the position of another voice response device, can be obtained.

[0021] Also, its own unique information can be transmitted to another voice response device. Furthermore, in the above voice response device, as in the invention of the seventh aspect, a response (for example, an affirmative response or a negative response) output by itself or another voice response device may be input as character information, and a response for making a counterargument to this response may be generated. That is, from the user's perspective, discussions based on opinions from both the supportive and opposing positions can be heard. And after hearing this discussion, the user can make a final judgment.

[0022] This configuration can be realized using one or more voice response devices. In this case, for multiple voice response devices to exchange voices, the voices may be directly input and output, or communication using wireless or the like may be utilized.

[0023] Also, in the invention of the eighth aspect, a voice response device that makes a voice response to the input character information, A means for acquiring personality information that acquires personality information that associates the characteristics of a person related to the user or a person related to the user with a predetermined classification, A response acquisition means for acquiring response candidates that represent multiple different responses to the aforementioned character information, A voice output means that selects a response to be output from a list of response candidates according to the aforementioned personality information and outputs the selected response, It is characterized by having the following features.

[0024] Such voice response devices can provide different responses depending on the personality of the user or those related to the user (stakeholders). Therefore, they can be made more user-friendly.

[0025] Furthermore, in the above-mentioned voice response device, as in the invention of the ninth phase, The system includes a first personality information generation means that generates personality information of the user or related party based on answers to a set number of pre-set questions, The personality information acquisition means acquires the personality information generated by the personality information generation means. You may do so.

[0026] Such a voice response device can generate personality information. When generating this personality information, well-known personality analysis techniques (such as the Rorschach test or the Szondi test) can be used. Alternatively, aptitude test techniques used by companies for recruitment examinations may also be used to generate this personality information.

[0027] Furthermore, in the above-mentioned voice response device, as in the invention of the 10th phase, The system includes a second personality information generation means that generates personality information of the user or related party based on the string contained in the input character information, The personality information acquisition means acquires the personality information generated by the personality information generation means. You may do so.

[0028] Such a voice response device can generate personality information during the user's experience using the device. Furthermore, in the above-mentioned voice response device, as in the invention of the 11th phase, The system includes a preference information generation means that generates preference information indicating the preference tendencies of the user or related party based on a string of characters contained in the text information, The voice output means selects a response to be output from the response candidates based on the preference information and outputs the selected response. You may do so.

[0029] Such voice response devices can respond according to the preferences of the user or related parties. Furthermore, in the above-mentioned voice response device, as in the invention of the 12th phase, the user's actions (conversation, places visited, things captured by the camera) may be learned (recorded and analyzed) and used to supplement any omissions in the user's conversation.

[0030] For example, in a conversation where the user responds to the question, "Is hamburger okay today?" with "I'd prefer curry," if this device adds, "We had hamburger yesterday," the reason why the user said they preferred curry becomes clear.

[0031] Furthermore, such a configuration can be implemented during a phone call, and it may also be configured to spontaneously join the user's conversation. Furthermore, in the above-mentioned voice response device, as in the invention of the 13th phase, A means for obtaining response candidates from a predetermined server or the internet. It may also be equipped with.

[0032] Such a voice response device can obtain response candidates not only from its own device and external devices, but also from any device connected via the internet, dedicated lines, etc. Furthermore, in the above-mentioned voice response device, as in the invention of the 14th phase, A character information generation means that converts user actions into character information. It may also be equipped with.

[0033] In this invention, the actions referred to include those resulting from muscular movements such as conversation, handwriting, or gestures (e.g., sign language). Such a voice response device can convert the user's actions into textual information.

[0034] Furthermore, in the above-mentioned voice response device, as in the invention of the 15th phase, The text information generation means converts the user's speech into text information and stores speech habits (such as pronunciation quirks) as learning information (capturing and recording these characteristics). You may do so.

[0035] Such a voice response device can generate text information based on learned data, thereby improving the accuracy of text information generation. Furthermore, in the above-mentioned voice response device, as in the invention of the 16th phase, Transfer means for transferring the learning information to another voice response device, It may also be equipped with.

[0036] With this type of voice response device, even when the user uses another voice response device, the learned information recorded by this device can be utilized. Therefore, the accuracy of text information generation can be improved even when using other voice response devices.

[0037] Furthermore, the above-mentioned voice response device may also detect either the user's actions or operations and generate learning information or personality information based on these, as in the invention of the 17th phase.

[0038] Such a voice response device could, for example, detect that the user has been jumping onto trains for several consecutive days, prompting them to leave home a few minutes earlier the following day. Or, if it detects from conversation that the user has a tendency to get angry easily, it could output calming voices or music.

[0039] Furthermore, in the above-mentioned voice response device, as in the invention of the 18th phase, It may also include means for acquiring information from other voice response devices to acquire information recorded in other voice response devices.

[0040] Such a voice response device can generate a response based on information recorded in other voice response devices. Furthermore, in the above-mentioned voice response device, as in the invention of the 19th phase, When the aforementioned character information is not input, a playback condition determination means determines whether the status of the voice response device matches the playback conditions set in advance as conditions for outputting sound, A message playback means that outputs a pre-set message when the aforementioned playback conditions are met, It may also be equipped with.

[0041] Such voice response devices can output voice even when no text information is entered (i.e., when the user does not speak). For example, by forcing the user to speak, it can be used as a measure to suppress drowsiness while driving. Also, by determining whether or not a person living alone responds, it can be used to check on their well-being.

[0042] Furthermore, in the above-mentioned voice response device, as in the invention of the 20th phase, The message playback mechanism acquires news information and outputs a message related to that news in the form of a question that prompts the user for a response. You may do so.

[0043] With such a voice response device, it is possible to have a conversation about the news, This helps prevent the conversation from becoming repetitive. For example, if you obtain information about a company's stock price, you could say, "Did you know that company XX's stock price went up by XX yen today?"

[0044] Furthermore, in the above-mentioned voice response device, as in the invention of the 21st phase, The audio output means or message playback means outputs a pre-set message with separately acquired external information (such as news or environmental information like temperature, weather, and location information) added to it. You may do so.

[0045] Such a voice response device can output a response that combines a predetermined message with acquired information. Furthermore, in the above-mentioned voice response device, as in the invention of the 22nd phase, The system retrieves multiple messages and selects which messages to play based on their playback frequency, then outputs them. You may do so.

[0046] Such voice response devices can introduce randomness into message playback by making it less likely to play frequently played messages, or they can encourage attention and memory retention by deliberately repeating frequently played messages.

[0047] Furthermore, in the above-mentioned voice response device, as in the invention of the 23rd phase, A means of sending a message to a pre-configured contact that identifies the user and indicates that a response was not received, in case no reply or answer is received. It may also be equipped with.

[0048] Such voice response devices can notify a contact person if no response is received. Therefore, for example, abnormalities in elderly people living alone can be reported early. Furthermore, in the above-mentioned voice response device, as in the invention of the 24th phase, The message playback mechanism memorizes the conversation content and asks questions to obtain the same information that was heard (memory confirmation process). You may do so.

[0049] Such voice response devices can be used to check the user's memory and to help solidify that memory. Furthermore, in the above-mentioned voice response device, as in the invention of the 25th phase, A speech accuracy detection means for detecting the degree of accuracy of pronunciation and accent of the voice input by the user, Accuracy output means that outputs the detected accuracy, It may also be equipped with.

[0050] Such voice response devices allow you to check the accuracy of your pronunciation and accent. They are particularly useful when practicing a foreign language. Furthermore, in the above-mentioned voice response device, as in the invention of the 26th phase, The accuracy output means outputs the audio containing the closest word when the accuracy is below a certain value. You may do so.

[0051] Such voice response devices allow users to verify the accuracy of their pronunciation and accent. Furthermore, in the above-mentioned voice response device, as in the invention of the 27th phase, The message playback mechanism may be configured to output the same question again if the accuracy is below a certain value.

[0052] Such voice response devices can obtain accurate answers by outputting the same question repeatedly. Furthermore, in the above-mentioned voice response device, as in the invention of the 28th phase, Connection control means that identifies the communication partner based on the input character information and connects the communication partner to a pre-configured communication destination for each communication partner. It may also be equipped with.

[0053] Such voice response devices can assist with reception duties and telephone answering. In particular, in the above-mentioned voice response device, as in the invention of the 29th phase, The connection control system identifies sales activities and visitors, and plays a message to decline if it is a sales activity. You may do so.

[0054] Such voice response devices allow users to exclude individuals who may disrupt their work without having to deal with them themselves. Furthermore, in the above-described voice response device, as in the invention of the 30th phase, keywords contained in the input text information (especially voice) may be extracted and connected to the destination corresponding to the keyword. For example, keywords such as the name of the other party and their corresponding destinations can be pre-associated.

[0055] Such voice response systems can assist with tasks such as transferring phone calls and calling reception staff. Furthermore, in the above-mentioned voice response device, as in the invention of the 31st phase, the device may recognize the requirements of what the other party is saying based on keywords and convey a summary of what the other party has said to the user.

[0056] Such voice response devices can assist in tasks such as liaising with customers. Furthermore, in the above-mentioned voice response device, as in the invention of the 32nd phase, Emotion determination means that reads emotions from the tone of voice input by the user and outputs which emotion it corresponds to, from among emotions including at least one of the following: normal, anger, joy, confusion, sadness, and exhilaration. It may also be equipped with.

[0057] Such voice response devices can output responses according to the user's emotions. Next, the invention of the 33rd phase is, A response generation means that generates a response corresponding to an image captured around the voice response device when the aforementioned character information is input, A voice output means for outputting the aforementioned response as voice, It is characterized by having the following features.

[0058] Such an audio response device can output a response in audio format according to the captured image. Therefore, it can improve usability compared to a configuration that generates a response from text information only.

[0059] Specific configurations of the present invention include, for example, a configuration in which text information is input to respond with what is recognized, and what (or who) is recognized from the captured image is output as audio.

[0060] By the way, in the above-mentioned voice response device, as in the invention of the 34th phase, A location identification means search means that searches for objects contained in text information from an captured image using image processing and identifies the location of the searched object, A guiding means for guiding the object to its position, It may also be equipped with.

[0061] Such a voice response device can guide the user to an object in the captured image. Furthermore, in the above-mentioned voice response device, as in the invention of the 35th phase, When inputting text information by voice, a video image is acquired capturing the shape of the user's mouth. Voice input video acquisition method, A text information conversion means that converts the aforementioned audio into text information and corrects the text information by estimating unclear parts of the audio based on the video image, It may also be equipped with.

[0062] Such a voice response device can estimate the content of speech from the shape of the mouth, thus enabling accurate estimation of unclear parts of speech. Furthermore, in the above-mentioned voice response device, as in the invention of the 36th phase, The message playback mechanism detects the user's frustration or agitation by detecting unexpected sounds, and generates a message to suppress that frustration or agitation. You may do so.

[0063] Such voice response devices can suppress frustration and agitation in the user, thereby reducing the likelihood of conflicts between the user and those around them. Furthermore, in the above-mentioned voice response device, as in the invention of the 37th phase, When providing directions to a destination, the system includes a route information acquisition means for acquiring route information such as weather, temperature, humidity, traffic information, and road surface conditions to the destination. The message playback method outputs route information as audio. You may do so.

[0064] Such voice response devices can notify the user of their progress (route information) towards their destination via voice. Furthermore, in the above-mentioned voice response device, as in the invention of the 38th phase, A gaze detection means for detecting the user's gaze, If the user's gaze does not move to a predetermined position in response to the call from the message playback means, the gaze movement request transmission means outputs an audio message requesting the user to move their gaze to the predetermined position. It may also be equipped with.

[0065] Such voice response devices allow the user to see a specific location. Therefore, safety checks during vehicle operation can be reliably performed. Furthermore, in the above-mentioned voice response device, as in the invention of the 39th phase, A change request transmission means observes the position of body parts and facial expressions, and if there is little change in response to the aforementioned call, outputs a voice requesting that the position of body parts and facial expressions be changed. It may also be equipped with.

[0066] Such a voice response device can be used to move parts of the user's body to specific locations or to guide them to make specific facial expressions. This invention can be used when driving a vehicle or during physical examinations.

[0067] Furthermore, in the above-mentioned voice response device, as in the invention of the 40th phase, A means for acquiring broadcast programs similar to the broadcast programs that the user is watching, A broadcast program completion means that, when a broadcast program is interrupted, compensates for the interrupted broadcast program by outputting the broadcast program it has acquired, It may also be equipped with.

[0068] Such voice response devices can compensate for interruptions in the broadcast program that the user is watching. Furthermore, in the above-mentioned voice response device, as in the invention of the 41st phase, The system includes a lyric-adding mechanism that, when a user adds lyrics to a song without lyrics and sings it, compares the song with lyrics and the lyrics added by the user, and outputs the lyrics as sound only in the parts where the user's lyrics are missing.

[0069] Such a voice response device can compensate for parts of a karaoke session that the user cannot sing (parts of the lyrics that are cut off). Furthermore, in the above-mentioned voice response device, as in the invention of the 42nd phase, When an image contains text, and the user asks for the pronunciation of this text, a pronunciation output means obtains information about the text from an external source and outputs the pronunciation of the text contained in this information as audio. It may also be equipped with.

[0070] Such voice response devices can teach users how to read characters. Furthermore, in the above-mentioned voice response device, as in the invention of the 43rd phase, It is equipped with behavioral environment detection means that detect the user's actions and the user's surrounding environment, The message generation means generates messages in response to detected behavior and the surrounding environment. You may do so.

[0071] Such voice response devices can alert users to dangerous areas or restricted zones. They can also detect unusual behavior by the user. Furthermore, in the above-mentioned voice response device, as in the invention of the 44th phase, A health status determination means that determines the health status based on captured images of the user, A health message generation means that generates messages according to the health status, It may also be equipped with.

[0072] Such voice response devices can be used to manage the user's health status. Furthermore, in the above-mentioned voice response device, as in the invention of the 45th phase, A notification system that notifies a designated contact person if the health status falls below the standard value. It may also be equipped with.

[0073] Such voice response devices can alert users if their health status falls below a certain threshold. Therefore, abnormalities can be notified to others more early. Furthermore, the above-mentioned voice response device may be configured to output information about the user in response to inquiries from persons other than the user, as in the invention of the 46th phase.

[0074] Such a voice response device can, for example, detect the user's diet and walking distance, and then answer questions on their behalf at hospitals and other medical facilities. It could also be programmed to learn the user's health status and self-introduction.

[0075] Furthermore, each invention in its respective phase does not need to presuppose other inventions, and should be as independent as possible. It is possible. [Brief explanation of the drawing]

[0076] [Figure 1] This is a block diagram showing a schematic configuration of a voice response system to which the present invention is applied. [Figure 2] This is a block diagram showing the schematic configuration of the terminal device. [Figure 3] This flowchart shows the voice response terminal processing performed by the MPU of the terminal device. [Figure 4] This is a flowchart showing the voice response server processing performed by the server's processing unit. [Figure 5] This is an explanatory diagram showing an example of a response candidate database. [Figure 6] This is a flowchart showing the automated conversation terminal processing performed by the MPU of the terminal device. [Figure 7] This is a flowchart showing the automated conversation server processing performed by the server's processing unit. [Figure 8] This is a flowchart showing the message terminal processing performed by the MPU of the terminal device. [Figure 9] This flowchart shows the message server processing performed by the server's processing unit. [Figure 10] This is a flowchart showing the induction terminal processing performed by the MPU of the terminal device. [Figure 11] This flowchart shows the induction server processing performed by the server's computing unit. [Figure 12] This flowchart shows the reception process executed by the server's processing unit. [Figure 13] This is a flowchart showing the information provision terminal processing performed by the MPU of the terminal device. [Figure 14] This is an explanatory diagram showing an example of a personality database. [Figure 15] This flowchart shows the personality information generation process performed by the MPU of the terminal device. [Figure 16] This is an explanatory diagram showing an example of a preference database. [Figure 17] This flowchart shows the preference information generation process performed by the server's processing unit. [Figure 18] This is an explanatory diagram showing examples of combinations between personality categories and preferences. [Figure 19] This flowchart shows the character input process performed by the server's processing unit. [Figure 20] This is a flowchart showing the processing performed by the server's processing unit when using other terminals. [Figure 21] This flowchart shows the memory verification process performed by the server's processing unit. [Figure 22] This flowchart shows the pronunciation determination process 1 executed by the server's processing unit. [Figure 23] This flowchart shows the pronunciation determination process 2 executed by the server's processing unit. [Figure 24] This flowchart shows the pronunciation determination process 3 executed by the server's processing unit. [Figure 25] This flowchart shows the emotion determination process performed by the server's processing unit. [Figure 26] This flowchart shows the emotion response generation process performed by the server's processing unit. [Figure 27] This flowchart shows the guidance process executed by the server's processing unit. [Figure 28] This is a flowchart showing the movement request processing 1 executed by the server's processing unit. [Figure 29] This flowchart shows the movement request processing 2 executed by the server's processing unit. [Figure 30] This flowchart shows the broadcast music completion process performed by the server's processing unit. [Figure 31] This flowchart shows the character interpretation process performed by the server's processing unit. [Figure 32] This is a flowchart showing the action response terminal processing performed by the server's processing unit. [Figure 33] This flowchart shows the action response server processing performed by the server's processing unit. [Modes for carrying out the invention]

[0077] Embodiments of the present invention will be described below with reference to the drawings. [First Embodiment] [Structure of this embodiment] The voice response system 100 to which the present invention is applied is configured such that, in response to voice input from a terminal device 1, a server 90 generates an appropriate response, and the terminal device 1 outputs the response as voice. More specifically, as shown in Figure 1, the voice response system 100 is configured such that multiple terminal devices 1 and the server 90 can communicate with each other via a communication base station 80 or the Internet network 85.

[0078] Server 90 has the functions of a normal server device. In particular, Server 90 has an arithmetic unit 101 and various databases (DBs). The arithmetic unit 101 is configured as a well-known arithmetic unit equipped with a CPU and memory such as ROM and RAM, and performs various processes based on programs in memory, such as communication with terminal devices 1 etc. via the internet network 85, reading and writing data in various DBs, and speech recognition and response generation for conversations with users using terminal devices 1.

[0079] As shown in Figure 1, the various databases include speech recognition DB102, predictive text DB103, voice DB104, response candidate DB105, personality DB106, learning DB107, preference DB108, news DB109, weather DB110, playback conditions DB111, handwriting / sign language DB112, terminal information DB113, emotion judgment DB114, health judgment DB115, karaoke DB116, notification destination DB117, sales DB118, client DB119, etc. Details of these databases will be described as needed in each processing explanation.

[0080] Next, as shown in Figure 2, the terminal device 1 is configured with a behavior sensor unit 10, a communication unit 50, a notification unit 60, and an operation unit 70, all housed in a predetermined enclosure. The behavioral sensor unit 10 is equipped with a well-known MPU 31 (microprocessor unit), memory 39 such as ROM and RAM, and various sensors. The MPU 31 performs processing, such as driving a heater to optimize the temperature of the sensor elements, so that the sensor elements constituting the various sensors can detect the object to be inspected (humidity, wind speed, etc.) in an effective manner.

[0081] The behavior sensor unit 10 includes various sensors, including a 3D acceleration sensor 11 (3DG sensor), a 3-axis gyroscope sensor 13, a temperature sensor 15 located on the back of the housing, a humidity sensor 17 located on the back of the housing, a temperature sensor 19 located on the front of the housing, a humidity sensor 21 located on the front of the housing, an illuminance sensor 23 located on the front of the housing, a wetness sensor 25 located on the back of the housing, a GPS receiver 27 for detecting the current location of the terminal device 1, and a wind speed sensor 29.

[0082] The behavior sensor unit 10 also includes various sensors, such as an electrocardiogram sensor 33, a heart sound sensor 35, a microphone 37, and a camera 41. The temperature sensors 15 and 19, and the humidity sensors 17 and 21 measure the temperature or humidity of the outside air of the housing as the object of inspection.

[0083] The 3D acceleration sensor 11 detects acceleration applied to the terminal device 1 in three mutually orthogonal directions (vertical direction (Z direction), width direction of the housing (Y direction), and thickness direction of the housing (X direction)) and outputs the detection result.

[0084] The 3-axis gyro sensor 13 detects angular acceleration (with counterclockwise velocities in each direction being positive) in the vertical direction (Z direction) and any two directions orthogonal to the vertical direction (the width direction of the housing (Y direction) and the thickness direction of the housing (X direction)) as angular velocity applied to the terminal device 1, and outputs the detection results.

[0085] The temperature sensors 15 and 19 are configured to include, for example, a thermistor element whose electrical resistance changes according to the temperature. In this embodiment, the temperature sensors 15 and 19 detect Celsius temperature, and all temperature displays described below will be in Celsius.

[0086] The humidity sensors 17 and 21 are configured, for example, as well-known polymer film humidity sensors. These polymer film humidity sensors are configured as capacitors in which the amount of water contained in the polymer film changes in response to changes in relative humidity, and the dielectric constant changes accordingly.

[0087] The illuminance sensor 23 is configured as a well-known illuminance sensor equipped with, for example, a phototransistor. The wind speed sensor 29 is, for example, a well-known wind speed sensor that calculates the wind speed from the power (heat dissipation) required to maintain the heater temperature at a predetermined temperature.

[0088] The heart sound sensor 35 is configured as a vibration sensor that captures vibrations caused by the user's heartbeat, and the MPU 31 considers the detection results from the heart sound sensor 35 and the heart sounds input from the microphone 37 to distinguish between vibrations and noises caused by heartbeats and other vibrations and noises.

[0089] The wetness sensor 25 detects water droplets on the surface of the housing, and the electrocardiogram sensor 33 detects the user's heartbeat. The camera 41 is positioned inside the housing of the terminal device 1 so that its imaging range is the area outside the terminal device 1.

[0090] The communication unit 50 comprises a well-known MPU 51, a wireless telephone unit 53, and a contact memory 55, and is configured to acquire detection signals from various sensors constituting the behavior sensor unit 10 via an input / output interface (not shown). The MPU 51 of the communication unit 50 then performs processing according to the detection results from the behavior sensor unit 10, input signals input via the operation unit 70, and programs stored in ROM (not shown).

[0091] Specifically, the MPU 51 of the communication unit 50 performs the following functions: an action detection device that detects specific actions performed by the user; a positional relationship detection device that detects the positional relationship with the user; an exercise load detection device that detects the load of exercise performed by the user; and a function that transmits the processing results from the MPU 51.

[0092] The wireless telephone unit 53 is configured to communicate with, for example, a mobile phone base station, and the MPU 51 of the communication unit 50 outputs the processing results of the MPU 51 to the notification unit 60 or transmits them via the wireless telephone unit 53 to a pre-set destination.

[0093] The contact memory 55 functions as a storage area for storing location information of the user's visited locations. This contact memory 55 also stores information about contacts (such as phone numbers) to be contacted in the event of an emergency involving the user.

[0094] The notification unit 60 includes, for example, a display 61 configured as an LCD or an organic EL display, an illumination 63 consisting of, for example, LEDs capable of emitting seven colors, and a speaker 65. Each component of the notification unit 60 is driven and controlled by the MPU 51 of the communication unit 50.

[0095] Next, the operating unit 70 includes a touchpad 71, a confirmation button 73, a fingerprint sensor 75, and a rescue request lever 77. The touchpad 71 outputs signals corresponding to the position and pressure when touched by the user (the user or their guardian, etc.).

[0096] The confirmation button 73 is configured such that when pressed by the user, the contacts of the built-in switch close, allowing the communication unit 50 to detect that the confirmation button 73 has been pressed.

[0097] The fingerprint sensor 75 is a well-known fingerprint sensor, configured to read fingerprints using, for example, an optical sensor. Alternatively, any means capable of recognizing human physical characteristics (means capable of biometric authentication: means capable of identifying an individual), such as a sensor that recognizes the shape of veins in the palm, can be used instead of the fingerprint sensor 75.

[0098] It also features a rescue request lever 77 that, when operated, connects to a designated contact person. [Processing in this embodiment] The processes performed in such a voice response system 100 are described below.

[0099] The voice response terminal processing performed by terminal device 1 involves receiving voice input from the user, sending this voice to server 90, and, upon receiving a response from server 90, playing this response aloud. This process is initiated when the user indicates via the operation unit 70 that they wish to perform voice input.

[0100] In detail, as shown in Figure 3, first, the system is set to accept input from the microphone 37 (ON state) (S2), and then imaging (recording) by the camera 41 is started (S4). Then, it is determined whether or not there is audio input (S6).

[0101] If there is no voice input (S6: NO), it is determined whether a timeout has occurred (S8). Here, a timeout indicates that the allowable time for waiting for processing has been exceeded, and in this case, the allowable time is set to approximately 5 seconds.

[0102] If a timeout occurs (S8:YES), the process proceeds to S30, which will be described later. If a timeout does not occur (S8:NO), the process returns to S6. If there is voice input (S6:YES), the voice is recorded in memory (S10), and it is determined whether the voice input has ended or not (S12). Here, it is determined that the voice input has ended if the voice is interrupted for a certain period of time or if an input to end the voice input is received via the operation unit 70.

[0103] If voice input has not been completed (S12: NO), the process returns to S10. If voice input has been completed (S12: YES), data such as an ID for self-identification, voice, and captured image is sent as a packet to server 90 (S14). Note that the data transmission process may be performed between S10 and S12.

[0104] Next, it is determined whether the data transmission is complete (S16). If the transmission is not complete (S16: NO), the process returns to S14. Furthermore, if transmission is complete (S16:YES), it is determined whether or not the data (packet) to be sent in the voice response server processing described later has been received (S18). If the data has not been received (S18:NO), it is determined whether or not a timeout has occurred (S20).

[0105] If a timeout occurs (S20:YES), the process proceeds to S30, which will be described later. If a timeout does not occur (S20:NO), the process returns to S18. Furthermore, if data has been received (S18:YES), a packet is received (S22). In this process, one or more different responses to the character information are obtained, each associated with a different tone of voice.

[0106] Then, it is determined whether reception is complete or not (S24). If reception is not complete (S24: NO), it is determined whether a timeout has occurred or not (S26). If a timeout occurs (S26:YES), an error is reported via the notification unit 60, and the voice response terminal processing is terminated. If a timeout does not occur (S26:NO), the process returns to S22.

[0107] Furthermore, if reception is complete (S24: YES), a voice response based on the received packet is output from speaker 65 (S28). In this process, if multiple responses are to be played, each response is played in a different voice. Once this process is complete, the voice response terminal process is terminated.

[0108] Next, the voice response server processing performed by server 90 (external device) will be explained using Figure 4. The voice response server processing is the process of receiving voice from terminal device 1, performing speech recognition to convert this voice into text information, and generating a response to the voice and returning it to terminal device 1. In particular, in this embodiment, multiple responses may be transmitted in association with voices of different tones.

[0109] As shown in Figure 4, the voice response server processing is as follows: First, it is determined whether or not a packet has been received from any of the terminal devices 1 (S42). If no packet has been received (S42: NO), the process in S42 is repeated.

[0110] Furthermore, if a packet has been received (S42: YES), the communication partner terminal device 1 is identified (S44). In this process, terminal device 1 is identified by the ID of terminal device 1 contained in the packet.

[0111] Next, the speech contained in the packet is recognized (S46). Here, in the speech recognition DB102, many speech waveforms are associated with many characters. In addition, in the predictive text DB103, words that tend to be used following a particular word are associated.

[0112] Therefore, in this process, by referring to the speech recognition DB102 and predictive text DB103, a well-known speech recognition process is performed to convert speech into text information. Next, the captured image is processed to identify objects within the image (S48). Then, the user's emotions are determined based on the waveform of the sound and the endings of words (S50).

[0113] In this process, the system refers to the emotion determination DB114, which associates speech waveforms (voice tone) and word endings with categories of emotions such as anger, joy, confusion, sadness, and exhilaration, to determine whether the user's emotion falls into one of these categories, and records this determination result in memory. Subsequently, the system refers to the learning DB107 to search for words that the user frequently speaks and corrects any ambiguous parts of the text information generated by speech recognition.

[0114] Furthermore, the learning database 107 records each user's characteristics, such as words they frequently use and their pronunciation habits. Data is also added to and modified in the learning database 107 during conversations with the user.

[0115] Next, the corrected character information is identified as the input character information (S54), and a response is obtained from the response candidate DB105 by searching for a sentence similar to the character information as input (S56). Here, the response candidate DB105 contains the input character information, the first output, the tone of voice of the first output, the second output, and the tone of voice of the second output, as shown in Figure 5. It is associated with righteousness.

[0116] For example, as shown in the first row of Figure 5, when the text information "Today's weather in *" is input, the first output "Today's weather in * is *" is output, corresponding to the voice of female 1. However, the "*" part is obtained by accessing the weather DB110, which associates the region name with the weather forecast for that region for several days.

[0117] Additionally, if the text information "Today's weather in *" is entered, the weather at the time the weather changes today is also retrieved from the weather DB110, and a second output, "However, * is *," is output, corresponding to the voice of Male 1. If the weather in Tokyo today is sunny and the weather tomorrow is rainy, and "Today's weather in Tokyo" is entered, the voice of Female 1 will output "Today's weather in Tokyo is sunny," and the voice of Male 1 will output "However, it will rain tomorrow."

[0118] In this embodiment, we have described the case where multiple responses are output, but if there is only one answer to the input, there will be only one response. Therefore, we determine whether there is only one response or not (S58). If there is only one response (S58: YES), we proceed to the process in S62 described later.

[0119] Furthermore, if there are multiple responses (S58:NO), the response content is associated with the tone of voice (S60). Here, the voice DB104 stores a database of artificial voices for each tone of voice, and in this process, the tone of voice set for each response is associated with the tone of voice in the database.

[0120] Next, the response content is converted into speech (S62). In this process, the response content (text information) is output as speech based on the database stored in the speech DB104.

[0121] Then, the generated response (voice) is sent as a packet to the communication partner's terminal device 1 (S64). Alternatively, the voice of the response may be generated while the packet is being sent. Next, the conversation content is recorded (S68). In this process, the input character information and the output response content are recorded as conversation content in the learning DB107. At this time, keywords included in the conversation content (words recorded in the speech recognition DB102) and characteristics of pronunciation are also recorded in the learning DB107.

[0122] Once this process is complete, the voice response server process will terminate. [Effects of this embodiment] As described above, the voice response system 100 is a system that provides voice responses to input character information, and the terminal device 1 (MPU 31) acquires multiple different responses to the character information and outputs each of the multiple different responses in a different voice tone.

[0123] With such a voice response system 100, multiple responses can be output in different tones of voice. Therefore, even when a single character of information cannot be uniquely identified, different solutions can be output to the user in a clear and understandable manner using different tones. Thus, the system can be made more user-friendly.

[0124] Furthermore, in the above-described voice response system 100, the terminal device 1 receives voice input from the user via the microphone 37, the server 90 (processing unit 101) converts the input voice into text information, generates multiple different responses to the text information and transmits them to the terminal device 1. The terminal device 1 then receives the response from the server 90.

[0125] With such a voice response system 100, the terminal device 1 can receive voice input, allowing for a configuration where text information is input via voice. Furthermore, since the server 90 can generate the response, the processing load on the voice response system 100 can be reduced.

[0126] Furthermore, in the above-mentioned voice response system 100, the server 90 converts the user's speech into text information and stores speech habits (such as pronunciation quirks) as learning information (capturing and recording these characteristics).

[0127] Such a voice response system 100 can generate text information based on learned information, thereby improving the accuracy of text information generation. Furthermore, in the above-mentioned voice response system 100, the server 90 reads the emotion from the tone of voice of the voice input by the user and outputs which emotion it corresponds to, from among emotions including at least one of the following: normal, anger, joy, confusion, sadness, and exhilaration.

[0128] Such a voice response system 100 can output a response according to the user's emotions. [Modified version of the first embodiment] In this embodiment, speech recognition was used as the input method for character information, but input may be performed using other input means (operation unit 70), such as a keyboard or touch panel, rather than being limited to speech recognition. Also, although the operation of "converting the input speech into character information" was performed by the server 90, it may also be performed by the terminal device 1.

[0129] Furthermore, in the above-described voice response system 100, the server 90 is provided with a response candidate DB 105 in which multiple different responses, including positive and negative responses for each of the multiple pieces of character information, are recorded. The terminal device 1 may acquire the positive and negative responses as multiple different responses and play them back in different tones for the positive and negative responses.

[0130] For example, as shown in the second row of Figure 5, when a voice input asks "Is it okay to buy this item?", positive information about the item, such as good reviews, is output in a female voice. On the other hand, negative information, such as bad reviews, is output in a different voice (in this case, a male voice) than the female voice associated with the positive information.

[0131] With such a voice response system 100, it is possible to reproduce responses from different positions, such as positive and negative responses, in different tones of voice, making it sound as if different people are speaking. Therefore, it is less likely that the user listening to the voice will feel any sense of unease.

[0132] Furthermore, the tone of voice may be changed depending on the type of response and the wording used in the response. For example, if the response is to be given in a gentle tone, it could be played in a calm female voice, and if the response is given in an aggressive tone, it could be played in a brave male voice. In other words, the response content should be associated with personality, and the tone of voice should be set according to the personality.

[0133] Furthermore, the voice response system 100 may also input responses (for example, positive or negative responses) output by its own terminal device 1 or other terminal devices 1 as text information and generate responses to counter these responses. In other words, from the user's perspective, they can hear arguments from both the pro and con viewpoints. After hearing these arguments, the user can then make a final decision.

[0134] This configuration can be implemented using one or more terminal devices 1. Multiple terminal devices 1 can exchange voice with each other by directly inputting and outputting voice, or by using wireless or other communication methods. When multiple terminal devices 1 communicate with the server 90, data can be sent to other terminal devices 1 during the processing in S66.

[0135] Furthermore, in the above-mentioned voice response system 100, the calculation unit 101 may learn (record and analyze) the user's actions (conversation, places moved, things captured by the camera) and supplement any omissions in the user's conversation.

[0136] For example, in a conversation where the user responds to the question, "Is hamburger okay today?" with "I'd prefer curry," if this device adds, "We had hamburger yesterday," the reason why the user said they preferred curry becomes clear.

[0137] Furthermore, such a configuration can be implemented during a phone call, and it may also be configured to spontaneously join the user's conversation. Furthermore, in the above-described voice response system 100, the server 90 may obtain response candidates from a predetermined server or from the Internet.

[0138] With such a voice response system 100, response candidates can be obtained not only from the server 90 but also from any device connected via the internet, a dedicated line, or the like. [Second Embodiment] [Processing in the second embodiment] Next, a different form of voice response system will be described. In this embodiment (second embodiment) and subsequent embodiments, only the parts that differ from the voice response system 100 of the first embodiment will be described in detail, and parts that are the same as those of the voice response system 100 of the first embodiment will be denoted by the same reference numerals and their description will be omitted.

[0139] In the second embodiment of the voice response system, voice is output even when the user does not input text information. Specifically, terminal device 1 performs the automatic conversation terminal processing shown in Figure 6. Automatic conversation terminal processing is a process that starts, for example, when the power to terminal device 1 is turned on, and is subsequently executed repeatedly.

[0140] In the automated conversation terminal processing, it is first determined whether the setting for automated conversation is turned ON (S82). Note that the user can set whether or not to perform automated conversation via the operation unit 70 or by inputting voice.

[0141] If the automatic conversation setting is OFF (S82: NO), the automatic conversation terminal process is terminated. If the automatic conversation setting is ON (S82: YES), a message is sent to the server 90 indicating that automatic conversation mode has been set, along with an ID to identify itself (S84).

[0142] Next, it is determined whether or not a packet has been received from server 90 (S86). If no packet has been received (S86: NO), the process in S86 is repeated. If a packet has been received (S86: YES), the same process as described in S22-S30 is performed, and once these processes are completed, the automated conversation terminal process is terminated.

[0143] Furthermore, server 90 executes the automated conversation server processing shown in Figure 7. This automated conversation server processing is initiated, for example, when server 90 is powered on, and is then repeatedly executed.

[0144] In the automated conversation server processing, it is first determined whether or not a message indicating that the automated conversation mode has been set has been received from terminal device 1 (S92). If no such message has been received (S92:NO), the process proceeds to S98.

[0145] If a message indicating that automatic conversation mode has been set is received (S92:YES), the terminal device 1 to be the communication partner is identified based on the ID contained in the received packet (S94), and the system is set to initiate an automatic conversation with this communication partner (S96). Subsequently, for each terminal device 1 that has been set to initiate an automatic conversation, it is determined whether the playback conditions are met (S98).

[0146] Here, the playback conditions refer to things like a certain amount of time having passed since the last conversation (voice input), a specific time of day, specific weather conditions, or when any sensor value indicates an abnormal value.

[0147] If the playback conditions are not met (S98:NO), the automated conversation server process is terminated. If the playback conditions are met (S98:YES), a message corresponding to the playback conditions is generated (S100).

[0148] Here, a message that corresponds to the playback conditions may be a standard phrase such as "Good morning" or "Hello," or it may be related to the latest news obtained from the News DB109, which is automatically updated with the latest news. If the message is related to the latest news, for example, if information on the stock price of a certain company is obtained, it could say, "The stock price of XX company rose by XX yen today. Did you know?"

[0149] Once this process is complete, the processes described in S42 to S54 above are performed. After the process in S54 is completed, it is determined whether or not a predetermined response has been received from the terminal device 1, which is the communication partner (S112). Here, the predetermined response may be, for example, some kind of voice, or a specific answer. A specific answer may be, for example, "Do you know?" the answer would be "Yes" or "No," and "What's the weather like now?" the answer would be "It's raining" or "It's sunny," or something that includes a word indicating the weather.

[0150] If the required response is received (S112: YES), the automated conversation server process terminates. If the required response is not received (S112: NO), the message sent in S100 is resent (S114). When resending the message in this way, the voice tone is changed, and a stronger, stern tone is generated.

[0151] Next, the system refers to the notification destination DB117, which has been pre-associated with the terminal device 1 and the notification destinations, and sends a message to the designated notification destination indicating that no response was received (S116). Once this process is complete, the automated conversation server process is terminated.

[0152] [Effects of the second embodiment] In the above-described voice response system 100, the server 90 determines whether the status of the voice response system 100 matches the playback conditions set in advance for outputting voice when no text information is input. If the playback conditions are met, the server outputs a pre-set message.

[0153] Such a voice response system 100 can output voice even when no text information is input (i.e., when the user does not speak). For example, by forcing the user to speak, it can be used as a measure to suppress drowsiness while driving. Furthermore, by determining whether or not a person living alone responds, it is possible to check on their well-being.

[0154] Furthermore, in the above-mentioned voice response system 100, the server 90 acquires news information and outputs a message related to the news in the form of a question that prompts the user for an answer. With such a voice response system 100, it is possible to have conversations about the news, thus preventing the conversation from becoming repetitive.

[0155] Furthermore, in the above-mentioned voice response system 100, the server 90 adds separately acquired external information (such as news and environmental information like temperature, weather, and location information) to a pre-configured message and outputs it.

[0156] Such a voice response system 100 can output a response that combines a predetermined message with acquired information. Furthermore, in the above-mentioned voice response system 100, if the server 90 does not receive a response or answer to a message, it sends information identifying the user and a statement that no response was received to a pre-configured contact.

[0157] With such a voice response system 100, if no response is received, a notification can be sent to a contact person. Therefore, for example, abnormalities in elderly people living alone can be reported early.

[0158] [Modified version of the second embodiment] Furthermore, in the above-described voice response system 100, the server 90 may acquire multiple messages and select and output a message to be played according to the frequency of message playback.

[0159] With such a voice response system 100, it is possible to create randomness in message playback by making it difficult to play frequently played messages, or to encourage attention and memory retention by deliberately playing frequently played messages repeatedly.

[0160] [Third Embodiment] [Processing in the third embodiment] Next, in the third embodiment of the voice response system, the terminal device 1 is configured to convey things that the user finds difficult to say directly to someone. For example, if the user tells the device before a date that they want to say a certain thing today, the voice response system 100 will speak on their behalf (play an audio recording) at an appropriate time (for example, at a pre-set time or after a certain amount of time has passed since the conversation ended).

[0161] In detail, terminal device 1 performs the message terminal processing shown in Figure 8, and server 90 performs the message server processing shown in Figure 9. The message terminal processing is initiated, for example, when terminal device 1 is powered on, and is subsequently executed repeatedly.

[0162] In the message terminal processing, as shown in Figure 8, it is first determined whether or not the message mode has been set by the user (S132). If the message mode has not been set (S132: NO), the process in S132 is repeated.

[0163] Furthermore, if message mode is set (S132: YES), processes S2 to S8 are performed, and if a positive determination is made in S6, the message mode flag is set to the ON state in the memory of terminal device 1 (S134). Then, processes S10 to S16 are performed.

[0164] If the result in S16 is positive, it is determined whether or not a packet has been received from server 90 (S136). If no packet has been received (S136: NO), the process in S136 is repeated. If a packet has been received (S136: YES), the processes in S24 to S30 are performed, and the message terminal process is terminated.

[0165] Next, the message server processing is a process that starts, for example, when the server 90 is powered on, and is then executed repeatedly. Specifically, it first determines whether or not a packet has been received from any terminal device 1 (S142). If no packet has been received (S142: NO), the process proceeds to S156, which will be described later.

[0166] Furthermore, if a packet has been received (S142:YES), the terminal device 1 of the communication partner is identified (S44), and it is determined whether or not the packet contains mode flags such as a message mode flag (S144). If there are no mode flags (S144:NO), the process proceeds to S148.

[0167] Furthermore, if a mode flag exists (S144: YES), the server 90 also sets the mode by setting the flag corresponding to the communication partner terminal device 1 to the ON state (S146). For example, if the message mode flag corresponds to the message mode, the processes S46 to S152 described later will be performed, and if the induction mode flag described later corresponds to the induction mode, S46 to S176 (see Figure 11) will be performed.

[0168] Next, it is determined whether the message flag is in the ON state (S148). If the message flag is in the ON state (S148: YES), the processes in S46 to S54 are performed, and then the message playback conditions are extracted (S150).

[0169] Here, the message playback conditions can be set in advance by the user via the operation unit 70 of the terminal device 1, and include, for example, time and location. The message playback conditions are sent to the server 90 when the message terminal processing packet is transmitted.

[0170] Next, the message and the voice (tone of voice) are associated and recorded in memory (S152), and the process proceeds to S156. If the message flag is OFF (S148: NO), processing related to other modes is performed (S154), and it is determined whether or not it is time for playback (S156). Here, playback timing refers to the content set in the message playback conditions.

[0171] If it is not the playback timing (S156: NO), the message server processing is terminated immediately. If it is the playback timing (S156: YES), the processes in S62 to S64 are executed, and the message server processing is terminated.

[0172] [Effects of the Third Embodiment] According to this third embodiment of the voice response system, the voice input by the user is not played back immediately, but rather can be played back after a certain period of time when the message playback conditions are met.

[0173] For example, as shown in the third row of Figure 5, if you input "Tell XX to XX," the message you want to convey will be played only after XX's voice is recognized (heard).

[0174] [Modified version of the third embodiment] In the third embodiment described above, the system was configured to play back what the user said, but to The setup may also include a phrase that serves as a trigger for conversation, such as, "By the way, didn't you say you were going to tell her something?" In detail, terminal device 1 performs the guidance terminal processing shown in Figure 10, and server 90 performs the guidance server processing shown in Figure 11.

[0175] The induction terminal processing is a process that starts when, for example, the power to terminal device 1 is turned on, and is then repeatedly executed.

[0176] In the induction terminal processing, as shown in Figure 10, it is first determined whether the induction mode has been set by the user (S162). If the induction mode has not been set (S162:NO), the process in S162 is repeated.

[0177] Furthermore, if induction mode is set (S162: YES), processes S2 to S8 are performed, and if a positive determination is made in S6, the induction mode flag is set to the ON state in the memory of terminal device 1 (S164). Then, processes S10 to S16 are performed.

[0178] If the result in S16 is positive, it is determined whether or not a packet from server 90 has been received (S166). If no packet has been received (S166: NO), the process in S166 is repeated. If a packet has been received (S166: YES), the processes in S24 to S30 are executed, and the guidance terminal process is terminated.

[0179] Next, the induction server processing is a process that starts, for example, when the power to server 90 is turned on, and is then executed repeatedly. Specifically, it executes the processes described in S142 to S146 above. Then, it is determined whether the induction flag is in the ON state or not (S172).

[0180] If the induction flag is ON (S172:YES), the processes in S46 to S54 are performed, followed by the extraction of induction regeneration conditions (S174). Here, the guidance playback conditions, like the message playback conditions, can be set in advance by the user via the operation unit 70 of the terminal device 1, and these conditions include, for example, time and location. The guidance playback conditions are transmitted to the server 90 when sending packets for message terminal processing.

[0181] Next, guidance content is generated, and this guidance content is associated with the voice (tone of voice) and recorded in memory (S176). Here, the guidance content is, for example, a search for words expressing desires such as "want to" or "hope" contained in the input text information, the keywords preceding these words are extracted, and words registered as guiding words for these keywords are output as guidance content. Note that the keywords and the words indicating the guidance content are pre-associated and recorded in the response candidate DB105.

[0182] Next, the processes described in S156 and below are performed, and the server processing is terminated. Also, if the induction flag is OFF (S172: NO), processing related to other modes is performed (S154), the processes described in S156 and below are performed, and the server processing is terminated.

[0183] According to this modified configuration of the third embodiment, instead of directly outputting the words the user wants to say, it is possible to guide the user to speak the words they want to say. [Fourth Embodiment] [Processing in the fourth embodiment] Next, an example of using terminal device 1 for reception work will be described. In this embodiment, terminal device 1 is installed in the company's reception area, etc. It can also be used for telephone reception, such as the company's main telephone line or telephone banking. In this embodiment, the process in S56 in the first embodiment is replaced with the reception process shown in Figure 12.

[0184] In the reception process, as shown in Figure 12, it is first determined whether or not the text information contains a company name (S192). This process determines whether or not it contains a common name or a company name (recorded in the speech recognition DB 102).

[0185] If the text information does not contain a company name or personal name (S192: YES), a response is generated to ask for the company name and personal name (S194), and the reception process is terminated. This process generates a response such as, "Please tell us your name and purpose of your inquiry."

[0186] If the text information contains a company name or individual name (S192: NO), this company name or individual name is extracted from the sales DB 118 and client DB 119 (S196). Here, the sales DB 118 records the names of companies and representatives who have made sales visits in the past, or the names of complainers who only make complaints. The client DB 119 records the company name, the representative of that company, the representative on the user side of terminal device 1 (our company's side), the schedule such as the scheduled meeting time, and contact information associated with each representative.

[0187] Next, it is determined whether the company name and individual name could be extracted from the sales database 118, that is, whether the company name and individual name contained in the text information were included in the sales database 118 (S198). If the company name and individual name could be extracted from the sales database 118 (S198: YES), a sales rejection response (a response refusing to forward the call) is generated (S200), and the reception process is terminated.

[0188] Furthermore, if the company name or individual name cannot be extracted from the sales DB118 (S198:NO), it is determined whether the person who came to the reception desk is scheduled to visit at a nearby time (for example, within one hour before or after the current time) according to the schedule in the client DB119 (S202). If the person is scheduled to visit at a nearby time (S202:YES), the contact information of the person in charge of this person is extracted from the client DB119, and the person who came to the reception desk is connected to this person so that they can communicate with them (S204). For this process, it is sufficient to connect to the person in charge's extension phone, mobile phone, etc.

[0189] Next, a reception response for the client is generated (S206). Here, the reception response for the client is generated as an example, such as, "Thank you for your continued patronage, Mr. / Ms. XX. We are connecting you to the appropriate person, please wait a moment." Once this process is complete, the reception process is terminated.

[0190] Furthermore, if the person is not scheduled to visit at the same time (S202: NO), the system connects to a pre-configured reception contact and connects the person to this receptionist so that they can communicate with the receptionist (S208). Then, a normal reception response is generated (S210).

[0191] At this point, a typical reception response would be something like, "We are connecting you to reception, please wait a moment." Once this process is complete, the reception process is terminated.

[0192] [Effects of the fourth embodiment] The above-described voice response system 100 is configured for use in workplaces or company reception areas. In this configuration, the name and company name of the person making the sales call are pre-recorded in the sales DB 118 on the server 90. When the person who comes to the reception desk states this name and company name, the system generates a response that plays a voice message declining their offer.

[0193] Furthermore, in the above voice response system 100, the server 90 will respond based on the input character information. It identifies the communication partner and connects the communication partner to a pre-configured communication destination for each communication partner. Such a voice response system 100 can assist with reception duties and telephone support. Furthermore, such a voice response system 100 can exclude individuals who may disrupt the user's work without the user having to deal with them directly.

[0194] Furthermore, in the voice response system 100 described above, the server 90 extracts keywords contained in the input text information (especially voice) and connects to the destination corresponding to the keyword. For example, keywords such as the name of the other party are pre-associated with their respective destinations.

[0195] Such a voice response system 100 can assist with tasks such as transferring phone calls and calling reception staff. [Modified version of the fourth embodiment] In the above embodiment, the connection destination was configured to be set according to the recipient. However, this technology could be applied to, for example, telephone reception for telephone banking or telephone shopping, where the system recognizes the requirements (keywords included in the text information) and changes the connection destination according to the requirements.

[0196] Furthermore, in the above-described voice response system 100, the server 90 may recognize the requirements of what the other party is saying based on keywords and convey a summary of what the other party has said to the user. Such a voice response system 100 can assist in the task of handling customer inquiries.

[0197] [Fifth Embodiment] [Processing of the fifth embodiment] Next, terminal device 1 may receive a request from another terminal device 1 and provide the information requested by that other terminal device 1.

[0198] In this configuration, during processing S56, server 90 requests the necessary information from other terminal devices 1, obtains the necessary information from other terminal devices 1, and then generates a response. The terminal device 1 that provides the necessary information then performs the information provision terminal processing shown in Figure 13. Information provision terminal processing is a process that is started, for example, when a request comes from server 90.

[0199] As shown in Figure 13, the information provision terminal processing first extracts the information recipient (S222). This information recipient indicates another terminal device 1 requesting information, and the ID for identifying this other terminal device 1 is included in the request from server 90.

[0200] Next, it is determined whether or not the recipient is authorized to receive information (S224). Here, the terminal information DB113 has in advance the IDs of recipients who are authorized to receive information, such as family and friends. This determination is made by referring to this terminal information DB113.

[0201] If the recipient is authorized to receive information (S224:YES), the device retrieves the requested information from its own memory 39 and various sensors (S226) and sends this data to the server 90 (S228). If the recipient is not authorized to receive information (S224:NO), the device sends a message to the server 90 refusing to provide the information (S230).

[0202] Once this process is complete, the information provision terminal process will terminate. In this configuration, for example, as shown in the fourth row of Figure 5, in response to the question "What is Mr. / Ms. XX doing?", the server 90 requests location information from Mr. / Ms. XX's terminal device 1, and this terminal device 1 returns the location information.

[0203] Then, server 90 recognizes the actions of Mr. ○○ based on the location information. For example, if moving on the line at a speed faster than the running speed of a human, it is determined that the person is moving while on a train, and a response such as "Mr. ○○ is on the train. It seems that he is on his way home." will be generated.

[0204] [Effect according to the fifth embodiment] In the voice response system 100, server 90 acquires information recorded in another terminal device 1 different from the terminal device 1 of the request source from another terminal device 1 and provides it to another terminal device 1. That is, in the voice response system 100, server 90 acquires information for generating a response to character information from another terminal device 1.

[0205] According to such a voice response system 100, a response can be generated based on the information recorded in another terminal device 1. Also, in the voice response system 100, when terminal device 1 is requested for information for generating a response to character information from another terminal device 1, it returns the information corresponding to this request.

[0206] In this configuration, terminal device 1 is equipped with sensors for detecting location information, temperature, humidity, illuminance, noise level, etc., and a database such as dictionary information, and extracts necessary information according to the request.

[0207] According to such a voice response system 100, information unique to another terminal device 1, such as the location of another terminal device 1, can be acquired. Also, its own unique information can be transmitted to another terminal device 1.

[0208] [Sixth embodiment] [Processing of the sixth embodiment] Next, in the voice response system of the sixth embodiment, a personality DB 106 in which personality information associating the personality of a user or a related person representing a person related to the user according to a preset category is recorded is prepared. The personality DB 106 is recorded, for example, as shown in FIG. 14, by associating the names of the user and related persons with the personality categories of these persons.

[0209] In addition, in the personality DB 106 shown in FIG. 14, a personality test is conducted on the user and related persons, and the test results are also recorded. Here, when generating personality information, well-known personality analysis techniques (such as the Rorschach test, the Szondi test, etc.) may be used. Also, when generating personality information, the techniques of aptitude tests used by companies and the like in employment tests may be used.

[0210] When generating personality information, for example, the personality information generation process shown in FIG. 15 is performed. The personality information generation process is, for example, a process that starts when an instruction to generate personality information is input using the operation unit 70 or the like in the terminal device 1.

[0211] In the personality information generation process, as shown in FIG. 15, first, the microphone 37 is turned on (S242), and one of the predetermined four-choice questions is output as audio (S244). At this time, for the four-choice questions, they may be obtained from the server 90, or the questions previously recorded in the memory 39 may be presented.

[0212] Subsequently, it is determined whether an audio response has been received from the subject (user or related person) (S246). If no response is received (S246: NO), the process of S246 is repeated. If a response is received (S246: YES), conversation parameters such as the ending of words and conversation speed are extracted (S248), and it is determined whether the current question is the last question (S250). If it is not the last question (S250: NO), the next question is selected (S252), and the process returns to S242.

[0213] If it is the last question (S250: YES), a personality analysis is performed based on the answers to the four-choice questions (S254), and a personality analysis is performed based on the conversation parameters (S256). Here, in the personality analysis based on the conversation parameters, it is possible to capture the tendency that people with confidence in themselves have a strong ending of words, while those without confidence have a weak ending of words, and the tendency that hasty people have a fast conversation speed and calm people have a slow conversation speed.

[0214] Next, these personality analysis results are comprehensively analyzed, including by weighting and averaging (S258), and then categorized into personality groups (S260). More specifically, the personality traits of the subjects obtained from the tests are scored, and then categorized into personality groups according to their scores.

[0215] Next, the subject and personality category are associated (S262) and recorded in the personality DB 106 (S264). In other words, the relationship between the subject and personality category is sent to the server 90. At this time, the test results are also sent to the server 90, and the server 90 constructs the personality DB 106 as shown in Figure 14. Once this process is complete, the personality information generation process is terminated.

[0216] When using the personality DB106 generated in this way, a corresponding set of responses different from the personality classification is prepared in the response candidate DB105. Then, in the S56 process, the server 90 obtains response candidates representing multiple different responses to the character information, selects a response to output from the response candidates according to the personality information, and outputs the selected response in the S60 and S64 processes.

[0217] [Effects of the 6th Embodiment] In the above-described voice response system 100, the terminal device 1 generates personality information of the user or related party based on the answers to a set number of questions, and acquires the generated personality information.

[0218] With such a voice response system 100, personality information can be generated in the server 90 and terminal device 1. Furthermore, in the above-mentioned voice response system 100, the calculation unit 101 generates personality information of the user or related party based on the string contained in the input character information.

[0219] With such a voice response system 100, the user can generate personality information during the process of using the voice response system 100. Furthermore, such a voice response system 100 can provide different responses depending on the personality of the user or those related to the user (stakeholders). Therefore, it can improve usability for the user.

[0220] [Modified version of the sixth embodiment] In the sixth embodiment described above, the response may be narrowed down to one depending on the personality before outputting, or different voice tones may be associated with each of the multiple responses and output.

[0221] Furthermore, the processing of S248, S254-S264 of the above-mentioned personality information generation process may be performed on the server 90. In this case, as in the first embodiment, the server 90 can be made to identify the terminal device 1, and voice and problems can be exchanged between the terminal device 1 and the server 90.

[0222] Furthermore, in the above-described voice response system 100, the server 90 may detect any of the user's actions and operations and generate learning information or personality information based on these.

[0223] According to such a voice response system 100, for example, if it detects that the user has been jumping onto a train for several consecutive days, it can prompt the user to leave home a few minutes earlier the following day, or if it detects from the conversation that the user has a tendency to get angry easily, it can output calming voices or music.

[0224] [Seventh Embodiment] [Processing of the 7th Embodiment] Next, in the voice response system of the seventh embodiment, a preference DB 108 is prepared, which records preference information that associates the preferences of users and related parties with pre-set categories. In the preference DB 108, for example, as shown in Figure 16, the names of users and related parties and their preferences are recorded, associated with each type of preference, such as food preferences (food), color preferences (color), hobbies, etc.

[0225] In particular, regarding food preferences, it is classified into sweet tooth (sweet), spicy tooth (spicy), and medium (neutral); regarding color preferences, it is classified into warm color system (warm), cool color system (cool), and medium (neutral); regarding hobbies, it is classified into indoor hobbies (indoor), outdoor hobbies (outdoor), and both indoor and outdoor hobbies (both indoor and outdoor).

[0226] When constructing such a preference DB 108, for example, the preference information generation process shown in FIG. 17 is executed. The preference information generation process is implemented, for example, between S48 and S54. Specifically, as shown in FIG. 17, keywords related to preferences are extracted from the character information (S282), and among the objects identified by image processing, those related to preferences are extracted (S284). Note that the keywords related to preferences are associated with the type of preference and the classification within that type (such as sweet, neutral, spicy for food preferences) in the preference DB 108. In these processes, when the extracted keywords or objects are included in the preference DB 108, they are extracted as those related to preferences.

[0227] Subsequently, the counter is incremented for each group of keywords related to preferences (S288). For example, when an object like kimchi is extracted, where the type of preference is "food preference" and the classification is "spicy", the counters corresponding to "food preference" and "spicy" are incremented.

[0228] Then, based on the counter values, the preference information (preference DB 108) is updated (S290). That is, for each "type of preference", the "classification" with the largest counter value is recorded in the preference DB 108 as the one that best matches the preferences of the user or relevant parties as the characteristics of their preferences. When such a process ends, the preference information generation process ends.

[0229] When using the preference DB108 generated in this way, a response candidate DB105 is prepared in which different responses are associated with each preference. In the S56 process, the server 90 obtains response candidates that represent multiple different responses to the text information, selects a response to output from the response candidates according to the preference information, and outputs the selected response in the S60 and S64 processes.

[0230] [Effects of the 7th Embodiment] In the above voice response system 100, the server 90 generates preference information indicating the user's or related party's preferences based on the string of characters contained in the text information. Then, based on the preference information, it selects a response to be output from the response candidates and outputs the selected response.

[0231] According to such a voice response system 100, responses are made according to the preferences of the user or related party. This allows for the following: For example, when a user is buying a gift for someone they know, they can ask terminal device 1, "What would so-and-so want?" and receive a response that matches their preferences.

[0232] [Variation of the 7th embodiment] In the response candidate DB105, a table that associates personality classifications with preference information may be included, as shown in Figure 18.

[0233] For example, in the example shown in Figure 18, personality categories are correlated with color preferences, and products that are estimated to be appreciated as gifts by women are arranged in a matrix. The S56 process can also generate responses by taking both personality and preferences into account.

[0234] [Eighth Embodiment] [Processing of the 8th embodiment] In the above embodiment, audio was converted into text information, but it is also possible to convert user actions into text information.

[0235] In detail, terminal device 1 captures the user's actions as an image and transmits it to server 90, where server 90 can, for example, perform the action character input process shown in Figure 19. The action character input process is initiated when a part of the user's body is captured in the image during processing S48.

[0236] In the motion character input process, as shown in Figure 19, first, an image is acquired (S302). Then, it is determined whether the user is trying to input characters by handwriting or by sign language (S304, S308).

[0237] In these processes, for example, if the captured image shows the user's upper body along with their face, it is determined that the user is attempting to input characters using sign language. If the captured image shows the user's hands but not their face, it is determined that the user is attempting to input characters using handwriting.

[0238] If the user is attempting to input characters by hand (S304: YES), the behavior of the fingertip or pen tip is recorded (S306), and this behavior is converted into character information (S312). Here, the handwritten character / sign language DB112 associates the behavior when writing characters by hand with the characters themselves, and also associates hand movements with characters expressed in sign language. In the process of S312, character information is generated by referring to the handwritten character / sign language DB112.

[0239] Furthermore, if the user is attempting to input text using sign language (S304: NO, S308: YES), the system will refer to the handwritten text / sign language DB112 to recognize the sign language content and perform the process described in S312. If the user is not attempting to input text using handwriting or sign language (S308: NO), the system will process input using another method (S314).

[0240] Next, the system associates the characters entered by action with the characters entered by voice and determines whether there is any similar voice (whether the degree of agreement between the reference waveform based on the character and the pronunciation waveform is above a certain threshold) (S316). If such voice input is found (S316: YES), the accent and pronunciation characteristics of the user when entering the character are recorded in the learning DB107 in association with the character (S318), and the action character input process is terminated.

[0241] Furthermore, if there is no such voice input (S316:NO), the character input process will terminate. ru. [Effects of the 8th Embodiment] In the above-described voice response system 100, user actions are converted into text information, allowing the user to input text information without speaking.

[0242] [Variation of the 8th embodiment] The operation of this embodiment can be caused not only by handwriting or gestures (e.g., sign language), but also by any movement of muscles.

[0243] [Ninth Embodiment] [Processing in the 9th embodiment] The contents of the learning DB 107 may be made available for use on a different terminal device 1 than the terminal device 1 that the user normally uses. In this case, the other terminal device 1 sends the ID and password of the normally used terminal device 1 to the server 90 along with the request for use.

[0244] Then, server 90 executes the other terminal usage process shown in Figure 20. The other terminal usage process is initiated when a usage request is received. In the process of using another terminal, as shown in Figure 20, it is first determined whether or not an ID and password have been entered (S332). If an ID and password have not been entered (S332: NO), the process in S332 is repeated.

[0245] Furthermore, if an ID and password have been entered (S332:YES), it is determined whether authentication using the ID and password has been completed (S334). If authentication is completed (S334:YES), a message indicating that authentication has been completed is sent to the other terminal device 1 (S336), and the other terminal device 1 is configured to use the learning DB 107 of terminal device 1 corresponding to the ID and password (S338).

[0246] If authentication is not completed (S334: NO), an error is sent to the other terminal device 1 (S340), and the process of using the other terminal is terminated. [Effects of the 9th Embodiment] Furthermore, in the above-mentioned voice response system 100, the server 90 transfers the learning information of one terminal device 1 to another terminal device 1.

[0247] With this voice response system 100, even when a user using one terminal device 1 uses another terminal device 1, they can utilize the learning information recorded on the terminal device 1 (learning information recorded on the server 90). Therefore, the accuracy of character information generation can be improved even when using other terminal devices 1. This is particularly effective when a user possesses multiple terminal devices 1.

[0248] Furthermore, in the above-mentioned voice response system 100, the server 90 outputs information about the user in response to inquiries from persons other than the user. With such a voice response system 100, for example, if the system detects the user's diet or walking distance, it can answer questions on behalf of the user at hospitals, etc. It may also be possible to learn the user's health status and self-introduction.

[0249] [Modified version of the 9th embodiment] Similar to the configuration of the ninth embodiment, upon receiving a request to terminate use and an ID and password, the system may terminate (prohibit) the use of the learning DB 107 for the terminal device 1 corresponding to the ID and password.

[0250] [Tenth Embodiment] [Processing of the 10th embodiment] In the voice response system of the tenth embodiment, the server 90 memorizes the conversation content and asks questions to obtain the same information it heard. Specifically, in S100 of the automatic conversation server processing shown in Figure 7, the memory confirmation process shown in Figure 21 is executed.

[0251] In the memory verification process, as shown in Figure 21, past conversation content is extracted from the learning DB 107 (S352), and a question is generated in which a keyword contained in one of the conversation contents is used as the answer (S353). Once this process is completed, the memory verification process is terminated.

[0252] For memory verification, you can ask questions such as, "What was on the menu for dinner yesterday?" or "Where did you go three days ago?" [Effects of the 10th Embodiment] Such a voice response system 100 allows for the confirmation of the user's memory and promotes memory retention. It is also considered effective in slowing the progression of dementia in the elderly.

[0253] [Embodiment No. 11] [Processing of the 11th Embodiment] Next, in the 11th embodiment of the voice response system, the terminal device 1 and the server 90 are configured to allow the user to practice a foreign language.

[0254] In detail, the pronunciation determination process 1 shown in Figure 22, the pronunciation determination process 2 shown in Figure 23, and the pronunciation determination process 3 shown in Figure 24 are executed in order. However, the server 90 executes one of the pronunciation determination processes 1 to 3 each time the voice response server process (Figure 2) is performed. In addition, each of the pronunciation determination processes 1 to 3 is executed as the process in S56 described above.

[0255] First, in pronunciation determination process 1, as shown in Figure 22, a response is generated instructing the user to input a predetermined sentence as voice (S362). In this process, for example, a model sentence in a foreign language is generated, and the user is prompted to imitate this sentence after the model. Once this process is complete, pronunciation determination process 1 is terminated.

[0256] Next, when audio is input in conjunction with pronunciation judgment process 1, pronunciation judgment process 2 is performed. In pronunciation judgment process 2, as shown in Figure 23, the accuracy of pronunciation and accent is scored (S372). In this process, the audio is treated as a waveform, and the degree of agreement between this waveform and the waveform of a model sentence is scored.

[0257] Then, this score is recorded in memory (S374), and pronunciation judgment process 2 is terminated. Next, pronunciation judgment process 3 is performed. In pronunciation judgment process 3, as shown in Figure 24, it is first determined whether the score is below a threshold (S382).

[0258] If the score is below the threshold (S382: YES), a response is generated instructing the user to enter the same sentence again (S384). This process may generate a response, for example, prompting the user to repeat the example sentence.

[0259] Furthermore, if the score is above a threshold (S382: NO), a response is generated acknowledging the good pronunciation and prompting the user to enter the next sentence (S386). For example, a response such as "Good pronunciation. Let's move on." is generated.

[0260] Once this process is complete, the pronunciation determination process 3 is terminated. [Effects of the 11th Embodiment] In the above voice response system 100, the server 90 detects the degree of accuracy of the pronunciation and accent of the voice input by the user and outputs the detected degree of accuracy.

[0261] Such a voice response system 100 allows for verification of the accuracy of pronunciation and accent. For example, it is effective when practicing a foreign language. Furthermore, in the above-mentioned voice response system 100, if the accuracy is below a certain value, the server 90 will output the same question again.

[0262] According to such a voice response system 100, an accurate answer can be obtained by outputting the same question. [Modified version of the 11th embodiment] In the above-described voice response system 100, the server 90 may output a voice message containing the word that is closest to the pronunciation made by the user, for verification purposes, if the accuracy is below a certain value.

[0263] With such a voice response system 100, the user can verify the accuracy of their pronunciation and accent. [Twelfth Embodiment] [Processing of the 12th embodiment] Next, the voice response system of the twelfth embodiment will be described. In the voice response system of the twelfth embodiment, the user's emotions are detected from the voice input by the user, and a response is generated to soothe the user according to those emotions.

[0264] In detail, the system performs the emotion determination process shown in Figure 25 and the emotion response generation process shown in Figure 26. The emotion determination process is performed as a detail of the process described in S50 above, and as shown in Figure 25, first, it scores emotions based on tone of voice, emphasis on the end of sentences, sentence length, conversation speed, unexpected words, etc. (S392), then it classifies emotions based on the score and records them in memory (S394).

[0265] Once this process is complete, the emotion determination process is terminated. Subsequently, the emotion response generation process is executed in the S56 process described above. In detail, as shown in Figure 26, first, the emotion category set in the emotion determination process is determined (S412). If the emotion category is normal (S412: normal), a normal greeting such as "hello" is generated as a response (message) (S414).

[0266] Furthermore, if the emotion is anger (S412: anger), the system generates a response such as "Did I offend you?" to calm the other person's emotions (S416). Additionally, if the emotion is joy (S412: joy), the system generates a response such as "It's a fun day, isn't it?" to have a brighter nuance compared to a normal greeting (S418).

[0267] Furthermore, if the emotion is confused (S412: confused), a greeting such as "Is something wrong?" is generated as a response to show concern for the other person (S420). Once this process is complete, the emotion response generation process ends.

[0268] [Effects of the 12th Embodiment] In the above voice response system 100, the server 90 detects the user's frustration or agitation by detecting unexpected sounds, and generates a message to suppress the frustration or agitation.

[0269] Such a voice response system 100 can suppress frustration and agitation in the user. Therefore, it can prevent the user from voicing conflicts with those around them.

[0270] [13th Embodiment] [Processing of the 13th Embodiment] Next, the voice response system of the 13th embodiment will be described. In the voice response system of the 13th embodiment, processing is performed to guide the user to an object in the captured image. This processing is performed in the server 90 as a detail of the processing in S56 described above.

[0271] When a voice command such as "Please guide me to the tower I can see" is input to terminal device 1, the guidance process is performed in S56. In the guidance process, as shown in Figure 27, terminal location information is first obtained from the GPS receiver 27 of terminal device 1 (S432).

[0272] Then, based on audio (text information) and image processing, the target object is identified from among the objects in the captured image, and its location is determined (S434). In this process, the location of the object is determined in map information (which may be obtained from an external source or held by the server 90) based on the shape of the object, its relative position, etc. For example, if a tower is visible in the captured image, the tower is identified on the map based on the location of terminal device 1 and the shape of the tower.

[0273] Next, the path to this object is searched (S436), and the path information is obtained (S438). This process can be implemented using the same process as that used in well-known cloud-based navigation devices.

[0274] Then, a response is generated to guide the user along the route (S440). In this process, the same response as the one provided by the navigation system should be generated. Once this process is complete, the guidance process will end. When the user is moving while receiving guidance, the automated conversation server can be used to play the message, with the user reaching the designated point as the playback condition.

[0275] [Effects of the 13th Embodiment] In the above-described voice response system 100, when text information is input, the server 90 generates a response corresponding to the captured image taken around the voice response system 100, and outputs this response as voice.

[0276] This voice response system 100 can output a response in voice according to the captured image. Therefore, it can improve usability compared to a configuration that generates a response from text information only.

[0277] Furthermore, in the above-mentioned voice response system 100, the server 90 searches for objects included in the text information from the captured image using image processing, identifies the location of the searched object, and guides the user to the location of the object.

[0278] Such a voice response system 100 can guide the user to an object in the captured image. Furthermore, in the above-mentioned voice response system 100, when the server 90 provides directions to a destination, it acquires route information such as weather, temperature, humidity, traffic information, and road surface conditions to the destination, and outputs the route information by voice.

[0279] Such a voice response system 100 can notify the user of the status (route information) on their way to their destination via voice. [Modified version of the 13th embodiment] In addition to the above configuration, it is also possible to input text information to respond with what has been recognized, and output a voice message indicating what (or who) has been recognized from the captured image.

[0280] Furthermore, in the above-described voice response system 100, the server 90 may, instead of processing in S48, acquire a video image capturing the shape of the user's mouth when inputting text information by voice. In this case, instead of processing in S52, the voice may be converted into text information, and the text information may be corrected by estimating unclear parts of the voice based on the video image.

[0281] With such a voice response system 100, the content of the speech can be estimated from the shape of the mouth, so unclear parts of the speech can be accurately estimated. [14th Embodiment] [Processing of the 14th Embodiment] Next, the voice response system of the 14th embodiment will be described. In the voice response system of the 14th embodiment, a predetermined action is requested from the user, and it is determined whether the user has performed the action as requested. In this configuration, in the automatic conversation terminal processing shown in Figure 6 and the automatic conversation server processing shown in Figure 7, the movement request processing 1 shown in Figure 28 and the movement request processing 2 shown in Figure 29 are performed in order as details of the S56 process described above.

[0282] First, once the process in S54 is completed, the movement request process 1 is started. In the movement request process 1, as shown in Figure 28, a response (message) is output instructing the user to move their gaze or head to a predetermined position (S452). Once this process is completed, the movement request process 1 is terminated.

[0283] Next, after the processing in S54 is completed, the movement request processing 2 is started. In movement request processing 2, as shown in Figure 29, it is determined whether the gaze and head position have moved as instructed (S462). In this process, the user's movements are detected by image processing of images captured by the camera and by using the detection results from various sensors of the terminal device 1. When detecting gaze by image processing, well-known gaze recognition techniques can be used.

[0284] If the gaze or head does not move as instructed (S462: NO), the response generated in S452 is output again (S464). If the gaze or head moves as instructed (S462: YES), another arbitrary response is generated (S466).

[0285] Once this process is complete, the move request process 2 is terminated. [Effects of the 14th Embodiment] In the above-described voice response system 100, the system detects the user's gaze and, if the user's gaze does not move to a predetermined position in response to a call, outputs a voice prompt requesting the user to move their gaze to the predetermined position.

[0286] Such a voice response system 100 allows the user to see a specific location. Therefore, safety checks during vehicle operation can be reliably performed. In the above-mentioned voice response system 100, the server 90 observes the position of body parts and facial expressions, and if there is little change in response to the call, it outputs a voice requesting that the position of body parts and facial expressions be changed.

[0287] Such a voice response system 100 can guide the user to move parts of their body to specific locations or to make specific facial expressions. The present invention can be used when driving a vehicle or during physical examinations.

[0288] [15th Embodiment] [Processing of the 15th Embodiment] Next, the voice response system of the 15th embodiment will be described. In the voice response system of the 15th embodiment, when a user inputs a broadcast program or music as voice, processing is performed to fill in the gaps if the broadcast program or music is interrupted.

[0289] In this configuration, as a detail of S56 mentioned above, the broadcast music completion process shown in Figure 30 is performed. In the broadcast music completion process, as shown in Figure 30, it is first determined whether the broadcast program or music (or the song if sung by the user) has been interrupted (S482).

[0290] If there is an interruption (S482:YES), the synchronized broadcast program or song is set as the response content in the process described in S492 (S484), and the broadcast song completion process is terminated. If there is no interruption (S482:NO), the broadcast program is acquired if a broadcast program is being viewed (S486), and the corresponding song is acquired if a song is being played (S488).

[0291] In this case, the karaoke database 116 stores songs and lyrics in association, and when retrieving a song in this process, the song with lyrics attached is retrieved. Next, the system identifies the broadcast program or music that the user is watching (S490). Then, it retrieves this broadcast program or music and prepares it for playback in sync with the broadcast program or music the user is watching (S492), and the broadcast music completion process ends.

[0292] [Effects of the 15th Embodiment] In the above voice response system 100, the server 90 acquires a broadcast program similar to the one the user is watching, and if the broadcast program is interrupted, it outputs the broadcast program it has acquired to fill in the gaps in the broadcast program.

[0293] Such a voice response system 100 can compensate for interruptions in the broadcast program that the user is watching. Furthermore, in the above-mentioned voice response system 100, when a user adds lyrics to a song without lyrics and sings it, the server 90 compares the song with lyrics with the lyrics added by the user and outputs the lyrics as voice only in the parts where only the user's lyrics are missing.

[0294] Such a voice response system 100 can compensate for parts of a song that a user of a karaoke machine cannot sing (parts of the lyrics that are cut off). [16th Embodiment] [Processing of the 16th Embodiment] Next, the voice response system of the 16th embodiment will be described. In the voice response system of the 16th embodiment, when characters are included in the captured image, the terminal device 1 receives a question from the user about how to read these characters, obtains information about these characters from an external source, and outputs the reading of the characters contained in this information as voice.

[0295] In this configuration, as a detail of S56 mentioned above, the character explanation process shown in Figure 31 is performed. In the character explanation process, as shown in Figure 31, first, it is determined whether or not a question about the reading has been received, for example, "How to read" (S502). If a question about the reading has been received (S502: YES), the reading of the image-recognized character is searched for from other servers connected via the Internet network 85 (S504), the obtained reading is set as the response (S506), and the character explanation process is terminated.

[0296] If it's not a question about pronunciation (S502:NO), then it's a "word" like what's written in a Japanese dictionary. The system determines whether or not a question about the meaning of a character (word) has been received (S508). If a question about the meaning has been received, the system searches for the meaning of the recognized character (word) from other servers connected via the Internet network 85 (S510), sets the obtained meaning as the response (S512), and terminates the character explanation process.

[0297] [Effects of the 16th Embodiment] According to this voice response system 100, the reading of the character recognized from the image is searched for on another server, etc., and the obtained reading is set as the response, so that the user can be taught how to read the character and the meaning of the word.

[0298] [17th Embodiment] [Processing of Embodiment 17] Next, the voice response system of the 17th embodiment will be described. In the voice response system of the 17th embodiment, the server 90 detects abnormal behavior or status of the user of the terminal device 1 based on sensor values ​​detected by the terminal device 1, and performs a process to send a notification if an abnormality is detected.

[0299] In detail, terminal device 1 performs the action response terminal processing shown in Figure 32, and server 90 performs the action response server processing. In the action response terminal processing, as shown in Figure 32, first, outputs from various sensors mounted on terminal device 1 are acquired (S522), and images captured by camera 41 are acquired (S524). Then, the acquired outputs from the various sensors and captured images are transmitted as packets to server 90 (S526), ​​and the action response terminal processing is terminated.

[0300] Next, in the behavior response server processing, as shown in Figure 33, the processes described in S42 to S44 are performed first. Subsequently, based on the location information of terminal device 1 (detection result by GPS receiver 27), behavior such as wandering is identified (S532), and the user's environment is detected based on the detection results from temperature sensors 15, 19, etc. (S534). Then, an abnormality is detected (S536).

[0301] This process detects anomalies based on changes in location information and the environment. For example, if the user does not move in a place with high or low temperatures, or if the user is in a place they do not normally go to, it is detected as an anomaly (S536). Alternatively, the location information and environment are scored, and if this score falls below a standard value (outside the standard range), it is determined to be an anomaly.

[0302] Next, it is determined whether or not an anomaly has been detected (S538). If no anomaly has been detected (S538: NO), the behavioral response server processing is terminated. If an anomaly has been detected (S538: YES), a message indicating that an anomaly has occurred is generated (S540), and a notification is sent to the designated contact (S542). Then, the processes described in S62 to S68 (excluding S66) are performed, and the behavioral response server processing is terminated.

[0303] [Effects of the 17th Embodiment] In the above voice response system 100, the server 90 detects the user's actions and the user's surrounding environment, and generates a message according to the detected actions and surrounding environment.

[0304] Such a voice response system 100 can notify users of dangerous locations or restricted areas. It can also detect unusual behavior by the user.

[0305] Furthermore, in the above voice response system 100, the server 90 captures an image of the user. Based on the image, the system determines the user's health status and generates a message accordingly. Such a voice response system 100 can manage the user's health status.

[0306] Furthermore, in the above-mentioned voice response system 100, the server 90 notifies a designated contact person if the health status falls below a standard value. According to this voice response system 100, an alert can be issued if the user's health condition falls below a certain threshold. Therefore, abnormalities can be notified to others at an earlier stage.

[0307] [Other embodiments] The embodiments of the present invention are not limited in any way to the embodiments described above, and various forms can be taken as long as they fall within the technical scope of the present invention.

[0308] For example, the voice response system 100 may be configured to mediate communication between two or multiple parties. More specifically, in cases where vehicles need to yield the right of way at an intersection, the terminal devices 1 may negotiate with each other to determine which vehicle will enter the intersection first. In this case, each terminal device 1 transmits information to the server 90 regarding its direction of movement and approach speed to the intersection. The server 90 then sets a priority for each terminal device 1 according to its direction of movement and approach speed, and generates and outputs voice messages such as "Stop" or "Entry Allowed" according to the priority.

[0309] Furthermore, when the terminal device 1 receives incoming calls for communications that require real-time responses, such as voice communications, it may be configured to accept calls only when it is convenient for the user. Specifically, the system may configure the system to accept calls when the user's face can be captured by the camera 41, considering this to be a convenient time for the user.

[0310] Furthermore, some people become annoyed when the other party does not respond when called during voice communication. To mitigate such feelings, it may be helpful to inform users waiting for a response about the other party's status. For example, terminal device 1 could manage the user's schedule, and if the user does not respond to an incoming call, it could search for what the user is doing or their available time slots and inform the user when they will be able to respond.

[0311] Furthermore, if the user does not respond to an incoming call, the caller may be informed of the user's location. For example, if the user is connected to the internet via a smartphone or computer, it is possible to determine which device is being used. This information could then be used to identify the user's location and inform the caller.

[0312] Furthermore, the system may use location information, such as GPS, to determine whether the user is able to respond to an incoming call. Based on location information, it is possible to determine whether the user is in a car, at home, etc. For example, if the user is traveling or in bed, it can be determined that they are unable to respond to the call because they are either in a public place or asleep. In cases where the user is unable to respond to an incoming call, it is possible to inform the caller of what the user is doing, as mentioned above.

[0313] Furthermore, a configuration utilizing security cameras can be considered to acquire location information. In recent years, security cameras have been installed in various locations, so it is possible to use these security cameras and configurations for identifying the user, such as facial recognition, to determine the user's location. It is also possible to use security cameras to determine what the user is doing (whether or not they are in a situation where they can answer the phone). In addition, whether or not they can answer an incoming call can be determined based on conditions such as whether they are using another landline (they cannot answer an incoming call while using the landline). Cut.

[0314] Furthermore, if the user of terminal device 1 wishes to converse with someone, the system may use the user's personality learning results to call up a terminal device from a large, unspecified group that is estimated to be a good match for the user. In such cases, the system may also initiate a conversation with the user about topics that are likely to be engaging (topics that both users are interested in, extracted using the learning results).

[0315] Furthermore, if the voice response device is not used for an extended period (i.e., the user has not spoken for a certain amount of time), the voice response device may be configured to say something to the user. In this case, you may use location information such as GPS to select the words to speak to the user.

[0316] [Relationship between each means described in the claims or means for solving the problem (the present invention) and the configuration in the embodiment] The terminal device 1 and server 90 in the above embodiment correspond to an example of the voice response device of the present invention. Furthermore, the processing in 22 and S56 in the above embodiment corresponds to an example of the response acquisition means of the present invention.

[0317] Furthermore, the processing in S28, S60, and S64 in the above embodiment corresponds to an example of the audio output means of the present invention. Also, the processing in S2 and S6 in the above embodiment corresponds to an example of the audio input means of the present invention.

[0318] Furthermore, the processing in S14 in the above embodiment corresponds to an example of the voice transmission means of the present invention. Also, the response candidate DB105 in the above embodiment corresponds to an example of the response recording means of the present invention.

[0319] Furthermore, the processing in S56 in the above embodiment corresponds to an example of the personality information acquisition means of the present invention. Also, the processing in S22 and S56 in the above embodiment corresponds to an example of the response acquisition means of the present invention.

[0320] Furthermore, the processing in S28, S60, and S64 in the above embodiment corresponds to an example of the audio output means of the present invention. Also, the processing in S254, S258, and S260 in the above embodiment corresponds to an example of the first personality information generation means and the second personality information generation means of the present invention. Also, the processing in S56 in the above embodiment corresponds to an example of the personality information acquisition means of the present invention.

[0321] Furthermore, the processing in S22 and S56 in the above embodiment corresponds to the response acquisition means of the present invention. Also, the processing in S28, S60, and S64 in the above embodiment corresponds to an example of the audio output means of the present invention.

[0322] Furthermore, the processing in S254, S258, and S260 in the above embodiment corresponds to an example of the first personality information generation means and the second personality information generation means of the present invention. Furthermore, the processing in S48 and S56 in the above embodiment corresponds to an example of the response generation means of the present invention. Also, the processing in S28, S60, and S64 in the above embodiment corresponds to an example of the audio output means of the present invention.

[0323] Furthermore, a modified example in the above embodiment: the processing in S48 corresponds to an example of the voice input video acquisition means of the present invention. Also, the processing in S52 in the above embodiment corresponds to an example of the character information conversion means of the present invention.

[0324] Furthermore, the preference information generation process in the above embodiment is an example of the preference information generation means of the present invention. It corresponds to this. Furthermore, the processing in S56 in the above embodiment corresponds to an example of the response candidate acquisition means of the present invention.

[0325] Furthermore, the character input processing in the above embodiment corresponds to an example of the character information generation means of the present invention. Also, the other device information acquisition means, the other terminal utilization processing in the above embodiment corresponds to an example of the transfer means of the present invention.

[0326] Furthermore, the processing in S98 in the above embodiment corresponds to an example of the playback condition determination means of the present invention. Also, the processing in S100 in the above embodiment corresponds to an example of the message playback means of the present invention.

[0327] Furthermore, the processing in S116 in the above embodiment corresponds to an example of the non-response transmission means of the present invention. Also, the processing in S372 in the above embodiment corresponds to an example of the speech accuracy detection means of the present invention.

[0328] Furthermore, the processing in S374 in the above embodiment corresponds to an example of the accuracy output means of the present invention. Also, the processing in S204 in the above embodiment corresponds to an example of the connection control means of the present invention.

[0329] Furthermore, the processing in S50 in the above embodiment corresponds to an example of the emotion determination means of the present invention. Also, the processing in S438 in the above embodiment corresponds to an example of the route information acquisition means of the present invention.

[0330] Furthermore, the processing in S462 in the above embodiment corresponds to an example of the gaze detection means of the present invention. Also, the processing in S464 in the above embodiment corresponds to an example of the gaze movement request transmission means of the present invention.

[0331] Furthermore, the processing in S464 in the above embodiment corresponds to an example of the change request transmission means of the present invention. Also, the processing in S486 in the above embodiment corresponds to an example of the broadcast program acquisition means of the present invention.

[0332] Furthermore, the processing in S484 in the above embodiment corresponds to an example of the broadcast program supplementation means and lyric addition means of the present invention. Also, the processing in S504 and S506 in the above embodiment corresponds to an example of the reading output means of the present invention. Also, the processing in S522 and S524 in the above embodiment corresponds to an example of the behavioral environment detection means of the present invention.

[0333] Furthermore, the process in S538 in the above embodiment corresponds to an example of the health status determination means of the present invention. Also, the process in S540 in the above embodiment corresponds to an example of the health message generation means of the present invention.

[0334] Furthermore, the processing in S542 in the above embodiment corresponds to an example of the notification means of the present invention. [Explanation of symbols]

[0335] 1...Terminal device, 10...Motion sensor unit, 11...Dimensional acceleration sensor, 13...Axis gyro sensor, 15...Temperature sensor, 17...Humidity sensor, 19...Temperature sensor, 21...Humidity sensor, 23...Illuminance sensor, 25...Wet sensor, 27...GPS receiver, 29...Wind speed sensor, 33...Electrocardiogram sensor, 35...Heart sound sensor, 37...Microphone, 39...Memory, 41...Camera, 50...Communication unit, 53...Wireless telephone unit, 55...Contact memory, 60...Notification unit, 61...Display, 63...Illumination, 65...Speaker, 70...Operation unit, 71...Touchpad, 73...Confirmation button, 75...Fingerprint sensor, 77...Rescue request lever, 80...Communication base station, 85...Internet Network, 90...Server, 100...Voice response system, 101...Calculation unit, 102...Voice recognition DB, 103...Predictive text DB, 104...Voice DB, 105...Response candidate DB, 106...Personality DB, 107...Learning DB, 108...Preference DB, 109...News DB, 110...Weather DB, 111...Playback conditions DB, 112...Handwriting / sign language DB, 113...Terminal information DB, 114...Emotion judgment DB, 115...Health judgment DB, 116...Karaoke DB, 117...Report destination DB, 118...Sales DB, 119...Client DB.

Claims

[Claim 1] A voice response system comprising a requesting device that initiates a request, a providing device that provides information, and a server capable of communicating with the requesting device and the providing device, The aforementioned provider device is configured to allow setting whether or not to permit the provision of information. The server is configured to receive a request from the requesting device, and if the request includes a request specifying the providing device, to send the request to the specified providing device, and if the provision of information is permitted, to receive the provided information from the providing device. A providing unit is configured to generate an audio response based on the provided information as a response to a request from the requesting device, and to provide the response to the requesting device. A voice response system equipped with the following features.