Voice response system, server for voice response apparatus, terminal, and car
The voice response system addresses the limitation of single-answer systems by allowing multiple tone-based responses and personalized interaction, improving user experience and efficiency.
Patent Information
- Application Number
- JP2025176730
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-06-18
- Filing Date
- 2025-10-20
- Publication Date
- 2026-01-21
- Estimated Expiration
- 2033-05-29
AI Technical Summary
Existing voice response systems provide only one predefined answer to a question, lacking user-friendliness and flexibility in response generation.
A voice response system that includes a requesting device, a providing device, and a server, allowing for multiple responses to be generated and output in different tones of voice, with the ability to set information provision permissions, and capable of converting voice input to text and generating responses based on personality and preference information.
Enables multiple responses in different tones, enhancing user understanding and usability, reducing processing load, and providing personalized responses based on user characteristics and preferences.
Smart Images

Figure 2026010141000001_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This international application is a continuation of Japanese Patent Application No. 2 filed with the Japan Patent Office on June 18, 2012. This international application claims priority to Japanese Patent Application Nos. 2012-137065, 2012-137066, and 2012-137067, the entire contents of which are incorporated herein by reference. [Technical Field]
[0002] The present invention relates to a voice response system that allows responses to be made by voice. [Background technology]
[0003] Known voice response devices include those that search for an answer to an input question from a dictionary and output the searched answer by voice (see, for example, Patent Document 1). Also known is a technology that generates an answer to a question based on the content of a dialogue with a user (see, for example, Patent Document 2). [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 4832097 [Patent Document 2] Patent No. 4924950 Summary of the Invention [Problem to be solved by the invention]
[0005] The above technology is simply set to provide one answer specified by a dictionary to one question. One aspect of the present invention is to provide a voice response system that is more user-friendly for users. [Means for solving the problem]
[0006] In the first aspect of the invention, A voice response system including a requesting device that is a request source, a providing device that is an information providing source, and a server that can communicate with the requesting device and the providing device, the providing device is configured to be able to set whether or not to permit information provision; an information acquisition unit configured in the server to acquire a request from the request source device, and if the request includes a request specifying the providing source device, to transmit the request to the specified providing source device, and if the information provision is permitted, to acquire the provided information provided by the providing source device; a providing unit configured to generate a voice response based on the provided information as a response to a request from the requesting device and provide the response to the requesting device; Equipped with. In another aspect of the invention, A voice response device that responds to input text information by voice, a response acquisition means for acquiring a plurality of different responses to the character information; a voice output means for outputting the plurality of different responses in different voice tones; The present invention is characterized by the following features.
[0007] According to such a voice response device, multiple responses can be output in different tones of voice, so even if a single answer cannot be identified for one piece of text information, different answers can be output in different tones of voice in a way that is easy for the user to understand, thereby making it easier for the user to use.
[0008] The voice response device of the present invention may be configured as, for example, a terminal device carried by a user, or as a server that communicates with this terminal device. Furthermore, text information may be input using an input means such as a keyboard, or may be input by converting voice into text information.
[0009] In the above voice response device, as in the second aspect of the invention, A voice input means for a user to input voice, and a means for converting the input voice into text information; a voice transmission means for transmitting to an external device a plurality of different responses to the character information and transmitting the responses to the voice response device; Equipped with The response acquisition means acquires the response from the external device. This may be done.
[0010] According to such a voice response device, since voice can be input into the voice response device, text information can be input by voice. Also, since a response can be generated in an external device, the processing load on the voice response device can be reduced.
[0011] In the voice transmission means, the operation of "converting input voice into text information" may be performed by the voice response unit or by an external device. Furthermore, in the voice response device, as in the third aspect of the invention, The voice response unit or the external device is provided with a response recording means in which a plurality of different responses including positive and negative responses to each of a plurality of pieces of text information are recorded, the response acquisition means acquires the positive response and the negative response as the plurality of different responses; The audio output means may reproduce the positive response and the negative response in different tones of voice.
[0012] Such a voice response device can reproduce responses with different positions, such as positive and negative responses, in different tones of voice, making it possible to reproduce the voice as if it were being spoken by a different person, thereby making it less likely that the user hearing the voice will feel uncomfortable.
[0013] The tone of voice may be changed depending on the type of response or the language used when responding. For example, if the response is to be made in a gentle tone, a calm female voice may be used, and if the response is to be made in a fierce tone, a brave male voice may be used. In other words, the content of the response may be associated with a personality, and the tone of voice may be set according to the personality.
[0014] Furthermore, the voice response device can be configured to be used at the reception desk of a workplace or company, as in the fourth aspect of the invention, or to convey something that the user finds difficult to say directly to someone on their behalf.
[0015] When using a voice response device at the reception, the name of the salesperson and the name of the company are recorded in advance on the voice response device or an external device, and the person who comes to the reception announces this name and company name. If so, a response can be generated to play a voice refusing the call.
[0016] Furthermore, if you want to configure the device to communicate something that is difficult to say, for example, before a date, you can talk to the device about what you would like to say today, and the voice response device will speak for you (play back audio) at an appropriate time (for example, at a preset time or after a certain amount of time has passed since the conversation stopped).
[0017] Alternatively, the voice may be configured to speak words that trigger something difficult to say, such as, "By the way, didn't you say you were going to tell her something?" In other words, instead of outputting a response immediately, the voice may be configured to output a response when a playback condition is met, such as after a certain period of time has passed.
[0018] Furthermore, in the above voice response device, as in a fifth aspect of the invention, the external device or the voice response device may acquire information for generating a response to text information from another voice response device. Also, as in a sixth aspect of the invention, when the voice response device is requested by another voice response device for information for generating a response to text information, the voice response device may return information in response to this request.
[0019] In this case, the voice response unit may be provided with sensors for detecting location information, temperature, humidity, illuminance, noise level, etc., and a database of dictionary information, etc., so that the necessary information can be extracted in response to a request.
[0020] Such a voice response unit (external device) can acquire information for generating a response from another voice response unit, such as the location of the other voice response unit, and other information specific to the other voice response unit.
[0021] Furthermore, it is possible to transmit its own unique information to other voice response units. Furthermore, in the above voice response device, as in the seventh aspect of the invention, a response (e.g., an affirmative response or a negative response) output by the voice response device itself or another voice response device may be input as text information, and a response for refuting this response may be generated. In other words, from the user's perspective, it is possible to hear arguments from both pro and con positions. Then, after listening to this argument, the user can make a final decision.
[0022] This configuration can be realized using one or more voice response units. In this case, the voices can be exchanged between the voice response units by direct input and output, or by wireless communication or the like.
[0023] In addition, in the eighth aspect of the invention, A voice response device that responds to input text information by voice, a personality information acquiring means for acquiring personality information in which the personality of a user or a person related to the user is associated with the user in accordance with a predetermined classification; a response acquisition means for acquiring response candidates representing a plurality of different responses to the text information; a voice output means for selecting a response to be output from among response candidates in accordance with the personality information and outputting the selected response; The present invention is characterized by the following features.
[0024] Such a voice response device can provide different responses depending on the characteristics of the user and those related to the user (stakeholders), thereby improving usability for the user.
[0025] In addition, in the above voice response device, as in the invention of a ninth aspect, a first personality information generating means for generating personality information of the user or the related person based on answers to a plurality of preset questions; The personality information acquisition means acquires the personality information generated by the personality information generation means. This may be done.
[0026] Such a voice response device can generate personality information within the voice response device. Well-known personality analysis techniques (such as the Rorschach test and the Szondi test) can be used to generate the personality information. Furthermore, aptitude test techniques used by companies and other organizations for employment examinations can also be used to generate the personality information.
[0027] Furthermore, in the voice response device, as in a tenth aspect of the invention, a second personality information generating means for generating personality information of the user or the related person based on a character string included in the input character information; The personality information acquisition means acquires the personality information generated by the personality information generation means. This may be done.
[0028] According to such a voice response device, personality information can be generated while the user is using the voice response device. In addition, in the above voice response device, as in an eleventh aspect of the invention, preference information generating means for generating preference information indicating the tendency of preferences of the user or the related person based on a character string included in the character information; The voice output means selects a response to be output from the response candidates based on the preference information, and outputs the selected response. This may be done.
[0029] Such a voice response device can respond according to the preferences of the user or the person concerned. Furthermore, in the above voice response device, as in the invention of the 12th aspect, the user's behavior (conversations, places traveled, and what is captured on camera) may be learned (recorded and analyzed) in advance to compensate for any shortcomings in the user's conversation.
[0030] For example, in a conversation in which the user responds to the question "Is hamburger steak okay today?" with "Curry would be good," the device can add, "Because we had hamburger steak yesterday," thereby conveying the reason why the user said curry would be good.
[0031] Such configuration can also be implemented during a phone call and may be configured to participate in the user's conversation without permission. Furthermore, in the above voice response device, as in the invention of a thirteenth aspect, a response candidate acquisition means for acquiring response candidates from a predetermined server or the Internet; The device may also include:
[0032] Such a voice response device can obtain candidate responses not only from its own device or an external device, but also from any device connected via the Internet, a dedicated line, or the like. In addition, in the above voice response device, as in the invention of a fourteenth aspect, a character information generating means for converting a user's action into character information; The device may also include:
[0033] Here, the movement referred to in the present invention corresponds to conversation, handwriting, gestures (for example, sign language) and other movements resulting from muscle movements. Such a voice response device can convert the user's actions into text information.
[0034] Furthermore, in the above voice response device, as in the invention of a fifteenth aspect, The text information generating means converts the user's speech into text information and accumulates speech habits (pronunciation habits, etc.) as learning information (capturing and recording the characteristics). This may be done.
[0035] According to such a voice response device, since it is possible to generate text information based on learning information, it is possible to improve the accuracy of generating text information. In addition, in the above voice response device, as in the invention of a sixteenth aspect, a transfer means for transferring the learning information to another voice response unit; The device may also include:
[0036] According to this voice response device, even when a user uses another voice response device, the learning information recorded by this voice response device can be used, thereby improving the accuracy of generating text information even when using another voice response device.
[0037] Furthermore, in the voice response device, as in a seventeenth aspect of the invention, either the behavior or the operation of the user may be detected, and learning information or personality information may be generated based on the detected behavior or the operation.
[0038] With such a voice response device, for example, if it detects that the user has been jumping on trains for several days in a row, it can urge the user to leave home a few minutes earlier from the next day, or if it detects from conversation that the user has a tendency to get angry easily, it can output voice or music to calm the user down.
[0039] In the voice response device, as in an eighteenth aspect of the invention, The voice response unit may further include an other device information acquisition unit that acquires information recorded in another voice response unit from the other voice response unit.
[0040] Such a voice response unit can generate a response based on information recorded in another voice response unit. Furthermore, in the above voice response device, as in the invention of a nineteenth aspect, a reproduction condition determination means for determining whether or not the state of the voice response unit matches a reproduction condition set in advance as a condition for outputting a voice when the character information is not input; a message reproducing means for outputting a preset message when the reproducing condition is met; The device may also include:
[0041] Such a voice response device can output voice even when no text information is input (i.e., when the user does not speak). For example, by forcing the user to speak, it can be used as a measure to suppress drowsiness while driving a car. In addition, by determining whether a person living alone responds, it is possible to check the safety of the person.
[0042] In addition, in the above voice response device, as in the invention of the twentieth aspect, The message reproducing means acquires news information and outputs a message relating to the news in the form of a question requesting a response from the user. This may be done.
[0043] Such a voice response device allows conversations about the news, For example, if information about a company's stock price is obtained, the conversation can be something like, "The stock price of company X went up by XX yen today. Did you know?"
[0044] Furthermore, in the above voice response device, as in the invention of a twenty-first aspect, The voice output means or message playback means outputs a preset message by adding separately acquired external information (such as news or environmental information (temperature, weather, location information, etc.)). This may be done.
[0045] Such a voice response unit can output a response that combines a predetermined message with the acquired information. In addition, in the above voice response device, as in the invention of the twenty-second aspect, Get multiple messages and select and output the message to be played depending on the message playback frequency. This may be done.
[0046] With such a voice response device, it is possible to make it difficult to play messages that are played frequently, thereby creating a sense of randomness when playing messages, or to repeatedly play messages that are played frequently to draw attention or help solidify memories.
[0047] Furthermore, in the above voice response device, as in the invention of a twenty-third aspect, a no-answer transmission means for transmitting information identifying the user and a message indicating that no answer has been received to a pre-set contact point when no response or reply to the message is received; The device may also include:
[0048] Such a voice response device can notify a contact person when no response is received, thereby enabling early notification of an abnormality in, for example, an elderly person living alone. In addition, in the above voice response device, as in the invention of a twenty-fourth aspect, The message playback means memorizes the conversation contents and asks questions to obtain the same contents as what was heard (memory confirmation process). This may be done.
[0049] Such a voice response device can check the memory ability of the user and also help the user to solidify the memory. Furthermore, in the above voice response device, as in the invention of a 25th aspect, a speech accuracy detection means for detecting the accuracy of pronunciation and accent of the voice input by the user; an accuracy degree output means for outputting the detected accuracy degree; The device may also include:
[0050] Such a voice response device makes it possible to check the accuracy of pronunciation and accent, which is useful when practicing a foreign language, for example. In addition, in the above voice response device, as in the invention of a 26th aspect, The accuracy output means outputs a voice containing the closest word when the accuracy is equal to or less than a certain value. This may be done.
[0051] Such a voice response device allows the user to check the accuracy of pronunciation and accent. Furthermore, in the above voice response device, as in the invention of a 27th aspect, The message reproducing means may be configured to output the same question again if the degree of accuracy is equal to or less than a certain value.
[0052] Such a voice response device can obtain an accurate answer by outputting the same question. In addition, in the above voice response device, as in the invention of a 28th aspect, a connection control means for identifying a communication partner based on input character information and connecting the communication partner to a communication destination preset for each communication partner; The device may also include:
[0053] Such a voice response device can assist with reception work and telephone response. In particular, in the above voice response device, as in the invention of the 29th aspect, The connection control means distinguishes between sales activities and visitors, and plays a message to decline if the call is for sales activities. This may be done.
[0054] Such a voice response device allows a user to exclude a person who may be causing a disruption to the user's work without the user having to deal with the person. Furthermore, in the voice response device, as in the invention of the thirtieth aspect, a keyword contained in the input character information (particularly voice) may be extracted and a connection may be made to a connection destination corresponding to the keyword. Note that a keyword such as the name of a destination and the connection destination may be associated in advance.
[0055] Such a voice response device can assist with tasks such as transferring calls and calling the reception desk. Furthermore, in the voice response device, as in the invention of the thirty-first aspect, the requirements of what the other party is saying may be recognized based on keywords, and an outline of what the other party has said may be conveyed to the user.
[0056] Such a voice response device can assist in the intermediary work with customers. Furthermore, in the above voice response device, as in the invention of a thirty-second aspect, An emotion determination means for reading emotions from the tone of voice input by the user and outputting which emotion corresponds to the input from among emotions including at least one of normal, anger, joy, confusion, sadness, and elation. The device may also include:
[0057] Such a voice response device can output a response according to the user's emotions. Next, the invention of the 33rd aspect is as follows: a response generating means for generating a response according to an image captured by capturing an image of the surroundings of the voice response device when the character information is input; a voice output means for outputting the response by voice; The present invention is characterized by the following features.
[0058] Such a voice response device can output a voice response in response to a captured image, which improves usability compared to a configuration in which a response is generated from text information only.
[0059] A specific configuration of the present invention is, for example, to input text information to respond as to what has been recognized, and output a voice message indicating what (someone) has been recognized from the captured image.
[0060] In the above voice response device, as in the invention of the 34th aspect, a position specifying means for searching for an object included in the character information from the captured image by image processing and for specifying the position of the object that has been searched for; a guide means for guiding the object to the position; The device may also include:
[0061] Such a voice response device can guide the user to an object in a captured image. Furthermore, in the above voice response device, as in the invention of a 35th aspect, A sound system that captures a moving image of the user's mouth shape when inputting text information by voice. A voice input video acquisition means; a character information conversion means for converting the voice into character information and correcting the character information by estimating unclear parts of the voice based on the moving image; The device may also include:
[0062] According to such a voice response device, the content of the speech can be estimated from the shape of the mouth, so that unclear parts of the speech can be estimated well. In addition, in the above voice response device, as in the invention of a 36th aspect, The message reproducing means detects the user's irritation or agitation by detecting unexpected sounds, and generates a message to suppress the user's irritation or agitation. This may be done.
[0063] Such a voice response device can suppress irritation or agitation of the user when the user becomes irritated or agitated, thereby preventing trouble between the user and those around him or her. Furthermore, in the above voice response device, as in the invention of a 37th aspect, a route information acquisition means for acquiring route information such as weather, temperature, humidity, traffic information, and road surface conditions to the destination when providing guidance to the destination; The message reproducing means outputs the route information by voice. This may be done.
[0064] Such a voice response device can notify the user of the status (route information) to the destination by voice. In addition, in the above voice response device, as in the invention of a 38th aspect, A gaze detection means for detecting a user's gaze; a gaze movement request transmitting means for outputting a voice requesting the user to move their gaze to a predetermined position when the user does not move their gaze to a predetermined position in response to the call from the message reproducing means; The device may also include:
[0065] Such a voice response device allows the user to view a specific location, thereby enabling the user to reliably check for safety when driving a vehicle. In the above voice response device, as in the invention of the 39th aspect, a change request transmitting means for observing the position of a body part or a facial expression, and outputting a voice requesting a change in the position of a body part or a facial expression if there is little change in response to the call; The device may also include:
[0066] The voice response device can move the user's body part to a specific position, induce the user to make a specific facial expression, etc. The present invention can be used when driving a vehicle or during a physical examination.
[0067] Furthermore, in the voice response device, as in the invention of a fortieth aspect, a broadcast program acquisition means for acquiring a broadcast program similar to the broadcast program viewed by the user; a broadcast program supplementing means for, when a broadcast program is interrupted, outputting the broadcast program acquired by itself to supplement the interrupted broadcast program; The device may also include:
[0068] Such a voice response device can compensate for the interruption of a broadcast program being viewed by a user. In addition, in the above voice response device, as in the invention of the forty-first aspect, The device is provided with a lyrics adding means for comparing a song with lyrics with the lyrics added by the user when the user sings the song without lyrics and outputting the lyrics by voice in the part where only the lyrics of the user are missing.
[0069] Such a voice response device can compensate for the part that the user cannot sing in so-called karaoke (the part where the lyrics are broken). Furthermore, in the above voice response device, as in the invention of a forty-second aspect, a pronunciation output means for externally acquiring information on characters when the captured image contains characters and a user asks how to read the characters, and for outputting the pronunciation of the characters contained in the information by voice; The device may also include:
[0070] Such a voice response device can teach the user how to read the characters. In addition, in the above voice response device, as in the invention of the 43rd aspect, An action environment detection means is provided to detect the action of the user and the surrounding environment of the user, The message generating means generates a message according to the detected behavior and the surrounding environment. This may be done.
[0071] Such a voice response device can warn of dangerous places, restricted areas, etc. It can also detect abnormal behavior of the user. Furthermore, in the voice response device, as in the invention of a 44th aspect, a health condition determining means for determining a health condition of a user based on a captured image of the user; a health message generating means for generating a message according to a health condition; The device may also include:
[0072] Such a voice response device makes it possible to manage the health condition of the user. In addition, in the above voice response device, as in the invention of the 45th aspect, A notification means for notifying a predetermined contact when the health condition falls below a standard value; The device may also include:
[0073] Such a voice response device can issue a notification when the user's health condition is below a reference value, thereby making it possible to notify others of an abnormality at an earlier stage. Furthermore, in the voice response unit, as in a forty-sixth aspect of the invention, information about the user may be output in response to an inquiry from a person other than the user.
[0074] Such a voice response device can, for example, detect the user's dietary habits and walking distances, and then answer questions on behalf of the user at a hospital, etc. It may also be configured to learn the user's health condition and self-introduction.
[0075] It should be noted that each aspect of the invention does not need to be premised on other inventions, and should be treated as an independent invention as much as possible. It is possible. [Brief explanation of the drawings]
[0076] [Figure 1] 1 is a block diagram showing a schematic configuration of a voice response system to which the present invention is applied; [Figure 2] FIG. 2 is a block diagram showing a schematic configuration of a terminal device. [Figure 3] 10 is a flowchart showing a voice response terminal process executed by an MPU of the terminal device. [Figure 4] 10 is a flowchart showing a voice response server process executed by a calculation unit of the server. [Figure 5] FIG. 10 is an explanatory diagram illustrating an example of a response candidate DB. [Figure 6] 10 is a flowchart showing an automatic conversation terminal process executed by an MPU of a terminal device. [Figure 7] 10 is a flowchart showing an automatic conversation server process executed by a calculation unit of the server. [Figure 8] 10 is a flowchart showing a message terminal process executed by the MPU of the terminal device. [Figure 9] 10 is a flowchart showing a message server process executed by a calculation unit of the server. [Figure 10] 10 is a flowchart showing a guidance terminal process executed by the MPU of the terminal device. [Figure 11] 10 is a flowchart showing a guidance server process executed by a calculation unit of the server. [Figure 12] 10 is a flowchart showing a reception process executed by a calculation unit of the server. [Figure 13] 10 is a flowchart showing an information providing terminal process executed by an MPU of the terminal device. [Figure 14] FIG. 10 is an explanatory diagram showing an example of a personality DB. [Figure 15] 10 is a flowchart showing a personality information generation process executed by an MPU of a terminal device. [Figure 16] FIG. 2 is an explanatory diagram illustrating an example of a preference DB. [Figure 17] 10 is a flowchart showing a preference information generating process executed by a calculation unit of the server. [Figure 18] FIG. 10 is an explanatory diagram showing an example of a combination of personality categories and preferences. [Figure 19] 10 is a flowchart showing a motion character input process executed by a calculation unit of the server. [Figure 20] 10 is a flowchart showing another terminal use process executed by a calculation unit of the server. [Figure 21] 10 is a flowchart showing a storage confirmation process executed by a calculation unit of the server. [Figure 22] 10 is a flowchart showing pronunciation determination processing 1 executed by a calculation unit of the server. [Figure 23] 10 is a flowchart showing pronunciation determination processing 2 executed by the calculation unit of the server. [Figure 24] 10 is a flowchart showing pronunciation determination processing 3 executed by the calculation unit of the server. [Figure 25] 10 is a flowchart showing emotion determination processing executed by a calculation unit of the server. [Figure 26] 10 is a flowchart showing an emotional response generation process executed by a calculation unit of the server. [Figure 27] 10 is a flowchart showing a guidance process executed by a calculation unit of the server. [Figure 28] 10 is a flowchart showing a movement request process 1 executed by a calculation unit of the server. [Figure 29] 10 is a flowchart showing a movement request process 2 executed by a calculation unit of the server. [Figure 30] 10 is a flowchart showing a broadcast music complementing process executed by a calculation unit of the server. [Figure 31] 10 is a flowchart showing a character explanation process executed by a calculation unit of the server. [Figure 32] 10 is a flowchart showing an action response terminal process executed by a calculation unit of the server. [Figure 33] 10 is a flowchart showing an action response server process executed by a calculation unit of the server. DETAILED DESCRIPTION OF THE INVENTION
[0077] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. [First embodiment] [Configuration of this embodiment] A voice response system 100 to which the present invention is applied is a system configured to generate an appropriate response in a server 90 in response to a voice input in a terminal device 1, and to output the response by voice in the terminal device 1. In detail, as shown in Fig. 1, the voice response system 100 is configured so that a plurality of terminal devices 1 and the server 90 can communicate with each other via a communication base station 80 and an Internet network 85.
[0078] The server 90 has the functions of a normal server device. In particular, the server 90 has a calculation unit 101 and various databases (DB). The calculation unit 101 is configured as a well-known calculation unit having a CPU and memories such as a ROM and a RAM, and performs various processes based on programs stored in the memory, such as communication with the terminal device 1 via the Internet 85, reading and writing data in various DBs, and speech recognition and response generation for conversation with the user of the terminal device 1.
[0079] 1, the various DBs include a speech recognition DB 102, a predictive conversion DB 103, a speech DB 104, a response candidate DB 105, a personality DB 106, a learning DB 107, a preference DB 108, a news DB 109, a weather DB 110, a playback condition DB 111, a handwritten character / sign language DB 112, a terminal information DB 113, an emotion determination DB 114, a health determination DB 115, a karaoke DB 116, a report destination DB 117, a sales DB 118, and a client DB 119. Details of these DBs will be described in each explanation of the processing.
[0080] Next, as shown in FIG. 2, the terminal device 1 is configured by providing the behavior sensor unit 10, a communication unit 50, a notification unit 60, and an operation unit 70 in a predetermined housing. The behavior sensor unit 10 is equipped with a well-known MPU 31 (microprocessor unit), memory 39 such as ROM and RAM, and various sensors, and the MPU 31 performs processing such as driving a heater to optimize the temperature of the sensor element so that the sensor elements that make up the various sensors can properly detect the object being tested (humidity, wind speed, etc.).
[0081] The behavior sensor unit 10 is equipped with various sensors, including a three-dimensional acceleration sensor 11 (3DG sensor), a three-axis gyro sensor 13, a temperature sensor 15 arranged on the back of the housing, a humidity sensor 17 arranged on the back of the housing, a temperature sensor 19 arranged on the front of the housing, a humidity sensor 21 arranged on the front of the housing, an illuminance sensor 23 arranged on the front of the housing, a wetness sensor 25 arranged on the back of the housing, a GPS receiver 27 that detects the current location of the terminal device 1, and a wind speed sensor 29.
[0082] The behavior sensor unit 10 also includes various sensors, such as an electrocardiogram sensor 33, a heart sound sensor 35, a microphone 37, and a camera 41. The temperature sensors 15 and 19 and the humidity sensors 17 and 21 measure the temperature or humidity of the air outside the housing as the test object.
[0083] The three-dimensional acceleration sensor 11 detects acceleration applied to the terminal device 1 in three mutually perpendicular directions (vertical direction (Z direction), width direction of the housing (Y direction), and thickness direction of the housing (X direction)), and outputs the detection results.
[0084] The three-axis gyro sensor 13 detects angular acceleration (counterclockwise speed in each direction is considered positive) in the vertical direction (Z direction) and any two directions perpendicular to the vertical direction (the width direction of the housing (Y direction) and the thickness direction of the housing (X direction)) as angular velocities applied to the terminal device 1, and outputs the detection results.
[0085] The temperature sensors 15 and 19 are configured to include, for example, thermistor elements whose electrical resistance changes depending on the temperature. In this embodiment, the temperature sensors 15 and 19 detect temperatures in degrees Celsius, and all temperature indications in the following description are in degrees Celsius.
[0086] The humidity sensors 17 and 21 are configured as, for example, well-known polymer film humidity sensors. These polymer film humidity sensors are configured as capacitors whose dielectric constant changes as the amount of moisture contained in the polymer film changes in response to changes in relative humidity.
[0087] The illuminance sensor 23 is configured as a well-known illuminance sensor including, for example, a phototransistor. The wind speed sensor 29 is, for example, a known wind speed sensor, and calculates the wind speed from the power (amount of heat radiation) required to maintain the heater temperature at a predetermined temperature.
[0088] The heart sound sensor 35 is configured as a vibration sensor that captures vibrations caused by the beating of the user's heart, and the MPU 31 distinguishes between vibrations and noise caused by beating and other vibrations and noise by taking into account the detection results from the heart sound sensor 35 and the heart sounds input from the microphone 37.
[0089] The wetness sensor 25 detects water droplets on the surface of the housing, and the electrocardiogram sensor 33 detects the heartbeat of the user. The camera 41 is disposed inside the housing of the terminal device 1 so that the outside of the terminal device 1 is within its imaging range.
[0090] The communication unit 50 includes a well-known MPU 51, a wireless telephone unit 53, and a contacts memory 55, and is configured to be able to acquire detection signals from various sensors constituting the behavior sensor unit 10 via an input / output interface (not shown). The MPU 51 of the communication unit 50 then executes processing according to the detection results from the behavior sensor unit 10, input signals input via the operation unit 70, and programs stored in a ROM (not shown).
[0091] Specifically, the MPU 51 of the communication unit 50 performs the functions of a motion detection device that detects specific motions performed by the user, a positional relationship detection device that detects the positional relationship with the user, an exercise load detection device that detects the load of exercise performed by the user, and a function of transmitting the processing results by the MPU 51.
[0092] The wireless telephone unit 53 is configured to be able to communicate with, for example, a mobile phone base station, and the MPU 51 of the communication unit 50 outputs the processing results of the MPU 51 to the notification unit 60 or transmits them to a pre-set destination via the wireless telephone unit 53.
[0093] The contacts memory 55 functions as a storage area for storing location information of places visited by the user. This contacts memory 55 stores information on contacts (such as telephone numbers) to be contacted in the event of an emergency involving the user.
[0094] The notification unit 60 includes a display 61 configured as, for example, an LCD or an organic EL display, illumination 63 configured from, for example, an LED capable of emitting light in seven colors, and a speaker 65. Each component constituting the notification unit 60 is driven and controlled by the MPU 51 of the communication unit 50.
[0095] Next, the operation unit 70 includes a touch pad 71, a confirmation button 73, a fingerprint sensor 75, and a rescue request lever 77. The touch pad 71 outputs a signal according to the position and pressure of the touch by the user (the user, the user's guardian, etc.).
[0096] The confirmation button 73 is configured so that when pressed by the user, the contacts of a built-in switch close, and the communication unit 50 can detect that the confirmation button 73 has been pressed.
[0097] The fingerprint sensor 75 is a well-known fingerprint sensor, and is configured to be able to read fingerprints using, for example, an optical sensor. Note that instead of the fingerprint sensor 75, any means capable of recognizing a person's physical characteristics (means capable of biometric authentication: means capable of identifying an individual), such as a sensor that recognizes the shape of veins in the palm of the hand, can be used.
[0098] It also has a rescue request lever 77 that, when operated, connects to a predetermined contact point. [Processing of this embodiment] The processing performed in such a voice response system 100 will be described below.
[0099] The voice response terminal processing performed by the terminal device 1 is a processing for accepting voice input by the user, sending this voice to the server 90, and, upon receiving a response to be output from the server 90, reproducing this response as voice. This processing is started when the user inputs via the operation unit 70 that he or she will make a voice input.
[0100] 3, first, the microphone 37 is set to a state in which it can receive input (ON state) (S2), and then image capture (recording) by the camera 41 is started (S4). Then, it is determined whether or not there is audio input (S6).
[0101] If there is no voice input (S6: NO), it is determined whether a timeout has occurred (S8). Here, a timeout means that the allowable time for waiting for processing has been exceeded, and in this case the allowable time is set to, for example, about 5 seconds.
[0102] If the time has expired (S8: YES), the process proceeds to S30, which will be described later. If the time has not expired (S8: NO), the process returns to S6. If there is voice input (S6: YES), the voice is recorded in memory (S10), and it is determined whether the voice input has ended (S12). Here, if the voice is interrupted for a certain period of time or if an instruction to end the voice input is input via the operation unit 70, it is determined that the voice input has ended.
[0103] If the voice input has not ended (S12: NO), the process returns to S10. If the voice input has ended (S12: YES), the device transmits data such as an ID for identifying itself, the voice, and the captured image as a packet to the server 90 (S14). The data transmission process may be performed between S10 and S12.
[0104] Next, it is determined whether or not the data transmission is complete (S16). If the transmission is not complete (S16: NO), the process returns to S14. If the transmission is completed (S16: YES), it is determined whether data (packets) transmitted in the voice response server process described below have been received (S18). If data has not been received (S18: NO), it is determined whether a timeout has occurred (S20).
[0105] If the time has expired (S20: YES), the process proceeds to S30, which will be described later. If the time has not expired (S20: NO), the process returns to S18. If data has been received (S18: YES), the packet is received (S22). In this process, one or more different responses to the text information, each associated with a different tone of voice, are acquired.
[0106] Then, it is determined whether or not the reception is complete (S24). If the reception is not complete (S24: NO), it is determined whether or not a timeout has occurred (S26). If a timeout has occurred (S26: YES), an error is reported via the notification unit 60, and the voice response terminal process is terminated. If a timeout has not occurred (S26: NO), the process returns to S22.
[0107] If reception is complete (S24: YES), a response based on the received packet is output as voice from the speaker 65 (S28). In this process, if multiple responses are to be played back, each of the multiple responses is played back in a different voice tone. When this process is completed, the voice response terminal process ends.
[0108] Next, the voice response server processing performed by the server 90 (external device) will be described with reference to Fig. 4. The voice response server processing is a process of receiving a voice from the terminal device 1, performing voice recognition to convert the voice into text information, and generating a response to the voice and returning it to the terminal device 1. In particular, in this embodiment, there are cases where a plurality of responses are transmitted in association with voices of different tones.
[0109] 4, the voice response server process first determines whether or not a packet has been received from any of the terminal devices 1 (S42). If no packet has been received (S42: NO), the process of S42 is repeated.
[0110] If a packet has been received (S42: YES), the terminal device 1 of the communication partner is identified (S44). In this process, the terminal device 1 is identified by the ID of the terminal device 1 contained in the packet.
[0111] Next, the speech contained in the packet is recognized (S46). Here, a large number of speech waveforms are associated with a large number of characters in the speech recognition DB 102. Also, a predictive conversion DB 103 is associated with words that tend to be used after a certain word.
[0112] Therefore, in this process, by referring to the speech recognition DB 102 and the predictive conversion DB 103, a well-known speech recognition process is performed to convert the speech into text information. Next, the captured image is processed to identify objects in the captured image (S48), and the user's emotions are determined based on the waveform of the voice and the endings of words (S50).
[0113] In this process, by referring to the emotion determination DB 114, which associates the waveform of the voice (tone of voice) and the endings of words with emotion categories such as normal, anger, joy, confusion, sadness, and elation, it is determined which category the user's emotion falls into, and this determination result is recorded in memory. Next, by referring to the learning DB 107, words that the user frequently speaks are searched for, and parts of the character information generated by voice recognition that were ambiguous are corrected.
[0114] The learning DB 107 records the characteristics of each user, such as the words the user frequently uses and pronunciation habits. Data is added and corrected to the learning DB 107 during conversations with the user.
[0115] Next, the corrected character information is identified as the input character information (S54), and a sentence similar to the character information is input and searched from the answer candidate DB 105 to obtain a response (S56). Here, the answer candidate DB 105 contains the input character information, the first output, the tone of voice of the first output, the second output, and the tone of voice of the second output, as shown in FIG. It is associated with righteousness.
[0116] For example, as shown in the first row of Fig. 5, when text information "Today's weather at *" is input, a first output "Today's weather at * is *" is output in association with the voice tone of Woman 1. However, the "*" part is acquired by accessing weather DB 110, which associates area names with weather forecasts for the next few days for each area.
[0117] Furthermore, when the text information "Today's weather for *" is input, the weather at the timing when today's weather changes is also obtained from the weather DB 110, and the second output "However, * is *" is output in association with the voice tone of Man 1. If the weather in Tokyo today is sunny and tomorrow's weather is rainy, and "Today's weather in Tokyo" is input, "Today's weather in Tokyo is sunny" is output in Woman 1's voice tone, and "However, tomorrow it will rain" is output in Man 1's voice tone.
[0118] In this embodiment, a case where multiple responses are output has been described, but if there is only one answer to the input, there will only be one response. Therefore, it is determined whether there is one response (S58). If there is only one response (S58: YES), the process proceeds to the processing of S62, which will be described later.
[0119] If there are multiple responses (S58: NO), the response contents are associated with the voice tones (S60). Here, the voice DB 104 stores a database of artificial voices for each voice tone, and in this process, the voice tones set for each response are associated with the voice tones in the database.
[0120] Next, the response content is converted into voice (S62). In this process, the response content (text information) is output as voice based on the database stored in the voice DB 104.
[0121] Then, the generated response (voice) is transmitted as a packet to the terminal device 1 of the other party of communication (S64). Note that the voice of the response content may be transmitted as a packet while being generated. Next, the conversation content is recorded (S68). In this process, the input character information and the output response content are recorded as the conversation content in the learning DB 107. At this time, keywords (words recorded in the speech recognition DB 102) included in the conversation content, pronunciation characteristics, etc. are also recorded in the learning DB 107.
[0122] When this process is completed, the voice response server process is terminated. [Effects of this embodiment] The voice response system 100 described above in detail is a system that responds to input text information by voice, and the terminal device 1 (MPU 31) acquires multiple different responses to the text information and outputs the multiple different responses in different voice tones.
[0123] According to such a voice response system 100, multiple responses can be output in different tones of voice, so even if a single answer cannot be identified for one piece of text information, different answers can be output in different tones of voice in a way that is easy for the user to understand, thereby making it easier for the user to use.
[0124] In the voice response system 100, the terminal device 1 receives input of a user's voice via the microphone 37, and the server 90 (the calculation unit 101) converts the input voice into text information, generates a plurality of different responses to the text information, and transmits them to the terminal device 1. Then, the terminal device 1 obtains the responses from the server 90.
[0125] According to the voice response system 100, since voice can be input to the terminal device 1, it is possible to configure the system so that text information can be input by voice. Also, since it is possible to configure the system so that a response is generated in the server 90, it is possible to reduce the processing load on the voice response system 100.
[0126] Furthermore, in the voice response system 100, the server 90 converts the user's speech into text information and accumulates speech habits (pronunciation habits, etc.) as learning information (capturing and recording characteristics).
[0127] According to such a voice response system 100, text information can be generated based on learning information, and therefore the accuracy of generating text information can be improved. Furthermore, in the voice response system 100, the server 90 reads the emotion from the tone of voice input by the user and outputs which emotion corresponds to the voice, usually including at least one of anger, joy, confusion, sadness, and elation.
[0128] According to such a voice response system 100, it is possible to output a response according to the user's emotions. [Modification of the first embodiment] In this embodiment, voice recognition is used as a configuration for inputting text information, but it is not limited to voice recognition and input may also be made using input means (operation unit 70) such as a keyboard or a touch panel. Also, the operation of "converting input voice into text information" is performed by server 90, but it may also be performed by terminal device 1.
[0129] Furthermore, in the above voice response system 100, the server 90 may be provided with a candidate response DB 105 in which a plurality of different responses, including positive and negative responses to each piece of text information, are recorded, and the terminal device 1 may acquire the positive and negative responses as the plurality of different responses and play the positive and negative responses in different tones of voice.
[0130] For example, as shown in the second row of Figure 5, when a voice is input saying "May I buy this?" for some item, positive information about the item, such as a good reputation, is output in a female voice. On the other hand, negative information, such as a bad reputation, is output in a voice tone different from the female voice associated with the positive information (here, a male voice).
[0131] According to the voice response system 100, responses with different positions, such as positive and negative responses, can be reproduced in different tones of voice, making it possible to reproduce the voice as if it were being spoken by a different person, thereby making it less likely that the user listening to the voice will feel uncomfortable.
[0132] The tone of voice may be changed depending on the type of response or the language used when responding. For example, if the response is to be made in a gentle tone, a calm female voice may be used, and if the response is to be made in a fierce tone, a brave male voice may be used. In other words, the content of the response may be associated with a personality, and the tone of voice may be set according to the personality.
[0133] Furthermore, in the voice response system 100, a response (for example, an affirmative response or a negative response) output by the user's own terminal device 1 or another terminal device 1 may be input as text information, and a response for refuting this response may be generated. In other words, from the user's perspective, it is possible to hear arguments from both pro and con positions. Then, after listening to this argument, the user can make a final decision.
[0134] This configuration can be realized using one or more terminal devices 1. For multiple terminal devices 1 to exchange audio with each other, audio may be input and output directly, or wireless communication may be used. When multiple terminal devices 1 communicate with the server 90, data may be sent to the other terminal devices 1 in the processing of S66.
[0135] Furthermore, in the voice response system 100, the calculation unit 101 may learn (record and analyze) the user's behavior (conversations, places traveled, and what is captured on camera) and compensate for any shortcomings in the user's conversation.
[0136] For example, in a conversation in which the user responds to the question "Is hamburger steak okay today?" with "Curry would be good," the device can add, "Because we had hamburger steak yesterday," thereby conveying the reason why the user said curry would be good.
[0137] Such configuration can also be implemented during a phone call and may be configured to participate in the user's conversation without permission. Furthermore, in the voice response system 100, the server 90 may acquire candidate responses from a predetermined server or from the Internet.
[0138] According to such a voice response system 100, response candidates can be obtained not only from the server 90 but also from any device connected via the Internet, a dedicated line, or the like. [Second embodiment] [Processing of the second embodiment] Next, another embodiment of the voice response system will be described. In the following embodiments, only the differences from the voice response system 100 of the first embodiment will be described in detail, and the same parts as those of the voice response system 100 of the first embodiment will be assigned the same reference numerals and will not be described again.
[0139] In the voice response system of the second embodiment, voice is output even when the user does not input text information. Specifically, the terminal device 1 performs the automatic conversation terminal process shown in Fig. 6. The automatic conversation terminal process is started, for example, when the terminal device 1 is powered on, and is then repeatedly executed.
[0140] In the automatic conversation terminal process, first, it is determined whether or not the setting for automatic conversation is ON (S82). Note that the user can set whether or not to perform automatic conversation via the operation unit 70 or by inputting voice.
[0141] If the automatic conversation setting is OFF (S82: NO), the automatic conversation terminal process ends. If the automatic conversation setting is ON (S82: YES), the device transmits to the server 90 a message indicating that the device is in automatic conversation mode together with its own ID (S84).
[0142] Next, it is determined whether or not a packet has been received from the server 90 (S86). If a packet has not been received (S86: NO), the process of S86 is repeated. If a packet has been received (S86: YES), the same processes as those of S22 to S30 described above are carried out, and when these processes are completed, the automatic conversation terminal process is terminated.
[0143] The server 90 also executes an automatic conversation server process shown in Fig. 7. The automatic conversation server process is started, for example, when the server 90 is powered on, and is then repeatedly executed.
[0144] In the automatic conversation server process, first, it is determined whether or not a notification that the automatic conversation mode has been set has been received from the terminal device 1 (S92). If the notification that the automatic conversation mode has been set has not been received (S92: NO), the process proceeds to S98.
[0145] If the notification that the automatic conversation mode has been set has been received (S92: YES), the terminal device 1 to be the communication partner is identified based on the ID included in the received packet (S94), and an automatic conversation is set for this communication partner (S96). Next, for each of the terminal devices 1 set to the automatic conversation mode, it is determined whether the playback conditions are met (S98).
[0146] Here, playback conditions include, for example, a certain amount of time having passed since the last conversation (voice input), a certain time of day, specific weather, or when any sensor value indicates an abnormality.
[0147] If the playback conditions are not met (S98: NO), the automatic conversation server process is terminated. If the playback conditions are met (S98: YES), a message according to the playback conditions is generated (S100).
[0148] Here, the message according to the playback conditions may be, for example, a fixed phrase such as "Good morning" or "Hello," or may be about the latest news obtained from the automatically updated news DB 109. When the message is about the latest news, for example, if information about the stock price of a certain company is obtained, the message may be, "The stock price of XX company went up by XX yen today. Did you know?"
[0149] When this process is completed, the processes of S42 to S54 described above are performed. Then, when the process of S54 is completed, it is determined whether or not a predetermined response has been obtained from the terminal device 1 that is the communication partner (S112). Here, the predetermined response may be, for example, some kind of voice or a specific answer. For example, a specific answer would be an answer such as "I know" or "I don't know" in response to the question "Do you know?", and an answer containing words indicating the weather, such as "It's raining" or "It's sunny" in response to the question "What's the weather like now?"
[0150] If there is a predetermined response (S112: YES), the automatic conversation server process ends. If there is no predetermined response (S112: NO), the message sent in S100 is resent (S114). When resending the message in this way, the tone of voice is changed to produce a stronger, harsher tone of voice.
[0151] Next, the server 100 refers to the report destination DB 117 in which the terminal device 1 and the report destination are previously associated with each other, and sends a message to the predetermined report destination indicating that there was no reply (S116). When this process is completed, the automatic conversation server process is terminated.
[0152] [Effects of the second embodiment] In the above voice response system 100, when no text information is input, the server 90 determines whether the situation of the voice response system 100 matches a playback condition that is set in advance as a condition for outputting a voice. If the playback condition is met, the server 90 outputs a preset message.
[0153] According to such a voice response system 100, it is possible to output voice even when no text information is input (i.e., when the user does not speak). For example, by forcing the user to speak, it can be used as a measure to suppress drowsiness while driving a car. In addition, by determining whether or not a person living alone responds, it is possible to confirm their safety.
[0154] In the voice response system 100, the server 90 acquires news information and outputs a message relating to the news in the form of a question requesting a response from the user. According to such a voice response system 100, conversations about the news can be held, which can prevent conversations from becoming the same over and over again.
[0155] Furthermore, in the voice response system 100, the server 90 adds separately acquired external information (such as news and environmental information (temperature, weather, location information, etc.)) to a preset message and outputs it.
[0156] According to such a voice response system 100, it is possible to output a response that combines a predetermined message with the acquired information. Furthermore, in the voice response system 100, if no response to a response or message is received, the server 90 transmits to a pre-set contact point information identifying the user and a message indicating that no response was received.
[0157] According to such a voice response system 100, if no answer is received, it is possible to notify the contact person. Therefore, for example, an abnormality in an elderly person living alone can be reported early.
[0158] [Modification of the second embodiment] Furthermore, in the voice response system 100, the server 90 may acquire a plurality of messages and select and output a message to be reproduced depending on the frequency of message reproduction.
[0159] According to such a voice response system 100, by making it difficult to play messages that are played frequently, it is possible to create a sense of randomness when playing messages, or by deliberately repeatedly playing messages that are played frequently, it is possible to draw attention and encourage memory retention.
[0160] [Third embodiment] [Processing of the third embodiment] Next, in the voice response system of the third embodiment, the terminal device 1 is configured to convey on behalf of the user things that the user finds difficult to say to someone directly. For example, before a date, if the user speaks to the device about what he or she would like to say today, the voice response system 100 will speak on behalf of the user (play back audio) at an appropriate timing (for example, at a preset time or when a certain amount of time has passed since the conversation stopped).
[0161] In detail, the terminal device 1 performs the message terminal process shown in Fig. 8, and the server 90 performs the message server process shown in Fig. 9. The message terminal process is started, for example, when the power of the terminal device 1 is turned on, and is then repeatedly executed.
[0162] In the message terminal process, first, it is determined whether or not the message mode has been set by the user (S132), as shown in Fig. 8. If the message mode has not been set (S132: NO), the process of S132 is repeated.
[0163] If the message mode is set (S132: YES), the processes of S2 to S8 are carried out, and if the determination in S6 is affirmative, the message mode flag is set to ON in the memory of the terminal device 1 (S134), and the processes of S10 to S16 are then carried out.
[0164] If the determination in S16 is affirmative, it is determined whether or not a packet has been received from the server 90 (S136). If a packet has not been received (S136: NO), the process of S136 is repeated. If a packet has been received (S136: YES), the processes of S24 to S30 are carried out, and the message terminal process is terminated.
[0165] Next, the message server process is a process that starts, for example, when the server 90 is powered on, and is executed repeatedly thereafter. In detail, first, it is determined whether or not a packet has been received from any of the terminal devices 1 (S142). If no packet has been received (S142: NO), the process proceeds to the process of S156, which will be described later.
[0166] If a packet has been received (S142: YES), the terminal device 1 of the communication partner is identified (S44), and it is determined whether or not the packet contains a mode flag such as a message mode flag (S144).If no mode flag is present (S144: NO), the process proceeds to S148.
[0167] Furthermore, if there is a mode flag (S144: YES), the server 90 also sets the mode by setting the flag corresponding to the terminal device 1 of the communication partner to ON (S146). For example, if the message mode flag is the corresponding message mode, the processes of S46 to S152 described later are performed, and if the guidance mode flag described later is the corresponding guidance mode, the processes of S46 to S176 (see FIG. 11) are performed.
[0168] Next, it is determined whether the message flag is ON (S148). If the message flag is ON (S148: YES), the processes of S46 to S54 are carried out, and then the message replay conditions are extracted (S150).
[0169] Here, the message playback conditions can be set in advance by the user via the operation unit 70 of the terminal device 1, and correspond to, for example, the time and the location. The message playback conditions are transmitted to the server 90 when packets for message terminal processing are transmitted.
[0170] Next, the message and the voice (tone of voice) are associated and recorded in memory (S152), and the process proceeds to S156. If the message flag is OFF (S148: NO), the process proceeds to S154 for other modes, and it is determined whether playback timing has arrived (S156). Here, playback timing refers to the content set in the message playback conditions.
[0171] If it is not the playback timing (S156: NO), the message server process is immediately terminated. If it is the playback timing (S156: YES), the processes of S62 to S64 are carried out, and the message server process is terminated.
[0172] [Effects of the third embodiment] According to the voice response system of the third embodiment, the voice input by the user is not played back immediately, but can be played back after a certain time has elapsed when the message playback condition is met.
[0173] For example, as shown in the third row of Figure 5, if you input "Please tell Mr. / Ms. XX that XX," the sentence you want to convey will be played after Mr. / Ms. XX's voice is recognized (heard).
[0174] [Modification of the third embodiment] In the third embodiment, the content of what the user said is reproduced. The terminal device 1 may be configured to speak words that will trigger the conversation, such as, "By the way, didn't you say you were going to tell her something?" In detail, the terminal device 1 performs the guidance terminal process shown in Fig. 10, and the server 90 performs the guidance server process shown in Fig. 11.
[0175] The guiding terminal process is started, for example, when the terminal device 1 is powered on, and is then repeatedly executed.The guiding terminal process is started, for example, when the terminal device 1 is powered on, and is then repeatedly executed.
[0176] In the guidance terminal process, first, it is determined whether or not the guidance mode has been set by the user (S162), as shown in Fig. 10. If the guidance mode has not been set (S162: NO), the process of S162 is repeated.
[0177] If the guidance mode is set (S162: YES), the processes of S2 to S8 are performed, and if the determination in S6 is affirmative, the guidance mode flag is set to ON in the memory of the terminal device 1 (S164). Then, the processes of S10 to S16 are performed.
[0178] If the determination in S16 is affirmative, it is determined whether or not a packet has been received from the server 90 (S166). If a packet has not been received (S166: NO), the process of S166 is repeated. If a packet has been received (S166: YES), the processes of S24 to S30 are performed, and the guiding terminal process is terminated.
[0179] Next, the guidance server process is started, for example, when the server 90 is powered on, and is thereafter repeatedly executed. Specifically, the process executes the above-mentioned processes of S142 to S146. Then, it is determined whether the guidance flag is in the ON state (S172).
[0180] If the guidance flag is ON (S172: YES), the processes of S46 to S54 are carried out, and then the guidance regeneration conditions are extracted (S174). Here, like the message playback conditions, the user can set the guided playback conditions in advance via the operation unit 70 of the terminal device 1, and these conditions include, for example, the time and the location. The guided playback conditions are transmitted to the server 90 when a packet for message terminal processing is transmitted.
[0181] Next, the system generates prompting content, associates the prompting content with the voice (tone of voice), and records the associated content in memory (S176). For example, the system searches for words expressing desires, such as "want to do" or "hope," included in the input text information, extracts keywords before these words, and outputs words registered as prompting words for these keywords as prompting content. Keywords and words representing prompting content are associated with each other and recorded in the response candidate DB 105 in advance.
[0182] Next, the process from S156 onwards is carried out, and the server process is terminated. If the guidance flag is OFF (S172: NO), the process for another mode is carried out (S154), and the process from S156 onwards is carried out, and the server process is terminated.
[0183] According to the configuration of this modified example of the third embodiment, the user can be guided to speak the words he or she wants to say, rather than directly outputting the words he or she wants to say. [Fourth embodiment] [Processing of the fourth embodiment] Next, an example of using the terminal device 1 for reception work will be described. In this embodiment, the terminal device 1 is installed at a company reception desk or the like. It can also be used to receive calls from a company's main phone line or for telephone banking. Here, in this embodiment, the process of S56 in the first embodiment is realized by replacing it with the reception process shown in FIG. 12.
[0184] In the reception process, as shown in Fig. 12, first, it is determined whether or not the character information includes a company name (S192). In this process, it is determined whether or not a general name or a company name (recorded in the voice recognition DB 102) is included.
[0185] If the text information does not include a company name or a personal name (S192: YES), a response is generated to ask for the company name and the personal name (S194), and the reception process ends. In this process, a response such as "Please tell us your name and business" is generated.
[0186] If the text information contains a company name or an individual name (S192: NO), the company name or individual name is extracted from the sales DB 118 and the client DB 119 (S196). The sales DB 118 stores the names of companies and their representatives who have made sales calls in the past, or the names of complainers who only complain. The client DB 119 stores the names of companies, their representatives, the representatives on the user's side of the terminal device 1 (their own company), schedules such as scheduled meeting times, and contact information for each representative.
[0187] Next, it is determined whether or not the company name or personal name can be extracted from the sales DB 118, that is, whether or not the company name or personal name included in the character information is included in the sales DB 118 (S198). If the company name or personal name can be extracted from the sales DB 118 (S198: YES), a sales refusal response (a response refusing the transfer) is generated to decline the sale (S200), and the reception process ends.
[0188] If the company name or individual name cannot be extracted from the sales DB 118 (S198: NO), it is determined whether the person at the reception desk is scheduled to visit at a nearby time (for example, within one hour before or after the current time) according to the schedule in the client DB 119 (S202). If the person is scheduled to visit at a nearby time (S202: YES), the contact information of the person in charge of this person is extracted from the client DB 119, and the person at the reception desk is connected to this person so that they can talk (S204). In this process, a connection can be made to the person's extension phone, mobile phone, etc.
[0189] Next, a reception response for the client is generated (S206). Here, an example of the reception response for the client is generated as follows: "Dear Mr. / Ms. XX, Thank you for your continued support. We are connecting you to the person in charge, so please wait a moment." When this process is completed, the reception process ends.
[0190] If the visitor is not coming soon (S202: NO), the system connects the caller to a pre-set contact for reception, and connects the caller to the receptionist so that the person can talk to the receptionist (S208).Then, a normal reception response is generated (S210).
[0191] Here, as a normal reception response, for example, a response such as "We are connecting to the reception, so please wait for a while" is generated. When this process is completed, the reception process ends.
[0192] [Effects of the fourth embodiment] The voice response system 100 is configured to be used at the reception desk of a workplace or company. In this configuration, the name of a salesperson and the name of the company are pre-recorded in the sales DB 118 of the server 90, and when a salesperson gives this name and company name at the reception desk, a response is generated to play back a voice refusing the salesperson's request.
[0193] In the voice response system 100, the server 90 responds to the input text information. The communication partner is identified by the above and a predetermined communication destination for each communication partner is connected to the communication partner. Such a voice response system 100 can assist with reception work and telephone correspondence. Also, such a voice response system 100 can remove people who may be interfering with the user's work without the user having to respond to them.
[0194] Furthermore, in the voice response system 100, the server 90 extracts keywords contained in the input character information (particularly voice) and connects to the connection destination corresponding to the keyword. Note that keywords such as the name of the other party and the connection destination are associated in advance.
[0195] Such a voice response system 100 can assist with tasks such as transferring calls and calling the reception desk. [Modification of the fourth embodiment] In the above embodiment, the connection destination is set according to the other party, but this technology may be applied to, for example, when accepting calls for telephone banking, telephone shopping, etc., by recognizing requirements (keywords contained in text information) and changing the connection destination according to the requirements.
[0196] In the voice response system 100, the server 90 may recognize the requirements of the other party's speech based on keywords and convey to the user an outline of what the other party has said. Such a voice response system 100 can assist in the intermediary work with customers.
[0197] [Fifth embodiment] [Processing of the fifth embodiment] Next, the terminal device 1 may receive a request from another terminal device 1 and provide the information that the other terminal device 1 requests.
[0198] When configured in this manner, the server 90 requests the necessary information from the other terminal device 1 in the process of S56, and generates a response after obtaining the necessary information from the other terminal device 1. Then, the terminal device 1 that provides the necessary information performs the information providing terminal process shown in Fig. 13. The information providing terminal process is a process that is started, for example, when a request is received from the server 90.
[0199] 13, the information providing terminal process first extracts an information providing destination (S222). This information providing destination indicates another terminal device 1 requesting information, and an ID for identifying this other terminal device 1 is included in the request from the server 90.
[0200] Next, it is determined whether the person is a party to whom information provision is permitted (S224). Here, IDs of parties to whom information provision is permitted, such as family and friends, are pre-recorded in the terminal information DB 113. In this process, the determination is made by referring to this terminal information DB 113.
[0201] If the information provision is permitted (S224: YES), the requested information is acquired from its own memory 39, various sensors, etc. (S226), and this data is transmitted to the server 90 (S228). If the information provision is not permitted (S224: NO), a message to the effect that the information provision is rejected is transmitted to the server 90 (S230).
[0202] When this processing is completed, the information providing terminal processing is terminated. In this configuration, for example, as shown in the fourth row of Figure 5, in response to the question "What is Mr. / Ms. XX doing?", the server 90 requests location information from Mr. / Ms. XX's terminal device 1, and the terminal device 1 returns the location information.
[0203] The server 90 then recognizes Mr. / Ms. XX's actions based on the location information. For example, if the person is moving on the tracks at a speed faster than a human can run, the server 90 determines that the person is traveling on a train and generates a response such as, "Mr. / Ms. XX is on the train. It appears that he / she is on his / her way home."
[0204] [Effects of the fifth embodiment] In the voice response system 100, the server 90 acquires information recorded in another terminal device 1 different from the requesting terminal device 1, and provides the information to the other terminal device 1. That is, in the voice response system 100, the server 90 acquires information for generating a response to text information from the other terminal device 1.
[0205] According to such a voice response system 100, a response can be generated based on information recorded in another terminal device 1. Furthermore, in the voice response system 100, when a terminal device 1 is requested by another terminal device 1 to provide information for generating a response to text information, the terminal device 1 returns information in response to this request.
[0206] In this configuration, the terminal device 1 is provided with sensors for detecting location information, temperature, humidity, illuminance, noise level, etc., and a database of dictionary information, etc., and extracts necessary information in response to a request.
[0207] According to such a voice response system 100, it is possible to obtain information specific to other terminal devices 1, such as the location of other terminal devices 1. It is also possible to transmit information specific to one terminal device 1 to other terminal devices 1.
[0208] [Sixth embodiment] [Processing of the Sixth Embodiment] Next, in the voice response system of the sixth embodiment, a personality DB 106 is prepared, which stores personality information in which the personalities of users or related parties representing people related to the user are associated with predetermined categories. For example, as shown in Fig. 14, the personality DB 106 stores the names of users and related parties in association with their personality categories.
[0209] 14, personality tests are administered to users and related parties, and the test results are also recorded. Well-known personality analysis techniques (such as the Rorschach test and the Szondi test) may be used to generate personality information. Furthermore, aptitude testing techniques used by companies and other organizations for employment testing may also be used to generate personality information.
[0210] When generating personality information, for example, a personality information generation process shown in Fig. 15 is performed. The personality information generation process is started when, for example, a command to generate personality information is input using the operation unit 70 or the like in the terminal device 1.
[0211] 15, in the personality information generation process, first, the microphone 37 is turned on (S242), and one of the predetermined multiple-choice questions is output by voice (S244). At this time, the multiple-choice question may be obtained from the server 90, or a question previously stored in the memory 39 may be posed.
[0212] Next, it is determined whether or not there is a voice response from the subject (the user or a related person) (S246). If there is no response (S246: NO), the process of S246 is repeated. If there is an answer (S246: YES), conversation parameters such as the ending of words and conversation speed are extracted (S248), and it is determined whether the current question is the final question (S250). If it is not the final question (S250: NO), the next question is selected (S252), and the process returns to S242.
[0213] If it is the final question (S250: YES), a personality analysis is performed by answering a multiple-choice question (S254), and a personality analysis is performed using conversation parameters (S256). Here, the personality analysis using conversation parameters can capture the tendency for confident people to speak with stronger accents and unconfident people to speak with weaker accents, or for impatient people to speak quickly and calm people to speak slowly, etc.
[0214] Next, these personality analysis results are analyzed comprehensively, such as by taking a weighted average (S258), and then assigned to a personality category (S260). Specifically, the subject's personality obtained through the test is converted into a score, and each score is assigned to a personality category.
[0215] Next, the subject and the personality category are associated (S262) and recorded in the personality DB 106 (S264). That is, the relationship between the subject and the personality category is transmitted to the server 90. At this time, the test results are also transmitted to the server 90, and the server 90 constructs the personality DB 106 as shown in Fig. 14. When this processing is completed, the personality information generation processing is terminated.
[0216] When using the personality DB 106 generated in this way, the personality categories are associated with different responses and are prepared in the response candidate DB 105. Then, in the process of S56, the server 90 acquires response candidates representing a plurality of different responses to the text information, selects a response to be output from the response candidates according to the personality information, and in the processes of S60 and S64, outputs the selected response.
[0217] [Effects of the sixth embodiment] In the voice response system 100, the terminal device 1 generates personality information of the user or related person based on answers to a plurality of questions set in advance, and acquires the generated personality information.
[0218] According to such a voice response system 100, the personality information can be generated in the server 90 or the terminal device 1. Furthermore, in the voice response system 100, the calculation unit 101 generates personality information of the user or related person based on the character string included in the input text information.
[0219] According to the voice response system 100, the user can generate personality information while using the voice response system 100. Furthermore, such a voice response system 100 can provide different responses depending on the personality of the user and those related to the user (stakeholders), thereby improving usability for the user.
[0220] [Modification of the sixth embodiment] In the sixth embodiment, the response may be narrowed down to one according to the personality and then output, or a plurality of responses may be output in correspondence with different tones of voice.
[0221] Furthermore, the processes of S248 and S254 to S264 in the personality information generation process may be performed by the server 90. In this case, similar to the first embodiment, the server 90 may identify the terminal device 1, and voice and questions may be exchanged between the terminal device 1 and the server 90.
[0222] Furthermore, in the voice response system 100, the server 90 may detect any of the user's actions and operations and generate learning information or personality information based on these.
[0223] For example, if such a voice response system 100 detects that a user has jumped on a train for several consecutive days, it can prompt the user to leave home a few minutes earlier from the next day, or if it detects from conversation that the user has a tendency to get angry easily, it can output voice or music to calm the user down.
[0224] [Seventh embodiment] [Processing of the Seventh Embodiment] Next, in the voice response system of the seventh embodiment, a preference DB 108 is prepared, which stores preference information in which the preferences of users and related parties are associated with preset categories. For example, as shown in Fig. 16, the preference DB 108 stores the names of users and related parties and their preferences in association with each of preference categories such as food preferences (food), color preferences (colors), hobbies, etc.
[0225] In particular, food preferences are classified as sweet tooth (sweet), spicy tooth (spicy), or something in between (average); color preferences are classified as warm colors (warm), cool colors (cold), or something in between (average); and hobbies are classified as indoor hobbies (indoor), outdoor hobbies (outdoor), or both indoor and outdoor hobbies (indoor and outdoor).
[0226] When constructing such a preference DB 108, for example, a preference information generation process shown in Fig. 17 is executed. The preference information generation process is carried out, for example, between S48 and S54. 17, keywords related to preferences are extracted from the text information (S282), and objects identified by image processing are extracted that relate to preferences (S284). Note that preference keywords are associated with preference types and classifications within those types (e.g., sweet, normal, spicy, etc. for food preferences) in preference DB 108, and in these processes, if the extracted keywords or objects are included in preference DB 108, they are extracted as being related to preferences.
[0227] Next, a counter is incremented for each group of keywords related to preferences (S288). For example, if a keyword such as kimchi is extracted, which has a preference type of "food preference" and a type of "spicy," the counters corresponding to "food preference" and "spicy" are incremented.
[0228] Then, the preference information (preference DB 108) is updated based on the counter value (S290). That is, for each "preference type," the "type" with the largest counter value is deemed to be the one that best matches the preference, and is recorded in preference DB 108 as a characteristic of the preference of the user or related person. When this process is completed, the preference information generation process is terminated.
[0229] When the preference DB 108 generated in this manner is used, different responses corresponding to each preference are prepared in the response candidate DB 105, and the server 90 acquires response candidates representing a plurality of different responses to the text information in the processing of S56, selects a response to be output from the response candidates according to the preference information, and outputs the selected response in the processing of S60 and S64.
[0230] [Effects of the Seventh Embodiment] In the voice response system 100, the server 90 generates preference information indicating the preferences of the user or related parties based on the character strings included in the text information, selects a response to be output from the response candidates based on the preference information, and outputs the selected response.
[0231] According to such a voice response system 100, responses can be made according to the preferences of the user or the person concerned. For example, when the user is buying a present for a related person, he or she can ask the terminal device 1, "What would Mr. / Ms. XX want?" and get a response according to the preference information.
[0232] [Modification of the Seventh Embodiment] The reply candidate DB 105 may have a table in which personality categories are associated with preference information, as shown in FIG.
[0233] For example, in the example shown in FIG. 18, personality categories are associated with color preferences, and products that are estimated to make women happy as presents are arranged in a matrix. In the process of S56, a response can be generated taking both personality and preferences into account.
[0234] [Eighth embodiment] [Processing of the Eighth Embodiment] In the above embodiment, voice is converted into text information, but a movement by the user may also be converted into text information.
[0235] In detail, the terminal device 1 captures the user's motion as a captured image and transmits it to the server 90, and the server 90 executes, for example, the motion character input process shown in Fig. 19. The motion character input process is started when a body part of the user is captured in the captured image in the process of S48.
[0236] In the motion character input process, as shown in Fig. 19, first, a captured image is acquired (S302), and then it is determined whether the user is trying to input characters by handwriting or by sign language (S304, S308).
[0237] In these processes, for example, if the captured image shows the user's upper body along with their face, it is determined that they are attempting to input characters using sign language, and if the captured image shows the user's hands but not their face, it is determined that they are attempting to input characters by hand.
[0238] If a character is being input by handwriting (S304: YES), the behavior of the fingertip or pen tip is recorded (S306), and the behavior is converted into character information based on this behavior (S312). Here, the handwritten character / sign language DB 112 associates the behavior when writing a character with the character, and also associates hand movements with characters expressed in sign language. In the process of S312, character information is generated by referring to the handwritten character / sign language DB 112.
[0239] If the user is attempting to input characters in sign language (S304: NO, S308: YES), the system refers to the handwritten character / sign language DB 112 to recognize the sign language content and executes the process of S312 described above. If the user is not attempting to input characters by handwriting or sign language (S308: NO), the system executes input processing using another method (S314).
[0240] Next, the character input by action and the character input by voice are associated with each other, and it is determined whether there is a similar voice (whether the degree of match between the reference waveform based on the character and the pronunciation waveform is equal to or greater than a reference value) (S316). If such a voice input is found (S316: YES), the accent and pronunciation characteristics when the user inputs this character are associated with the character and recorded in learning DB 107 (S318), and the action character input process is terminated.
[0241] If there is no such voice input (S316: NO), the motion character input process is terminated. do. [Effects of the Eighth Embodiment] In the voice response system 100, the user's actions are converted into text information, so that the user can input text information without speaking.
[0242] [Modification of the Eighth Embodiment] The movements in this embodiment may be not only handwritten characters or gestures (for example, sign language) but also movements resulting from muscle movements.
[0243] [Ninth embodiment] [Processing of the ninth embodiment] When a user uses a terminal device 1 other than the terminal device 1 that the user normally uses, the contents of the learning DB 107 may be made available to the other terminal device 1. In this case, the other terminal device 1 transmits the ID and password of the terminal device 1 that the user normally uses together with a use request to the server 90.
[0244] Then, the server 90 executes the other terminal use process shown in Fig. 20. The other terminal use process is started when a use request is received. In the other terminal use process, first, it is determined whether or not an ID and password have been input (S332), as shown in Fig. 20. If an ID and password have not been input (S332: NO), the process of S332 is repeated.
[0245] If an ID and password have been input (S332: YES), it is determined whether authentication using the ID and password has been completed (S334). If authentication has been completed (S334: YES), a message indicating that authentication has been completed is transmitted to the other terminal device 1 (S336), and the other terminal device 1 is set to use the learning DB 107 of the terminal device 1 whose ID and password correspond (S338).
[0246] If the authentication is not completed (S334: NO), an error message is sent to the other terminal device 1 (S340), and the other terminal use process is terminated. [Effects of the ninth embodiment] In the voice response system 100, the server 90 transfers learning information of a certain terminal device 1 to other terminal devices 1.
[0247] According to such a voice response system 100, even when a user of a certain terminal device 1 uses another terminal device 1, the user can use the learning information recorded in the certain terminal device 1 (the learning information recorded in the server 90). Therefore, the accuracy of generating text information can be improved even when the other terminal device 1 is used. This is particularly effective when a user has multiple terminal devices 1.
[0248] Furthermore, in the voice response system 100, the server 90 outputs information about the user in response to an inquiry from a person other than the user. According to such a voice response system 100, if the user's dietary habits, walking distance, etc. are detected, it can answer questions on behalf of the user at a hospital, etc. It may also be configured to learn the user's health condition, self-introduction, etc.
[0249] [Modification of the ninth embodiment] As in the configuration of the ninth embodiment, when a request to end use and an ID and password are received, use of the learning DB 107 for the terminal device 1 corresponding to the ID and password may be ended (prohibited).
[0250] [Tenth embodiment] [Processing of the 10th embodiment] In the voice response system of the tenth embodiment, the server 90 memorizes the conversation content and asks questions to obtain the same content about the content that was heard. More specifically, in S100 of the automatic conversation server process shown in Fig. 7, the memory confirmation process shown in Fig. 21 is executed.
[0251] In the memory confirmation process, as shown in Fig. 21, past conversation contents are extracted from the learning DB 107 (S352), and a question is generated with an answer being a keyword contained in any of the conversation contents (S353). When this process is completed, the memory confirmation process ends.
[0252] In memory verification processing, questions such as "What was on the menu for dinner last night?" or "Where did you go three days ago?" can be asked. [Effects of the Tenth Embodiment] Such a voice response system 100 can check the memory of the user and help the user to solidify the memory, which is also considered to be effective in preventing the progression of dementia in the elderly.
[0253] [Eleventh embodiment] [Processing of the 11th embodiment] Next, the voice response system of the eleventh embodiment is configured so that the user can practice a foreign language by using the terminal device 1 and the server 90.
[0254] In detail, pronunciation determination process 1 shown in Fig. 22, pronunciation determination process 2 shown in Fig. 23, and pronunciation determination process 3 shown in Fig. 24 are executed in this order. However, the server 90 executes one of the pronunciation determination processes 1 to 3 each time the voice response server process (Fig. 2) is executed. Furthermore, each of the pronunciation determination processes 1 to 3 is executed as the process of S56 described above.
[0255] First, in pronunciation determination process 1, a response is generated instructing the user to input a predetermined sentence by voice (S362), as shown in Fig. 22. In this process, for example, a model sentence in a foreign language is generated, and the user is prompted to speak the model sentence by imitating it. When this process is completed, pronunciation determination process 1 is terminated.
[0256] Next, when speech is input in accordance with pronunciation determination process 1, pronunciation determination process 2 is carried out. In pronunciation determination process 2, the accuracy of pronunciation and accent is scored (S372), as shown in Fig. 23. In this process, speech is treated as a waveform, and the degree of match between the waveform and a waveform of a model sentence is scored.
[0257] Then, this score is recorded in memory (S374), and pronunciation determination process 2 is terminated. Next, pronunciation determination process 3 is performed. In pronunciation determination process 3, as shown in Fig. 24, first, it is determined whether or not the score is less than a threshold value (S382).
[0258] If the score is less than the threshold (S382: YES), a response is generated to instruct the user to input a similar sentence again (S384). In this process, for example, a response is generated to prompt the user to speak by imitating the example again.
[0259] If the score is equal to or greater than the threshold (S382: NO), a response is generated to inform the user that the pronunciation was good and to prompt the user to enter the next sentence (S386). For example, a response such as "Good pronunciation. Let's move on to the next sentence" is generated.
[0260] When this process is completed, the pronunciation determination process 3 is completed. [Effects of the eleventh embodiment] In the voice response system 100, the server 90 detects the accuracy of the pronunciation and accent of the voice input by the user, and outputs the detected accuracy.
[0261] Such a voice response system 100 makes it possible to check the accuracy of pronunciation and accent, which is effective when practicing a foreign language, for example. Furthermore, in the voice response system 100, the server 90 outputs the same question again if the accuracy is equal to or less than a certain value.
[0262] According to such a voice response system 100, it is possible to obtain an accurate answer by outputting the same question. [Modification of the eleventh embodiment] In the voice response system 100, if the accuracy is equal to or less than a certain value, the server 90 may output, for confirmation, a voice containing a word that is closest to the pronunciation made by the user.
[0263] Such a voice response system 100 allows the user to check the accuracy of pronunciation and accent. [Twelfth embodiment] [Processing of the 12th embodiment] Next, a description will be given of a voice response system according to a twelfth embodiment. The voice response system according to the twelfth embodiment detects the emotion of a user from a voice input by the user and generates a response that soothes the user according to the emotion.
[0264] In detail, the emotion determination process shown in Fig. 25 and the emotion response generation process shown in Fig. 26 are executed. The emotion determination process is carried out as detailed processing of S50 described above, and as shown in Fig. 25, first, emotions are scored based on tone of voice, stress at the end of sentences, length of a sentence, speed of speech, unexpected words, etc. (S392), and then the emotions are classified according to the score and recorded in memory (S394).
[0265] When this process is completed, the emotion determination process ends. Next, in the process of S56 described above, the emotion response generation process is executed. In detail, as shown in Fig. 26, first, the emotion category set in the emotion determination process is determined (S412). If the emotion category is normal (S412: normal), a normal greeting such as "Hello" is generated as a response (message) (S414).
[0266] If the emotion category is anger (S412: anger), a response that calms the other person's emotions, such as "Are you offended?", is generated (S416). If the emotion category is joy (S412: joy), a greeting with a brighter nuance than a normal greeting, such as "I'm having a good time today," is generated (S418).
[0267] If the emotion category is confusion (S412: confusion), a greeting message showing concern for the other person, such as "What's wrong?", is generated as a response (S420). When this process is completed, the emotion response generation process ends.
[0268] [Effects of the twelfth embodiment] In the voice response system 100, the server 90 detects an unexpectedly uttered voice to detect irritation or agitation of the user, and generates a message to suppress the irritation or agitation.
[0269] According to the voice response system 100, when the user becomes irritated or upset, it is possible to suppress these, and therefore it is possible to suppress the user from vocalizing trouble between the user and those around him.
[0270] [Thirteenth embodiment] [Processing of the 13th embodiment] Next, a voice response system according to a thirteenth embodiment will be described. The voice response system according to the thirteenth embodiment performs a process of guiding a user to an object in a captured image. This process is performed by the server 90 as details of the process of S56 described above.
[0271] When a user vocally inputs something like "Please show me the way to the tower I can see" into the terminal device 1, the guidance process is executed in the process of S56. In the guidance process, as shown in Fig. 27, first, the terminal location information is acquired from the GPS receiver 27 of the terminal device 1 or the like (S432).
[0272] Then, based on the voice (text information) and image processing, the target object is identified from among the objects in the captured image, and its position is identified (S434). In this process, the position of the object is identified in map information (which may be acquired from an external source or may be held by the server 90) based on the shape, relative position, etc. of the object. For example, if a tower is captured in the captured image, the tower is identified on the map based on the position of the terminal device 1 and the shape of the tower.
[0273] Next, a route to this object is searched for (S436), and route information is acquired (S438). This process can be realized using a process similar to that in a well-known cloud-based navigation device.
[0274] Then, a response for providing route guidance is generated (S440). In this process, the response generated may be similar to the response provided by the navigation device. When this process is completed, the guidance process ends. When the guidance process is performed while the user is moving, the message can be played back using the automatic conversation server process, with the playback condition being that the user reaches the point to be guided.
[0275] [Effects of the thirteenth embodiment] In the voice response system 100, when text information is input, the server 90 generates a response according to a captured image of the surroundings of the voice response system 100, and outputs this response by voice.
[0276] According to the voice response system 100, a voice response can be output in response to a captured image, which improves usability compared to a configuration in which a response is generated only from text information.
[0277] Furthermore, in the voice response system 100, the server 90 searches for an object contained in the text information in the captured image by image processing, identifies the position of the object found, and provides guidance to the position of the object.
[0278] Such a voice response system 100 can guide the user to an object in a captured image. Furthermore, in the voice response system 100, when providing guidance to a destination, the server 90 acquires route information such as weather, temperature, humidity, traffic information, road surface conditions, etc. to the destination, and outputs the route information by voice.
[0279] According to such a voice response system 100, the status (route information) to the destination can be notified to the user by voice. [Modification of the thirteenth embodiment] In addition to the above configuration, text information may be input to respond as to what has been recognized, and what (someone) has been recognized from the captured image may be output by voice.
[0280] Furthermore, in the above voice response system 100, instead of the process of S48, the server 90 may acquire a moving image capturing the shape of the user's mouth when inputting text information by voice. In this case, instead of the process of S52, the server 90 may convert the voice into text information and correct the text information by estimating unclear parts of the voice based on the moving image.
[0281] According to such a voice response system 100, the content of the speech can be estimated from the shape of the mouth, and therefore unclear parts of the speech can be estimated well. [Fourteenth embodiment] [Processing of the 14th embodiment] Next, a voice response system according to a fourteenth embodiment will be described. The voice response system according to the fourteenth embodiment requests a user to perform a predetermined action and determines whether the user has performed the action as requested. In this configuration, in the automatic conversation terminal process shown in FIG. 6 and the automatic conversation server process shown in FIG. 7, movement request process 1 shown in FIG. 28 and movement request process 2 shown in FIG. 29 are sequentially performed as details of the process of S56 described above.
[0282] First, when the process of S54 is completed, movement request process 1 is started, and in movement request process 1, a response (message) instructing the user to move their line of sight or head to a predetermined position is output (S452), as shown in Fig. 28. When this process is completed, movement request process 1 is terminated.
[0283] Next, when the processing of S54 is completed, movement request processing 2 is started, and in movement request processing 2, it is determined whether the line of sight and head position have moved as instructed (S462), as shown in Fig. 29. In this processing, the user's movements are detected by processing images captured by the camera and using detection results from various sensors of the terminal device 1. When detecting the line of sight by image processing, well-known line of sight recognition technology may be employed.
[0284] If the gaze or head movement is not as instructed (S462: NO), the response generated in S452 is output again (S464). If the gaze or head movement is as instructed (S462: YES), another arbitrary response is generated (S466).
[0285] When this processing is completed, the movement request processing 2 is completed. [Effects of the 14th embodiment] In the voice response system 100, the line of sight of the user is detected, and if the user does not move his / her line of sight to a predetermined position in response to a call, a voice requesting the user to move his / her line of sight to the predetermined position is output.
[0286] According to the voice response system 100, the user can be made to look at a specific location, thereby ensuring safety confirmation when driving a vehicle. In the voice response system 100, the server 90 observes the positions of body parts and facial expressions, and if there is little change in response to the call, outputs a voice requesting that the positions of body parts or facial expressions be changed.
[0287] The voice response system 100 can move the position of a body part of the user to a specific position or induce the user to make a specific facial expression. The present invention can be used when driving a vehicle or during a physical examination.
[0288] [Fifteenth embodiment] [Processing of the 15th embodiment] Next, a voice response system according to a fifteenth embodiment will be described. In the voice response system according to the fifteenth embodiment, when a user inputs a broadcast program or a piece of music as voice, if the broadcast program or the piece of music is interrupted, a process for complementing the input is performed.
[0289] In this configuration, as a detailed example of S56 described above, the broadcast music supplementation process shown in Fig. 30 is performed. As shown in Fig. 30, the broadcast music supplementation process first determines whether the broadcast program or music (or the song if sung by the user) has been interrupted (S482).
[0290] If there is an interruption (S482: YES), the synchronized broadcast program or song is set as the response content in the processing of S492 described later (S484), and the broadcast song supplementation processing ends. Also, if there is no interruption (S482: NO), if a broadcast program is being watched, the broadcast program is acquired (S486), and if a song is being played, the corresponding song is acquired (S488).
[0291] Here, the karaoke DB 116 stores music pieces and lyrics in association with each other, and when a music piece is acquired in this process, a music piece with lyrics is acquired. Next, the broadcast program or music piece that the user is viewing is identified (S490). Then, this broadcast program or music piece is acquired and prepared for playback in synchronization with the broadcast program or music piece that the user is viewing (S492), and the broadcast music supplementation process is completed.
[0292] [Effects of the 15th embodiment] In the voice response system 100, the server 90 acquires a broadcast program similar to the broadcast program that the user is watching, and when the broadcast program is interrupted, the server 90 outputs the broadcast program that it has acquired to complement the interrupted broadcast program.
[0293] According to such a voice response system 100, it is possible to compensate for the interruption of a broadcast program being viewed by a user. In addition, in the voice response system 100, when a user sings a song without lyrics by adding lyrics, the server 90 compares the song with lyrics with the lyrics added by the user and outputs the lyrics by voice in the part where only the lyrics of the user are missing.
[0294] According to such a voice response system 100, it is possible to compensate for the part that a user of a so-called karaoke machine cannot sing (a part where the lyrics are broken). [16th embodiment] [Processing of the 16th embodiment] Next, a voice response system according to a sixteenth embodiment will be described. In the voice response system according to the sixteenth embodiment, when a captured image contains characters and the terminal device 1 receives a question from the user about how to read the characters, the voice response system acquires information about the characters from an external device and outputs the reading of the characters included in the information by voice.
[0295] In this configuration, as details of S56 described above, the character explanation process shown in Fig. 31 is performed. In the character explanation process, as shown in Fig. 31, first, it is determined whether or not a question about the reading, such as "how to read," has been received (S502). If a question about the reading has been received (S502: YES), the reading of the image-recognized character is searched for from other servers connected via the Internet network 85 (S504), and the obtained reading is set in the response (S506), thereby terminating the character explanation process.
[0296] If it is not a question about reading (S502: NO), then it is a question about "words" such as those in a Japanese dictionary. It is determined whether or not a question about the meaning of " has been received (S508). If a question about the meaning has been received, the meaning of the image-recognized character (word) is searched for from other servers connected via the Internet network 85 (S510), the obtained meaning is set as the response (S512), and the character explanation process is terminated.
[0297] [Effects of the 16th embodiment] According to such a voice response system 100, the reading of characters recognized by image recognition is searched from other servers, etc., and the obtained reading is set in the response, so that the user can be informed of how to read characters, the meaning of words, etc.
[0298] [17th embodiment] [Processing of the 17th embodiment] Next, a voice response system according to a seventeenth embodiment will be described. In the voice response system according to the seventeenth embodiment, the server 90 detects abnormal behavior or a state of a user of the terminal device 1 based on sensor values detected by the terminal device 1, and performs a process of reporting an abnormality if any.
[0299] In detail, the terminal device 1 performs the behavior response terminal processing shown in Fig. 32, and the server 90 performs the behavior response server processing. In the behavior response terminal processing, as shown in Fig. 32, first, outputs from various sensors mounted on the terminal device 1 are acquired (S522), and captured images are acquired by the camera 41 (S524). Then, the acquired outputs from the various sensors and the captured images are transmitted as packets to the server 90 (S526), and the behavior response terminal processing ends.
[0300] Next, in the behavior response server process, first, the processes of S42 to S44 described above are performed, as shown in Fig. 33. Next, behavior such as wandering is identified based on the location information of the terminal device 1 (detection result by the GPS receiver 27) (S532), and the user's environment is detected based on the detection result by the temperature sensors 15, 19, etc. (S534). Then, an abnormality is detected (S536).
[0301] In this process, an abnormality is detected based on changes in location information and the environment. For example, if the user is motionless in a hot or cold place, or if the user is in a place the user does not usually go, an abnormality is detected (S536). Alternatively, the location information and the environment are converted into a score, and if this score is below a reference value (outside the reference range), an abnormality is determined.
[0302] Next, it is determined whether an abnormality has been detected (S538). If no abnormality has been detected (S538: NO), the behavior response server processing ends. If an abnormality has been detected (S538: YES), a message indicating the abnormality is generated (S540) and notified to a predetermined contact point (S542). Then, the processing of S62 to S68 (excluding S66) described above is performed, and the behavior response server processing ends.
[0303] [Effects of the 17th embodiment] In the voice response system 100, the server 90 detects the user's behavior and the user's surrounding environment, and generates a message in accordance with the detected behavior and surrounding environment.
[0304] Such a voice response system 100 can notify dangerous places, restricted areas, etc. It can also detect abnormal behavior of the user.
[0305] Furthermore, in the voice response system 100, the server 90 receives an image of the user. Based on the image, a health condition is determined and a message is generated according to the health condition. Such a voice response system 100 makes it possible to manage the health condition of the user.
[0306] Furthermore, in the voice response system 100, the server 90 notifies a predetermined contact when the health condition falls below a reference value. According to the voice response system 100, when the health condition of the user is below a reference value, a notification can be issued, thereby making it possible to notify others of an abnormality at an earlier stage.
[0307] [Other embodiments] The present invention is not limited to the above-described embodiment, and various other forms may be adopted as long as they fall within the technical scope of the present invention.
[0308] For example, the voice response system 100 may mediate communication between two or more parties. In particular, when vehicles need to give way to each other at an intersection or the like, the terminal devices 1 may negotiate with each other as to which vehicle will enter the intersection first. In this case, each terminal device 1 transmits information about the direction of travel when approaching the intersection and the approach speed to the intersection to the server 90, and the server 90 sets priorities for each terminal device 1 according to the direction of travel and the approach speed, and generates and outputs voice messages such as "Stop" or "Enter" according to the priority.
[0309] Furthermore, when the terminal device 1 accepts an incoming call (incoming call) for communication that requires a real-time response, such as voice communication, the call may be accepted only when it is convenient for the user. Specifically, when the camera 41 captures an image of the user's face, the call may be accepted as being convenient for the user.
[0310] Furthermore, during voice communication, some people become annoyed if the other party does not answer when they call. To suppress such feelings, the user who is waiting for a response from the other party may be informed of the other party's status. For example, the terminal device 1 may manage the user's schedule, and if the user does not answer an incoming call, the terminal device 1 may search for free time in the user's schedule and inform the user of when the user will be able to answer.
[0311] In addition, if the user does not answer the call, the user's location may be reported to the caller. For example, if the user is connected to the Internet via a smartphone or PC, it is possible to know which device is being operated. This information can be used to identify the user's location and report it to the caller.
[0312] Furthermore, whether or not the user is available to answer an incoming call may be determined using location information using GPS or the like. Based on the location information, it is possible to determine whether or not the user is in a car, at home, etc. For example, if the user is on the move or in bed, it may be determined that the user is in a public place or asleep and is therefore unable to answer the incoming call. In this case, if the user is unable to answer the incoming call, it may be possible to inform the caller of what the user is doing, as described above.
[0313] Another possible configuration for obtaining location information is to use security cameras. In recent years, security cameras have been installed in various places, and so it is possible to use these security cameras to recognize the user's location by utilizing a configuration for identifying the user, such as face recognition. Security cameras can also be used to determine the situation, such as what the user is doing (whether or not they are in a position to answer the phone). Whether or not they are able to answer an incoming call can also be determined based on conditions such as whether or not they are using another landline phone (incoming calls cannot be answered while the landline phone is in use). Cut.
[0314] Furthermore, when a user of the terminal device 1 wants to have a conversation with someone, the results of learning the user's personality may be used to call a terminal device that is estimated to be a good match for the user among an unspecified number of users. In such a case, the terminal device may be configured to talk to the user about a topic that is likely to be exciting (a topic that both users are interested in (extracted using the learning results)).
[0315] Furthermore, when the voice response unit is not used for a long time (when the user has not spoken for a reference time or longer), the voice response unit may say something to the user. At this time, location information such as GPS may be used to select the words to be spoken.
[0316] [Relationship between each means described in the claims or the summary of the invention and the configurations in the embodiments] The terminal device 1 and the server 90 in the above embodiment correspond to an example of a voice response unit of the present invention. The processes of 22 and S56 in the above embodiment correspond to an example of a response acquisition means of the present invention.
[0317] Furthermore, the processes of S28, S60, and S64 in the above embodiment correspond to an example of a voice output means of the present invention, and the processes of S2 and S6 in the above embodiment correspond to an example of a voice input means of the present invention.
[0318] Furthermore, the process of S14 in the above embodiment corresponds to an example of a voice transmitting means of the present invention, and the reply candidate DB 105 in the above embodiment corresponds to an example of a reply recording means of the present invention.
[0319] Furthermore, the process of S56 in the above embodiment corresponds to an example of the personality information acquisition means of the present invention. Furthermore, the processes of S22 and S56 in the above embodiment correspond to an example of the response acquisition means of the present invention.
[0320] Furthermore, the processes of S28, S60, and S64 in the above embodiment correspond to an example of a voice output means of the present invention. Furthermore, the processes of S254, S258, and S260 in the above embodiment correspond to an example of a first personality information generating means and a second personality information generating means of the present invention. Furthermore, the process of S56 in the above embodiment corresponds to an example of a personality information acquiring means of the present invention.
[0321] Furthermore, the processes of S22 and S56 in the above embodiment correspond to the response acquisition means of the present invention, and the processes of S28, S60, and S64 in the above embodiment correspond to an example of the voice output means of the present invention.
[0322] Furthermore, the processes of S254, S258, and S260 in the above embodiment correspond to an example of a first personality information generating means and a second personality information generating means of the present invention. Furthermore, the processes of S48 and S56 in the above embodiment correspond to an example of a response generating means of the present invention, and the processes of S28, S60, and S64 in the above embodiment correspond to an example of a voice output means of the present invention.
[0323] Furthermore, the process of S48 in the modified example of the above embodiment corresponds to an example of the voice-input moving image obtaining means of the present invention, and the process of S52 in the above embodiment corresponds to an example of the character information converting means of the present invention.
[0324] Furthermore, the preference information generating process in the above embodiment is an example of the preference information generating means of the present invention. The process of S56 in the above embodiment corresponds to an example of a response candidate acquisition means of the present invention.
[0325] Furthermore, the motion character input process in the above embodiment corresponds to an example of the character information generating means of the present invention. Also, the other device information obtaining means, or other terminal use process in the above embodiment corresponds to an example of the transferring means of the present invention.
[0326] Furthermore, the process of S98 in the above embodiment corresponds to an example of a reproduction condition determining means of the present invention, and the process of S100 in the above embodiment corresponds to an example of a message reproducing means of the present invention.
[0327] Furthermore, the process of S116 in the above embodiment corresponds to an example of a no-answer sending means of the present invention, and the process of S372 in the above embodiment corresponds to an example of an utterance accuracy detecting means of the present invention.
[0328] Furthermore, the process of S374 in the above embodiment corresponds to an example of the accuracy output means of the present invention, and the process of S204 in the above embodiment corresponds to an example of the connection control means of the present invention.
[0329] Furthermore, the process of S50 in the above embodiment corresponds to an example of the emotion determination means of the present invention, and the process of S438 in the above embodiment corresponds to an example of the route information acquisition means of the present invention.
[0330] Furthermore, the process of S462 in the above embodiment corresponds to an example of the line-of-sight detecting means of the present invention, and the process of S464 in the above embodiment corresponds to an example of the line-of-sight movement request transmitting means of the present invention.
[0331] Furthermore, the process of S464 in the above embodiment corresponds to an example of a change request sending means of the present invention, and the process of S486 in the above embodiment corresponds to an example of a broadcast program acquisition means of the present invention.
[0332] Furthermore, the process of S484 in the above embodiment corresponds to an example of a broadcast program supplementation means and a lyric addition means of the present invention. Furthermore, the processes of S504 and S506 in the above embodiment correspond to an example of a pronunciation output means of the present invention. Furthermore, the processes of S522 and S524 in the above embodiment correspond to an example of a behavioral environment detection means of the present invention.
[0333] Furthermore, the process of S538 in the above embodiment corresponds to an example of the health condition determining means of the present invention, and the process of S540 in the above embodiment corresponds to an example of the health message generating means of the present invention.
[0334] Furthermore, the process of S542 in the above embodiment corresponds to an example of the reporting means of the present invention. [Explanation of symbols]
[0335] 1...terminal device, 10...behavior sensor unit, 11...3-dimensional acceleration sensor, 13...axial gyro sensor, 15...temperature sensor, 17...humidity sensor, 19...temperature sensor, 21...humidity sensor, 23...illuminance sensor, 25...wetness sensor, 27...GPS receiver, 29...wind speed sensor, 33...electrocardiogram sensor, 35...heartbeat sensor, 37...microphone, 39...memory, 41...camera, 50...communication unit, 53...wireless telephone unit, 55...contact memory, 60...alarm unit, 61...display, 63...illumination, 65...speaker, 70...operation unit, 71...touchpad, 73...confirmation button, 75...fingerprint sensor, 77...rescues request lever, 80...communication base station, 85...Internet network, 90...server, 100...voice response system, 101...calculation unit, 102...voice recognition DB, 103...predictive conversion DB, 104...voice DB, 105...response candidate DB, 106...personality DB, 107...learning DB, 108...preference DB, 109...news DB, 110...weather DB, 111...playback condition DB, 112...handwritten characters / sign language DB, 113...terminal information DB, 114...emotion determination DB, 115...health determination DB, 116...karaoke DB, 117...report destination DB, 118...sales DB, 119...client DB.
Claims
[Claim 1] A voice response system including a requesting device that is a request source, a providing device that is an information providing source, and a server that can communicate with the requesting device and the providing device, the providing device is configured to be able to set whether or not to permit information provision; an information acquisition unit configured in the server to acquire a request from the request source device, and if the request includes a request specifying the providing source device, to transmit the request to the specified providing source device, and if the information provision is permitted, to acquire the provided information provided by the providing source device; a providing unit configured to generate a voice response based on the provided information as a response to a request from the requesting device and provide the response to the requesting device; A voice response system comprising:
Citation Information
Patent Citations
Notifying device for position of automobile telephone set
JP1991120995A
Method for transferring information with equipment, equipment with interactive function applying the same and life support system formed by combining equipment
JP2001256036A
System, method and program for providing information
JP2002342356A
Response message generation apparatus, and terminal device thereof
JP2003108376A
Life adviser support system, adviser side terminal system, authentication server, server, support method and program
JP2009151766A