Voice response device server, vehicle, and voice response system
The voice response system addresses the lack of flexibility in response output by enabling multiple responses in different voice colors and reducing processing load through voice input and server-based generation, enhancing user interaction.
Patent Information
- Application Number
- JP2025069688
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2012-06-18
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2033-05-29
AI Technical Summary
Existing voice response systems typically provide a single, specified answer to a question, lacking user-friendly interaction and flexibility in response output.
A voice response system that includes a server capable of communicating with requester and provider devices, allowing multiple responses to be output in different voice colors, and can input voice information to reduce processing load on the device.
Enables multiple responses to be output in different voice colors, improving user usability and reducing processing load by allowing voice input and response generation on the server.
Smart Images

Figure 2025108674000001_ABST
Abstract
Description
Cross - reference to related applications
[0001] This international application claims priority based on Japanese Patent Application No. 2012 - 137065, Japanese Patent Application No. 2012 - 137066, and Japanese Patent Application No. 2012 - 137067, which were filed with the Japan Patent Office on June 18, 2012, and incorporates by reference the entire contents of Japanese Patent Application No. 2012 - 137065, Japanese Patent Application No. 2012 - 137066, and Japanese Patent Application No. 2012 - 137067 into this international application.
Technical Field
[0002] The present invention relates to a voice response system that causes responses to be made by voice.
Background Art
[0003] As the above - mentioned voice response device, there is known one that searches a dictionary for an answer to an input question and outputs the searched answer by voice (see, for example, Patent Document 1). Also, a technique for generating an answer to a question based on the content of the dialogue with the user is known (see, for example, Patent Document 2).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] In the above - mentioned technology, it is simply set to give one answer specified by a dictionary to one question. In a voice response system, making it more user - friendly for the user is one aspect of the present invention.
Means for Solving the Problems
[0006] In the invention of the first aspect, a voice response system including a requester device serving as a requester of a request, a provider device serving as an information provider, and a server capable of communicating with the requester device and the provider device, the provider device is configured to be able to set whether to permit information provision, the server acquires a request from the requester device, and when the request includes a request designating the provider device, transmits the request to the designated provider device, and when information provision is permitted, is configured with an information acquisition unit to acquire provision information provided from the provider device, a provision unit configured to generate a voice response based on the provision information as a response to the request from the requester device and provide the response to the requester device, and includes. In the invention of another aspect, a voice response device that causes a response to input character information to be made by voice, response acquisition means for acquiring a plurality of different responses to the character information, voice output means for outputting the plurality of different responses in different voice colors, and is characterized by including.
[0007] According to such a voice response device, since a plurality of responses can be output in different voice colors, even when the solution for one piece of character information cannot be specified uniquely, different solutions can be output to the user in an easy-to-understand manner in different voice colors. Therefore, it is possible to improve the usability for the user.
[0008] Note that the voice response device of the present invention may be configured as, for example, a terminal device possessed by a user, or may be configured as a server that communicates with this terminal device. Further, the character information may be input using an input means such as a keyboard, or may be input by converting voice into character information.
[0009] Incidentally, in the above voice response device, like the invention of the second aspect, a voice input means for the user to input voice, and converts the input voice into character information, a voice transmission means for transmitting to an external device that generates a plurality of different responses to the character information and transmits them to the voice response device, is provided, the response acquisition means may acquire the response from the external device. It may be like this.
[0010] According to such a voice response device, since the voice response device can input voice, it can be configured to input character information by voice. Also, since it can be configured to generate a response in the external device, the processing load on the voice response device can be reduced.
[0011] Note that in the voice transmission means, the operation of "converting the input voice into character information" may be performed by the voice response device or by the external device. Furthermore, in the above voice response device, like the invention of the third aspect, the voice response device or the external device is provided with response recording means in which a plurality of different responses including an affirmative response and a negative response to each of a plurality of character information are recorded, the response acquisition means acquires the affirmative response and the negative response as the plurality of different responses, the voice output means may reproduce them in different voice colors for the affirmative response and the negative response.
[0012] According to such a voice response device, responses with different positions such as an affirmative response and a negative response can be reproduced in different voice colors, so that the voice can be reproduced as if another person is speaking. Therefore, it is possible to make it difficult for the user listening to the voice to feel a sense of discomfort.
[0013] Note that the tone of voice may be changed according to the type of response and the choice of words in the response. For example, when responding in a gentle tone, it may be played back with a calm female voice, and when responding in an intense tone, it may be responded with a heroic male voice. That is, the response content and the personality may be associated with each other, and the tone of voice may be set according to the personality.
[0014] Further, in the above voice response device, like the invention of the fourth aspect, it can be configured to be used at a workplace or a company reception, or to be configured to convey on behalf of the user something that is difficult to say directly to someone.
[0015] When using the voice response device at the reception, the name and company name of the person coming for sales are pre-recorded in the voice response device or an external device, and when the person coming to the reception states this name or company name, the response may be generated so as to play back the voice of the rejection phrase. If so, the response may be generated so as to play back the voice of the rejection phrase.
[0016] Further, when configured to convey on behalf of something that is difficult to say, for example, before a date, if the user talks to this device saying that they want to say such and such today, the voice response device may talk (play back the voice) on behalf of the user at an appropriate timing (for example, a preset time, or when a certain time has elapsed since the conversation has stopped).
[0017] Alternatively, it may be configured to say words that trigger something difficult to say, for example, words such as "Come to think of it, didn't I say I had something to talk to her about?" That is, instead of immediately outputting a response, the response may be output when the playback condition is satisfied, such as after a certain time has elapsed.
[0018] Furthermore, in the above voice response device, as in the invention of the fifth aspect, the external device or the voice response device may obtain information for generating a response to character information from another voice response device. Also, in the above voice response device, as in the invention of the sixth aspect, when requested for information for generating a response to character information from another voice response device, it may return information corresponding to this request.
[0019] In this case, the voice response device may be provided with sensors for detecting position information, temperature, humidity, illuminance, noise level, etc., and databases such as dictionary information, and extract necessary information according to the request.
[0020] According to such a voice response device (external device), information for generating a response can be obtained from another voice response device. In this case, information unique to another voice response device, such as the position of another voice response device, can be obtained.
[0021] Also, its own unique information can be transmitted to another voice response device. Furthermore, in the above voice response device, as in the invention of the seventh aspect, a response (for example, an affirmative response or a negative response) output by itself or another voice response device may be input as character information, and a response for making a counterargument to this response may be generated. That is, from the user's perspective, discussions based on opinions from both the supportive and opposing positions can be heard. And after hearing this discussion, the user can make a final judgment.
[0022] This configuration can be realized using one or a plurality of voice response devices. In this case, for a plurality of voice response devices to exchange voices, the voices may be directly input and output, or communication by wireless or the like may be used.
[0023] Also, in the invention of the eighth aspect, a voice response device that makes a response to the input character information in voice, A personality information acquisition means for acquiring personality information in which the personality of a user or a person related to the user is associated according to a preset classification; A response acquisition means for acquiring response candidates representing a plurality of different responses to the character information; Voice output means for selecting a response to be output from the response candidates according to the personality information and outputting the selected response; It is characterized by comprising the above.
[0024] According to such a voice response device, different responses can be made according to the personality of the user or a person related to the user (related person). Therefore, it can be made more convenient for the user.
[0025] Also, in the above voice response device, as in the invention of the ninth aspect, It is provided with a first personality information generation means for generating the personality information of the user or the related person based on the answers to a plurality of preset questions; The personality information acquisition means acquires the personality information generated by the personality information generation means. It may be like this.
[0026] According to such a voice response device, personality information can be generated in the voice response device. When generating personality information, well-known personality analysis techniques (such as the Rorschach test, the Thematic Apperception Test, etc.) may be used. Also, when generating personality information, techniques for aptitude tests used by companies, etc. in employment tests may be used.
[0027] Furthermore, in the above voice response device, as in the invention of the tenth aspect, It is provided with a second personality information generation means for generating the personality information of the user or the related person based on the character string included in the input character information; The personality information acquisition means acquires the personality information generated by the personality information generation means. It may be like this.
[0028] According to such a voice response device, personality information can be generated in the process of the user using the voice response device. Also, in the above voice response device, like the invention of the 11th aspect, it is provided with preference information generation means for generating preference information indicating the tendency of the preference of the user or the related person based on the character string included in the character information. The voice output means selects a response to be output from the response candidates based on the preference information, and outputs the selected response. It may be like this.
[0029] According to such a voice response device, a response can be made according to the preference of the user or the related person. Furthermore, in the above voice response device, like the invention of the 12th aspect, the behavior of the user (conversation, place moved, things reflected in the camera) may be learned (recorded and analyzed) so as to supplement the lack of words in the user's conversation.
[0030] For example, for a conversation where the user answers "Curry would be nice." to a question "Is hamburger okay today?", if this device supplements with "Because it was hamburger yesterday.", the reason why the user said that curry is nice will be conveyed.
[0031] Also, such a configuration can be implemented during a phone call, and it may be configured to participate in the user's conversation arbitrarily. Furthermore, in the above voice response device, like the invention of the 13th aspect, response candidate acquisition means for acquiring response candidates from a predetermined server or the Internet. It may be provided.
[0032] According to such a voice response device, response candidates can be acquired not only from the device itself or an external device, but also from any device connected by the Internet, a dedicated line, or the like. Also, in the above voice response device, like the invention of the 14th aspect, character information generation means for converting the operation by the user into character information. It may be provided with.
[0033] Here, the operations referred to in the present invention include those resulting from muscle movements such as conversations, handwritten characters, or gestures (e.g., sign language). According to such a voice response device, the operations of the user can be converted into character information.
[0034] Furthermore, in the above voice response device, as in the invention of the 15th aspect, the character information generation means converts the voice from the user's speech into character information and accumulates the habits (such as pronunciation habits) during vocalization as learning information (capture the characteristics and record this feature). It may be done in this way.
[0035] According to such a voice response device, since character information can be generated based on the learning information, the generation accuracy of the character information can be improved. Also, in the above voice response device, as in the invention of the 16th aspect, transfer means for transferring the learning information to another voice response device, It may be provided with.
[0036] According to such a voice response device, even when the user uses another voice response device, the learning information recorded in the present voice response device can be used. Therefore, even when using another voice response device, the generation accuracy of the character information can be improved.
[0037] Furthermore, in the above voice response device, as in the invention of the 17th aspect, any of the user's actions and operations may be detected, and learning information or personality information may be generated based on these.
[0038] According to such a voice response device, for example, when it detects that the user has boarded a train continuously for several days, it can prompt the user to leave home a few minutes earlier from the next day, or when it detects that the user has a tendency to get angry easily from the conversation, it can output a voice or music to calm the mood.
[0039] Also, in the above voice response device, like the invention of aspect 18, it may be provided with other device information acquisition means for acquiring information recorded in another voice response device from another voice response device.
[0040] According to such a voice response device, a response can be generated based on the information recorded in another voice response device. Furthermore, in the above voice response device, like the invention of aspect 19, reproduction condition determination means for determining whether or not the situation of the voice response device matches a reproduction condition set in advance as a condition for outputting voice when the character information is not input, message reproduction means for outputting a preset message when the reproduction condition is met, may be provided.
[0041] According to such a voice response device, even when character information is not input (that is, when the user does not speak), voice can be output. For example, by forcing the user to speak, it can be used as a countermeasure against drowsiness during driving. Also, by determining whether a person living alone responds or not, safety confirmation can be performed.
[0042] Also, in the above voice response device, like the invention of aspect 20, The message reproduction means may acquire news information and output a message regarding the news in a question format asking for the user's answer. in such a manner.
[0043] According to such a voice response device, a conversation regarding news can be conducted, so It is also possible to suppress always having the same conversation. For example, when information regarding the stock price of a certain company is obtained as the content of the conversation, it can be something like "The stock price of Company XX has increased by XX yen today. Were you aware?"
[0044] Furthermore, in the above voice response device, like the invention of the 21st aspect, the voice output means or the message playback means may output by adding externally acquired information (such as news, environment (temperature, weather, location information, etc.)) to a preset message. This may be done.
[0045] According to such a voice response device, it is possible to output a response combining a predetermined message and the acquired information. Also, in the above voice response device, like the invention of the 22nd aspect, a plurality of messages may be acquired, and a message to be played back may be selected according to the playback frequency of the messages and output. This may be done.
[0046] According to such a voice response device, by making it difficult to play back messages with a high playback frequency, randomness can be achieved during message playback, or by deliberately repeating messages with a high playback frequency, attention can be drawn or memory can be strengthened.
[0047] Furthermore, in the above voice response device, like the invention of the 23rd aspect, when no response or answer to a message is obtained, an unanswered transmission means for transmitting information identifying the user and the fact that no answer was obtained to a preset contact destination may be provided.
[0048] According to such a voice response device, it is possible to report to the contact destination when no answer is obtained. Thus, for example, it is possible to report abnormalities of the elderly living alone at an early stage. Also, in the above voice response device, like the invention of the 24th aspect, The message playback means may store the conversation content and ask a question for obtaining the same content about the heard content (memory confirmation process). It may be like this.
[0049] According to such a voice response device, it is possible to confirm the user's memory and promote the fixation of memory. Furthermore, in the above voice response device, like the invention of the 25th aspect, Speech accuracy detection means for detecting the pronunciation and accent accuracy of the voice input by the user, Accuracy output means for outputting the detected accuracy, may be provided.
[0050] According to such a voice response device, it is possible to confirm the pronunciation and accent accuracy. For example, it is effective when practicing a foreign language. Also, in the above voice response device, like the invention of the 26th aspect, The accuracy output means may output a voice including the closest word when the accuracy is below a certain value. It may be like this.
[0051] According to such a voice response device, the user can confirm the pronunciation and accent accuracy. Furthermore, in the above voice response device, like the invention of the 27th aspect, The message playback means may output the same question again when the accuracy is below a certain value.
[0052] According to such a voice response device, an accurate answer can be obtained by outputting the same question. Also, in the above voice response device, like the invention of the 28th aspect, Connection control means for identifying a communication partner based on the input character information and connecting the preset communication destination and the communication partner for each communication partner, may be provided.
[0053] According to such a voice response device, reception work and telephone response can be assisted. In particular, in the above voice response device, like the invention of the 29th aspect, the connection control means may identify sales activities (sales), visitors, and play a rejection message if it is a sales activity. It may be like this.
[0054] According to such a voice response device, a person who may interfere with the user's work can be excluded without the user having to handle it. Furthermore, in the above voice response device, like the invention of the 30th aspect, keywords included in the input character information (especially voice) may be extracted, and the connection may be made to the connection destination corresponding to the keyword. For example, the keywords such as the name of the other party and its connection destination may be associated in advance.
[0055] According to such a voice response device, operations such as telephone transfer and reception call can be assisted. Also, in the above voice response device, like the invention of the 31st aspect, the requirements for the other party to speak may be recognized based on the keyword, and the outline of what the other party said may be conveyed to the user.
[0056] According to such a voice response device, the business of intermediating with customers can be assisted. Furthermore, in the above voice response device, like the invention of the 32nd aspect, for the voice input by the user, an emotion determination means that reads the emotion from the voice tone and outputs which emotion among the emotions including at least one of anger, joy, confusion, sadness, and excitement the voice corresponds to may be provided.
[0057] According to such a voice response device, a response can be output according to the emotion of the user. Next, the invention of the 33rd aspect is Response generation means for generating a response according to a captured image that captures the surroundings of the voice response device when the text information is input; Voice output means for outputting the response as voice; It is characterized by comprising the above.
[0058] According to such a voice response device, a response can be output as voice according to a captured image. Therefore, the usability can be improved as compared with a configuration that generates a response only from text information.
[0059] As a specific configuration of the present invention, for example, there is a configuration in which text information is input so as to respond to what is recognized, and what (who) is recognized from a captured image is output as voice.
[0060] By the way, in the above voice response device, as in the invention of the 34th aspect, Position specifying means search means for searching for an object included in text information from a captured image by image processing and specifying the position of the searched object; Guiding means for guiding to the position of the object; It may be provided with the above.
[0061] According to such a voice response device, the user can be guided to an object in a captured image. Furthermore, in the above voice response device, as in the invention of the 35th aspect, Voice input moving image acquisition means for acquiring a moving image that captures the shape of the user's mouth when text information is input as voice ; Character information conversion means for converting the voice into character information and correcting the character information by estimating an unclear part of the voice based on the moving image; It may be provided with the above.
[0062] According to such a voice response device, the pronunciation content can be estimated from the shape of the mouth, so an unclear part of the voice can be estimated well. Also, in the above voice response device, as in the invention of the 36th aspect, The message playback means may detect the user's impatience or agitation by detecting an unexpected voice, and generate a message for suppressing the impatience or agitation. It may be like this.
[0063] According to such a voice response device, when the user is impatient or agitated, these can be suppressed. Therefore, the occurrence of troubles between the user and the surroundings can be suppressed. Furthermore, in the above voice response device, like the invention of the 37th aspect, when guiding to a destination, it is provided with route information acquisition means for acquiring route information such as weather, temperature, humidity, traffic information, road surface conditions, etc. to the destination, The message playback means may output the route information by voice. It may be like this.
[0064] According to such a voice response device, the situation (route information) to the destination can be notified to the user by voice. Also, in the above voice response device, like the invention of the 38th aspect, a line-of-sight detection means for detecting the user's line of sight, a line-of-sight movement request transmission means for outputting a voice requesting to move the line of sight to a predetermined position when the user's line of sight does not move to the predetermined position in response to the call by the message playback means, may be provided.
[0065] According to such a voice response device, the user can be made to look at a specific position. Therefore, safety confirmation during vehicle driving can be surely performed. Note that in the above voice response device, like the invention of the 39th aspect, a change request transmission means for observing the position of the body part and the facial expression, and outputting a voice requesting to change the position of the body part and the facial expression when there is little change in response to the call may be provided.
[0066] According to such a voice response device, it is possible to move the position of a part of the user's body to a specific position or to induce the user to make a specific expression. The present invention can be used during vehicle driving or physical examinations.
[0067] Furthermore, in the above voice response device, as in the invention of aspect 40, broadcast program acquisition means for acquiring a broadcast program similar to the broadcast program viewed by the user, broadcast program complementing means for complementing the interrupted broadcast program by outputting the broadcast program acquired by itself when the broadcast program is interrupted, may be provided.
[0068] According to such a voice response device, it is possible to supplement so that the broadcast program viewed by the user is not interrupted. Also, in the above voice response device, as in the invention of aspect 41, when the user sings by attaching lyrics to a song without lyrics, comparing the song with lyrics and the lyrics attached by the user, and lyric addition means for outputting the lyrics in voice at the part where only the user's lyrics are missing.
[0069] According to such a voice response device, it is possible to supplement the part that the user cannot sing (the part where the lyrics are interrupted) in so-called karaoke. Furthermore, in the above voice response device, as in the invention of aspect 42, when characters are included in the captured image, if a question about the reading of these characters is received from the user, reading output means for acquiring the information of these characters from the outside and outputting the reading of the characters included in this information in voice. may be provided.
[0070] According to such a voice response device, it is possible to teach the user the reading of the characters. Also, in the above voice response device, as in the invention of aspect 43, it is provided with action environment detection means for detecting the user's actions and the user's surrounding environment, The message generation means may generate a message according to the detected actions and the surrounding environment. It may be like this.
[0071] According to such a voice response device, it is possible to notify a dangerous place, a restricted area, etc. Further, it is possible to detect that the user has abnormal behavior or the like. Furthermore, in the above voice response device, like the invention of the 44th aspect, a health state determination means for determining a health state based on a captured image of the user, and a health message generation means for generating a message according to the health state, may be provided.
[0072] According to such a voice response device, the health state of the user can be managed. Also, in the above voice response device, like the invention of the 45th aspect, a notification means for making a notification to a predetermined contact when the health state falls below a reference value, may be provided.
[0073] According to such a voice response device, when the health state of the user is below the reference value, a notification can be made. Therefore, an abnormality can be notified to others earlier. Furthermore, in the above voice response device, like the invention of the 46th aspect, information about the user may be output in response to an inquiry from a person other than the user.
[0074] According to such a voice response device, for example, if the diet content of the user and the walking distance are detected, it is possible to answer questions at a hospital or the like on behalf of the user. Also, it may be learned about the health state, self-introduction, etc.
[0075] Note that the invention of each aspect does not need to be based on other inventions and can be made as independent an invention as possible. It can be done.
Brief Description of the Drawings
[0076]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Figure 33
Mode for Carrying Out the Invention
[0077] Embodiments of the present invention will be described below with reference to the drawings. [First Embodiment] [Configuration of This Embodiment] The voice response system 100 to which the present invention is applied is a system configured to generate an appropriate response at the server 90 for the voice input at the terminal device 1 and output the response as voice at the terminal device 1. Specifically, as shown in FIG. 1, the voice response system 100 is configured such that a plurality of terminal devices 1 and a server 90 can communicate with each other via a communication base station 80 or the Internet network 85.
[0078] The server 90 has functions as a normal server device. In particular, the server 90 includes an arithmetic unit 101 and various databases (DBs). The arithmetic unit 101 is configured as a well-known arithmetic device including a CPU and memories such as a ROM and a RAM, and based on the programs in the memories, communicates with the terminal device 1 etc. via the Internet network 85, reads and writes data in various DBs, or performs various processes such as voice recognition and response generation for conversation with the user using the terminal device 1.
[0079] As shown in FIG. 1, the various DBs include a voice recognition DB 102, a prediction conversion DB 103, a voice DB 104, a response candidate DB 105, a personality DB 106, a learning DB 107, a preference DB 108, a news DB 109, a weather DB 110, a reproduction condition DB 111, a handwritten character / sign language DB 112, a terminal information DB 113, an emotion determination DB 114, a health determination DB 115, a karaoke DB 116, a notification destination DB 117, a sales DB 118, a client DB 119, etc. The details of these DBs will be described each time the process is explained.
[0080] Next, as shown in FIG. 2, the terminal device 1 is configured to include an action sensor unit 10, a communication unit 50, a notification unit 60, and an operation unit 70 in a predetermined housing. The motion sensor unit 10 includes a well-known MPU31 (microprocessor unit), a memory 39 such as a ROM and a RAM, and various sensors. The MPU31 performs processes such as driving a heater for optimizing the temperature of the sensor elements so that the sensor elements constituting the various sensors can detect the inspection target (humidity, wind speed, etc.) well.
[0081] The motion sensor unit 10 includes, as various sensors, a three-axis acceleration sensor 11 (3DG sensor), a three-axis gyro sensor 13, a temperature sensor 15 arranged on the back surface of the housing, a humidity sensor 17 arranged on the back surface of the housing, a temperature sensor 19 arranged on the front surface of the housing, a humidity sensor 21 arranged on the front surface of the housing, an illuminance sensor 23 arranged on the front surface of the housing, a wetness sensor 25 arranged on the back surface of the housing, a GPS receiver 27 for detecting the current location of the terminal device 1, and a wind speed sensor 29.
[0082] In addition, the motion sensor unit 10 also includes an electrocardiogram sensor 33, a heart sound sensor 35, a microphone 37, and a camera 41 as various sensors. Note that each of the temperature sensors 15, 19 and each of the humidity sensors 17, 21 measure the temperature or humidity of the outside air of the housing as the inspection target.
[0083] The three-axis acceleration sensor 11 detects accelerations in three mutually perpendicular directions (vertical direction (Z direction), width direction (Y direction) of the housing, and thickness direction (X direction) of the housing) applied to the terminal device 1, and outputs the detection results.
[0084] The three-axis gyro sensor 13 detects angular accelerations (with the leftward speeds in each direction being positive) in the vertical direction (Z direction) and two arbitrary directions (width direction (Y direction) of the housing and thickness direction (X direction) of the housing) orthogonal to the vertical direction as angular velocities applied to the terminal device 1, and outputs the detection results.
[0085] The temperature sensors 15 and 19 are configured to include, for example, thermistor elements whose electrical resistance changes according to temperature. In this embodiment, the temperature sensors 15 and 19 detect Celsius temperature, and all temperature displays described in the following explanations are in Celsius temperature.
[0086] The humidity sensors 17 and 21 are configured as, for example, well-known polymer film humidity sensors. This polymer film humidity sensor is configured as a capacitor in which the amount of moisture contained in the polymer film changes according to a change in relative humidity, and the dielectric constant changes.
[0087] The illuminance sensor 23 is configured as, for example, a well-known illuminance sensor including a phototransistor. The wind speed sensor 29 is, for example, a well-known wind speed sensor that calculates the wind speed from the electric power (heat dissipation amount) required to maintain the heater temperature at a predetermined temperature.
[0088] The heart sound sensor 35 is configured as a vibration sensor that captures vibrations caused by the pulsation of the user's heart. The MPU 31 discriminates between vibrations and noises caused by pulsation and other vibrations and noises in view of the detection result by the heart sound sensor 35 and the heart sound input from the microphone 37.
[0089] The wetness sensor 25 detects water droplets on the surface of the housing, and the electrocardiogram sensor 33 detects the user's pulse. The camera 41 is arranged within the housing of the terminal device 1 so as to take the outside of the terminal device 1 as the imaging range.
[0090] The communication unit 50 includes a well-known MPU 51, a wireless telephone unit 53, and a contact memory 55, and is configured to be able to acquire detection signals from various sensors that constitute the motion sensor unit 10 via an input / output interface (not shown). Then, the MPU 51 of the communication unit 50 executes processing according to the detection results by this motion sensor unit 10, input signals input via the operation unit 70, and programs stored in a ROM (not shown).
[0091] Specifically, the MPU 51 of the communication unit 50 executes functions as an operation detection device that detects specific operations performed by the user, a positional relationship detection device that detects the positional relationship with the user, a motion load detection device that detects the load of the motion performed by the user, and a function of transmitting the processing result by the MPU 51.
[0092] The radio phone unit 53 is configured to be communicable with, for example, the base station of a mobile phone. The MPU 51 of the communication unit 50 outputs the processing result by the MPU 51 to the notification unit 60 or transmits it to a preset transmission destination via the radio phone unit 53.
[0093] The contact memory 55 functions as a storage area for storing the location information of the user's destination. Information of contacts (such as phone numbers) to be contacted in case of an abnormality of the user is recorded in this contact memory 55.
[0094] The notification unit 60 includes, for example, a display 61 configured as an LCD or an organic EL display, a lighting fixture 63 composed of LEDs that can emit light in, for example, seven colors, and a speaker 65. Each part constituting the notification unit 60 is driven and controlled by the MPU 51 of the communication unit 50.
[0095] Next, the operation unit 70 includes a touch pad 71, a confirmation button 73, a fingerprint sensor 75, and a rescue request lever 77. The touch pad 71 outputs a signal according to the position and pressure touched by the user (the user, the user's guardian, etc.).
[0096] The confirmation button 73 is configured such that when pressed by the user, the contacts of the built-in switch close, and the communication unit 50 can detect that the confirmation button 73 has been pressed.
[0097] The fingerprint sensor 75 is a well-known fingerprint sensor and is configured to be able to read fingerprints using, for example, an optical sensor. Note that instead of the fingerprint sensor 75, any means capable of recognizing human physical characteristics (means capable of performing biometric authentication: means capable of identifying an individual), such as a sensor for recognizing the shape of the veins in the palm, can be adopted.
[0098] Also, it is provided with a rescue request lever 77 that is connected to a predetermined contact when operated. [Processing of this Embodiment] The processing implemented in such a voice response system 100 will be described below.
[0099] The voice response terminal processing implemented in the terminal device 1 is a process of receiving voice input by the user, sending this voice to the server 90, and when receiving a response to be output from the server 90, playing back this response as voice. Note that this process is started when the user inputs an instruction to perform voice input via the operation unit 70.
[0100] Specifically, as shown in FIG. 3, first, it is set to the state (ON state) of receiving input from the microphone 37 (S2), and imaging (recording) by the camera 41 is started (S4). Then, it is determined whether there is voice input (S6).
[0101] If there is no voice input (S6: NO), it is determined whether a timeout has occurred (S8). Here, a timeout indicates that the allowable time during which the process waits has been exceeded. Here, the allowable time is set to about 5 seconds, for example.
[0102] If a timeout has occurred (S8: YES), the process proceeds to the process of S30 described later. If no timeout has occurred (S8: NO), the process returns to the process of S6. If there is voice input (S6: YES), the voice is recorded in the memory (S10), and it is determined whether the voice input has ended (S12). Here, it is determined that the voice input has ended when the voice has been interrupted for a certain period of time or when an instruction to end the voice input is input via the operation unit 70.
[0103] If the voice input has not ended (S12: NO), the process returns to the process of S10. Also, if the voice input has ended (S12: YES), data such as the ID for identifying itself, the voice, and the captured image is packet-transmitted to the server 90 (S14). Note that the process of transmitting the data may be performed between S10 and S12.
[0104] Subsequently, it is determined whether the data transmission has been completed (S16). If the transmission has not been completed (S16: NO), the process returns to the process of S14. Also, if the transmission has been completed (S16: YES), it is determined whether the data (packet) transmitted in the voice response server process described later has been received (S18). If the data has not been received (S18: NO), it is determined whether a timeout has occurred (S20).
[0105] If a timeout has occurred (S20: YES), the process proceeds to the process of S30 described later. Also, if a timeout has not occurred (S20: NO), the process returns to the process of S18. Also, if the data has been received (S18: YES), the packet is received (S22). In this process, one or more different responses to the character information are acquired, each associated with a different voice color.
[0106] Then, it is determined whether the reception has been completed (S24). If the reception has not been completed (S24: NO), it is determined whether a timeout has occurred (S26). If a timeout has occurred (S26: YES), an error occurrence is output via the notification unit 60, and the voice response terminal process ends. Also, if a timeout has not occurred (S26: NO), the process returns to the process of S22.
[0107] Also, if the reception is completed (S24: YES), a response based on the received packet is output from the speaker 65 as voice (S28). In this process, when playing multiple responses, the multiple responses are played in different voice colors. When such a process ends, the voice response terminal process ends.
[0108] Subsequently, the voice response server process executed by the server 90 (external device) will be described with reference to FIG. 4. The voice response server process is a process of receiving voice from the terminal device 1, performing voice recognition to convert this voice into character information, and generating a response to the voice and returning it to the terminal device 1. In particular, in this embodiment, there may be a case where a plurality of responses are transmitted in association with voices of different voice colors.
[0109] As details of the voice response server process, as shown in FIG. 4, first, it is determined whether a packet from any terminal device 1 has been received (S42). If no packet has been received (S42: NO), the process of S42 is repeated.
[0110] Also, if a packet has been received (S42: YES), the terminal device 1 of the communication partner is specified (S44). In this process, the terminal device 1 is specified by the ID of the terminal device 1 included in the packet.
[0111] Subsequently, the voice included in the packet is recognized (S46). Here, in the voice recognition DB 102, a large number of voice waveforms and a large number of characters are associated with each other. Also, in the prediction conversion DB 103, words that are likely to be used following a certain word are associated with each other.
[0112] Therefore, in this process, by referring to the voice recognition DB 102 and the prediction conversion DB 103, a well-known voice recognition process is performed to convert the voice into character information. Subsequently, by performing image processing on the captured image, the object in the captured image is specified (S48). Then, based on the waveform of the voice, the end of the word, etc., the emotion of the user is determined (S50).
[0113] In this process, by referring to the emotion determination DB114 in which the waveform (tone color) of the voice, the endings of words, etc. are usually associated with the classification of emotions such as anger, joy, confusion, sadness, excitement, etc., it is determined whether the user's emotion corresponds to any of the classifications, and the determination result is recorded in the memory. Subsequently, by referring to the learning DB107, the words frequently spoken by this user are searched, and the ambiguous parts in the character information generated by voice recognition are corrected.
[0114] Note that in the learning DB107, the characteristics of the user, such as the words frequently spoken by the user and the habits during pronunciation, are recorded for each user. Also, data is added to and modified in the learning DB107 during the conversation with the user.
[0115] Subsequently, the corrected character information is specified as the input character information (S54), and a response is obtained from the response candidate DB105 by searching the response candidate DB105 for a sentence similar to the character information as the input (S56). Here, in the response candidate DB105, as shown in FIG. 5, the input character information, the first output, the tone color of the first output, the second output, and the tone color of the second output are uniquely correspondingly associated.
[0116] For example, as shown in the first row of FIG. 5, when the character information "Today's ※ weather" is input, the first output "Today's ※ weather is ※" is output in association with the tone color of female 1. However, the part of "※" is obtained by accessing the weather DB110 in which the region name and the weather forecast for several days in that region are associated.
[0117] Also, when the character information "Today's ※ weather" is input, the weather at the timing when today's weather changes is also obtained from the weather DB110, and the second output "However, ※ is ※" is output in association with the tone color of male 1. When "Today's Tokyo weather" is input in the case where today's weather in Tokyo is sunny and tomorrow's weather is rainy, it is output in the tone color of female 1 as "Today's Tokyo weather is sunny.", and in the tone color of male 1 as "However, tomorrow will be rainy.".
[0118] In this embodiment, a case where a plurality of responses are output has been described. However, when there is only one response to the input, there will be only one response. Therefore, it is determined whether the number of responses is one (S58). If there is only one response (S58: YES), the process proceeds to S62, which will be described later.
[0119] On the other hand, if there are a plurality of responses (S58: NO), the response content and the tone and voice color are associated with each other (S60). Here, in the voice database 104, a database of artificial voice data is stored for each tone and voice color. In this process, the tone and voice color set for each response are associated with the tone and voice color in the database.
[0120] Subsequently, the response content is converted into voice (S62). In this process, based on the database stored in the voice database 104, a process of outputting the response content (character information) as voice is performed.
[0121] Then, the generated response (voice) is packet-transmitted to the terminal device 1 of the communication partner (S64). Note that the packet transmission may be performed while generating the voice of the response content. Subsequently, the conversation content is recorded (S68). In this process, the input character information and the output response content are recorded in the learning database 107 as the conversation content. At this time, keywords (words recorded in the voice recognition database 102) and pronunciation features included in the conversation content are recorded in the learning database 107.
[0122] When such a process is completed, the voice response server process is terminated. [Effect of this embodiment] As described in detail above, the voice response system 100 is a system that makes a voice response to the input character information. The terminal device 1 (MPU 31) acquires a plurality of different responses to the character information and outputs the plurality of different responses in different tones and voice colors.
[0123] According to such a voice response system 100, since a plurality of responses can be output with different voices, even when the solution for one piece of character information cannot be specified uniquely, different solutions can be output to the user clearly with different voices. Therefore, it is possible to improve the usability for the user.
[0124] Also, in the voice response system 100, the terminal device 1 inputs the voice of the user via the microphone 37, and the server 90 (the arithmetic unit 101) converts the input voice into character information, generates a plurality of different responses for the character information, and transmits them to the terminal device 1. Then, the terminal device 1 acquires the response from the server 90.
[0125] According to such a voice response system 100, since the terminal device 1 can input voice, it can be configured to input character information by voice. Also, since it can be configured to generate a response in the server 90, the processing load in the voice response system 100 can be reduced.
[0126] Furthermore, in the voice response system 100, the server 90 converts the voice from the user's utterance into character information, and accumulates the habits at the time of speaking (such as pronunciation habits) as learning information (capture the characteristics and record this characteristic).
[0127] According to such a voice response system 100, since character information can be generated based on the learning information, the generation accuracy of the character information can be improved. Furthermore, in the voice response system 100, the server 90 reads the emotion from the voice color for the voice input by the user, and outputs which emotion among the emotions usually including at least one of anger, joy, confusion, sadness, and excitement the voice corresponds to.
[0128] According to such a voice response system 100, a response can be output according to the emotion of the user. [Modification Example of the First Embodiment] In the present embodiment, voice recognition is used as a configuration for inputting character information. However, not limited to voice recognition, it may be input using input means (operation unit 70) such as a keyboard or a touch panel. Also, the operation of "converting the input voice into character information" was performed by the server 90, but it may be performed by the terminal device 1.
[0129] Furthermore, in the voice response system 100, the server 90 is provided with a response candidate DB 105 in which a plurality of different responses including an affirmative response and a negative response for each of the plurality of character information are recorded. The terminal device 1 may acquire an affirmative response and a negative response as a plurality of different responses and reproduce them in different voice colors for the affirmative response and the negative response.
[0130] For example, as shown in the second row in FIG. 5, when a voice such as "May I buy something?" is input, for this item, affirmative information such as a good reputation is output by associating it with a female voice. On the other hand, negative information such as a bad reputation is output in a voice color different from that of the female voice associated with the affirmative information (here, a male voice).
[0131] According to such a voice response system 100, responses with different positions such as an affirmative response and a negative response can be reproduced in different voice colors, so that the voice can be reproduced as if another person is speaking. Therefore, it is possible to make it difficult for the user listening to the voice to feel a sense of discomfort.
[0132] Note that the voice color may be changed according to the type of response or the diction at the time of response. For example, when responding in a gentle tone, it may be reproduced with a calm female voice, and when responding in an intense tone, it may be responded with a brave male voice, etc. That is, the response content and the personality may be associated with each other, and the voice color may be set according to the personality.
[0133] Furthermore, in the voice response system 100, a response (for example, an affirmative response or a negative response) output from its own terminal device 1 or another terminal device 1 may be input as character information, and a response for making a counterargument to this response may be generated. That is, from the user's perspective, it is possible to hear discussions based on both opinions of the supportive position and the opposing position. And after hearing this discussion, the user can make a final judgment.
[0134] This configuration can be realized by using one or a plurality of terminal devices 1. In order for a plurality of terminal devices 1 to exchange voices with each other, the voices may be directly input and output, or communication such as wireless communication may be used. When a plurality of terminal devices 1 communicate with the server 90, in the process of S66, it is only necessary to transmit data to another terminal device 1.
[0135] Furthermore, in the voice response system 100, the arithmetic unit 101 may learn (record and analyze) the actions of the user (conversation, the place where the user has moved, what is shown in the camera), and supplement the lack of words in the user's conversation.
[0136] For example, for a conversation where the user answers "Curry would be good." to a question "Is hamburger okay today?", if this device supplements with "Because we had hamburger yesterday.", the reason why the user said that curry would be good is conveyed.
[0137] Also, such a configuration can be implemented during a phone call, or it may be configured to arbitrarily participate in the user's conversation. Furthermore, in the voice response system 100, the server 90 may acquire response candidates from a predetermined server or from the Internet.
[0138] According to such a voice response system 100, response candidates can be acquired not only from the server 90 but also from any device connected by the Internet, a dedicated line, or the like. [Second Embodiment] [Processing of the Second Embodiment] Next, a different form of the voice response system will be described. In the following embodiments (the second embodiment), only the parts different from the voice response system 100 of the first embodiment will be described in detail, and for the parts similar to the voice response system 100 of the first embodiment, the same reference numerals will be used and the description will be omitted.
[0139] In the voice response system of the second embodiment, even when the user does not input character information, voice is output. Specifically, the terminal device 1 performs the automatic conversation terminal process shown in FIG. 6. The automatic conversation terminal process is a process that starts, for example, when the power of the terminal device 1 is turned on, and is a process that is repeatedly executed thereafter.
[0140] In the automatic conversation terminal process, first, it is determined whether the setting for having an automatic conversation is ON (S82). Note that whether to perform an automatic conversation can be set by the user via the operation unit 70 or by inputting voice.
[0141] If the setting for having an automatic conversation is OFF (S82: NO), the automatic conversation terminal process ends. Also, if the setting for having an automatic conversation is ON (S82: YES), the fact that the automatic conversation mode has been set is transmitted to the server 90 together with the ID for identifying itself (S84).
[0142] Subsequently, it is determined whether a packet has been received from the server 90 (S86). If no packet has been received (S86: NO), the process of S86 is repeated. Also, if a packet has been received (S86: YES), the same processes as those of S22 to S30 described above are performed, and when these processes are completed, the automatic conversation terminal process ends.
[0143] Also, the server 90 executes the automatic conversation server process shown in FIG. 7. The automatic conversation server process starts, for example, when the power of the server 90 is turned on, and is a process that is repeatedly executed thereafter.
[0144] In the automatic conversation server process, first, it is determined whether or not it has received from the terminal device 1 that it has been set to the automatic conversation mode (S92). If it has not received that it has been set to the automatic conversation mode (S92: NO), the process proceeds to the process of S98.
[0145] If it has received that it has been set to the automatic conversation mode (S92: YES), the terminal device 1 that becomes the communication partner is specified based on the ID included in the received packet (S94), and it is set that an automatic conversation will be made with this communication partner (S96). Subsequently, for each of the terminal devices 1 for which it has been set that an automatic conversation will be made, it is determined whether or not the reproduction conditions are satisfied (S98).
[0146] Here, the reproduction conditions indicate, for example, that a certain amount of time has elapsed since the previous conversation (voice input), a certain fixed time of the day, when the weather is specific, when any of the sensor values indicates an abnormal value, and so on.
[0147] If the reproduction conditions are not satisfied (S98: NO), the automatic conversation server process ends. Also, if the reproduction conditions are satisfied (S98: YES), a message corresponding to the reproduction conditions is generated (S100).
[0148] Here, the message corresponding to the reproduction conditions may be a fixed phrase such as "Good morning." or "Hello.", or may be related to the latest news obtained from the news DB 109 in which the latest news is automatically updated. When a message related to the latest news is used, for example, when information regarding the stock price of a certain company can be obtained, it can be something like "The stock price of ○○ company has increased by ○○ yen today. Did you know?"
[0149] When this process ends, the processes of S42 to S54 described above are executed. Then, when the process of S54 ends, it is determined whether a predetermined response has been obtained from the terminal device 1 that is the communication partner (S112). Here, the predetermined response may be, for example, some kind of voice or a specific answer. The specific answer means, for example, for a question such as "Do you know?" the answers "I know" or "I don't know" are applicable, and for a question such as "What's the weather like now?" those including words indicating the weather such as "It's raining" or "It's sunny" are applicable.
[0150] If there is a predetermined response (S112: YES), the automatic conversation server process ends. Also, if there is no predetermined response (S112: NO), the message transmitted at S100 is resent (S114). When resending the message in this way, the tone of voice is changed, and a voice with a strong and severe tone is generated.
[0151] Subsequently, the notification destination DB117 in which the terminal device 1 and the notification destination are associated in advance is referred to, and it is transmitted to the predetermined notification destination that there was no answer (S116). When such a process ends, the automatic conversation server process ends.
[0152] [Effect according to the second embodiment] In the above voice response system 100, the server 90 determines whether the situation of the voice response system 100 matches the reproduction condition set in advance as a condition for outputting voice when character information is not input. And when it matches the reproduction condition, a preset message is output.
[0153] According to such a voice response system 100, even when character information is not input (that is, when the user does not talk), voice can be output. For example, by forcibly making the user speak, it can be used for measures to suppress drowsiness during driving. Also, by determining whether a person living alone responds, safety confirmation can be performed.
[0154] Also, in the voice response system 100, the server 90 acquires news information and outputs a message regarding the news in a question format asking for the user's answer. According to such a voice response system 100, since conversations regarding news can be conducted, it is possible to suppress always having the same conversations.
[0155] Furthermore, in the voice response system 100, the server 90 adds externally acquired information (such as news, environment (temperature, weather, location information, etc.)) separately acquired to a preset message and outputs it.
[0156] According to such a voice response system 100, it is possible to output a response combining a predetermined message and the acquired information. Furthermore, in the voice response system 100, when no answer to a response or message is obtained, the server 90 transmits information identifying the user and the fact that no answer was obtained to a preset contact destination.
[0157] According to such a voice response system 100, it is possible to report to the contact destination when no answer is obtained. Therefore, for example, it is possible to report an abnormality of an elderly person living alone at an early stage.
[0158] [Modification Example of the Second Embodiment] Also, in the voice response system 100, the server 90 may acquire a plurality of messages and select and output a message to be reproduced according to the reproduction frequency of the messages.
[0159] According to such a voice response system 100, by making it difficult to reproduce a message with a high reproduction frequency, it is possible to achieve randomness when reproducing messages, or to promote attention and memory retention by repeatedly reproducing a message with a high reproduction frequency deliberately.
[0160] [Third Embodiment] [Processing of the Third Embodiment] Next, in the voice response system of the third embodiment, the terminal device 1 is configured to convey on behalf of the user something that is difficult to say directly to someone. For example, before a date, if the user talks to this device saying that they want to say such and such today, at an appropriate timing (for example, a preset time, or when a certain period of time has elapsed after the conversation has ended), the voice response system 100 will talk on behalf of the user (play the voice).
[0161] Specifically, the terminal device 1 performs the message terminal process shown in FIG. 8, and the server 90 performs the message server process shown in FIG. 9. The message terminal process is, for example, a process that starts when the power of the terminal device 1 is turned on and is then repeatedly executed.
[0162] In the message terminal process, as shown in FIG. 8, first, it is determined whether or not the message mode has been set by the user (S132). If the message mode has not been set (S132: NO), the process of S132 is repeated.
[0163] Also, if the message mode has been set (S132: YES), the processes of S2 to S8 are performed. If an affirmative determination is made in S6, the message mode flag is set to the ON state in the memory of the terminal device 1 (S134). Then, the processes of S10 to S16 are performed.
[0164] If an affirmative determination is made in S16, it is determined whether or not a packet has been received from the server 90 (S136). If a packet has not been received (S136: NO), the process of S136 is repeated. Also, if a packet has been received (S136: YES), the processes of S24 to S30 are performed, and the message terminal process ends.
[0165] Next, the message server process is a process that starts, for example, when the power of the server 90 is turned on and is then repeatedly executed. Specifically, first, it is determined whether or not a packet has been received from any of the terminal devices 1 (S142). If a packet has not been received (S142: NO), the process proceeds to S156, which will be described later.
[0166] Also, if a packet is being received (S142: YES), the terminal device 1 of the communication partner is identified (S44), and it is determined whether the packet contains a mode flag such as a message mode flag (S144). If there is no mode flag (S144: NO), the process proceeds to the process of S148.
[0167] Also, if there is a mode flag (S144: YES), the server 90 also sets the flag corresponding to the terminal device 1 of the communication partner to the ON state to perform mode setting (S146). For example, if the message mode flag corresponds to the message mode, the processes of S46 to S152 described later are executed, and if the guidance mode flag corresponds to the guidance mode described later, S46 to S176 (see FIG. 11) will be executed.
[0168] Subsequently, it is determined whether the message flag is in the ON state (S148). If the message flag is in the ON state (S148: YES), the processes of S46 to S54 are executed, and then, the message playback condition is extracted (S150).
[0169] Here, the message playback condition can be set in advance by the user via the operation unit 70 of the terminal device 1, and for example, the time and location are applicable. The message playback condition is transmitted to the server 90 when the packet of the message terminal process is transmitted.
[0170] Subsequently, the message is associated with the voice (voice color) and recorded in the memory (S152), and the process proceeds to the process of S156. Also, if the message flag is in the OFF state (S148: NO), the process related to another mode is performed (S154), and it is determined whether the playback timing has arrived (S156). Here, the playback timing indicates the content set by the message playback condition.
[0171] If it is not the playback timing (S156: NO), the message server process is immediately terminated. Also, if it is the playback timing (S156: YES), the processes of S62 to S64 are executed, and the message server process is terminated.
[0172] [Effect according to the third embodiment] According to the voice response system of the third embodiment, the voice input by the user is not played back immediately, but is played back when the message playback condition is met after a certain period of time.
[0173] For example, as shown in the third row of Figure 5, if you input "Please tell Mr. / Ms. XX that XX," the sentence you want to convey will be played after Mr. / Ms. XX's voice is recognized (heard).
[0174] [Modification of the third embodiment] In the third embodiment, the contents of what the user said are reproduced. The terminal device 1 may be configured to speak words that will trigger the conversation, such as, for example, "By the way, didn't you say that you were going to tell her something?" In detail, the terminal device 1 performs the guiding terminal process shown in Fig. 10, and the server 90 performs the guiding server process shown in Fig. 11.
[0175] The guiding terminal process is started, for example, when the terminal device 1 is powered on, and is then repeatedly executed. For example, the guiding terminal process is started, for example, when the terminal device 1 is powered on, and is then repeatedly executed.
[0176] In the guidance terminal process, as shown in Fig. 10, first, it is determined whether or not the guidance mode has been set by the user (S162). If the guidance mode has not been set (S162: NO), the process of S162 is repeated.
[0177] If the guidance mode is set (S162: YES), the processes of S2 to S8 are carried out, and if a positive determination is made in S6, a guidance mode flag is set to an ON state in the memory of the terminal device 1 (S164). Then, the processes of S10 to S16 are carried out.
[0178] If an affirmative determination is made in S16, it is determined whether a packet has been received from the server 90 (S166). If no packet has been received (S166: NO), the process of S166 is repeated. If a packet has been received (S166: YES), the processes of S24 to S30 are performed, and the guidance terminal process ends.
[0179] Next, the guidance server process is, for example, started when the power of the server 90 is turned on, and thereafter, it is a process that is repeatedly executed. Specifically, the processes of S142 to S146 described above are executed. Then, it is determined whether the guidance flag is in the ON state (S172).
[0180] If the guidance flag is in the ON state (S172: YES), the processes of S46 to S54 are performed, and subsequently, the guidance reproduction conditions are extracted (S174). Here, similar to the message reproduction conditions, the guidance reproduction conditions can be set in advance by the user via the operation unit 70 of the terminal device 1, and for example, the time and location are applicable. The guidance reproduction conditions are transmitted to the server 90 when the packet of the message terminal process is transmitted.
[0181] Subsequently, guidance content is generated, this guidance content is associated with voice (tone color), and recorded in the memory (S176). Here, as the guidance content, for example, words representing desires such as "want to" and "hope" included in the input character information are searched, the keywords before these words are extracted, and the words registered as the words for guiding these keywords are output as the guidance content. Note that the keywords and the words indicating the guidance content are associated in advance and recorded in the response candidate DB105.
[0182] Subsequently, the processes of S156 and below described above are performed, and the server process ends. If the guidance flag is in the OFF state (S172: NO), processing related to other modes is performed (S154), the processes of S156 and below described above are performed, and the server process ends.
[0183] According to the configuration of such a modification of the third embodiment, instead of directly outputting the words the user wants to say, it is possible to guide the user to be able to say the words they want to say. [Fourth Embodiment] [Processing of the Fourth Embodiment] Next, an example of using the terminal device 1 for reception work will be described. In this embodiment, the terminal device 1 is installed for company reception and the like. Note that it can also be adopted for telephone reception such as the company's representative telephone or telephone banking. Here, in this embodiment, the process of S56 in the first embodiment is realized by replacing it with the reception process shown in FIG. 12.
[0184] In the reception process, as shown in FIG. 12, first, it is determined whether the company name is included in the character information (S192). In this process, it is determined whether general names or company names (those recorded in the voice recognition DB102) are included.
[0185] If the company name or personal name is not included in the character information (S192: YES), a response for asking the company name and personal name is generated (S194), and the reception process ends. In this process, for example, a response such as "Please tell me your name and the matter." is generated.
[0186] If the company name or personal name is included in the character information (S192: NO), this company name and personal name are extracted from the sales DB118 and the client DB119 (S196). Here, in the sales DB118, the names of companies and persons in charge who have come for sales in the past, or the names of complainers who only talk about complaints, etc. are recorded. Also, in the client DB119, the company name, the person in charge of that company, the person in charge on the user side (own company side) of the terminal device 1, schedules such as the meeting scheduled time, and contact information are recorded in association with each person in charge.
[0187] Subsequently, it is determined whether the company name or personal name can be extracted from the sales DB118, that is, whether the company name or personal name included in the character information is included in the sales DB118 (S198). If the company name or personal name can be extracted from the sales DB118 (S198: YES), a sales rejection response (a response to reject the referral) indicating that the sales are rejected is generated (S200), and the reception process ends.
[0188] Also, if the company name or personal name cannot be extracted from the sales DB118 (S198: NO), it is determined whether the person who has come to the reception will visit at a nearby time (for example, within 1 hour before or after the current time) in the schedule in the client DB119 (S202). If the person will visit at a nearby time (S202: YES), the contact information of the person in charge of this person is extracted from the client DB119, and the person in charge is connected so that the person in charge and the person who has come to the reception can have a conversation (S204). In this process, it is sufficient to connect to the extension phone, mobile phone, etc. of the person in charge.
[0189] Subsequently, a reception response for the client is generated (S206). Here, as the reception response for the client, for example, a response such as "Thank you always, Mr. ○○. We have connected you to the person in charge, so please wait a moment." is generated. When such a process ends, the reception process ends.
[0190] Also, if the person is not the one who will visit at a nearby time (S202: NO), connect to the preset reception contact information, and connect to this reception person in charge so that the reception person in charge and the person who has come to the reception can have a conversation (S208). Then, a normal reception response is generated (S210).
[0191] Here, as the normal reception response, for example, a response such as "We have connected to the reception, so please wait a moment." is generated. When such a process ends, the reception process ends.
[0192] [Effects according to the fourth embodiment] In the above voice response system 100, it is configured to be used at a workplace or a company reception. In this configuration, the names and company names of the people coming to sales are pre-recorded in the sales DB 118 of the server 90, and when the person coming to the reception states these names or company names, a response is generated to play the voice of a rejection message.
[0193] Also, in the above voice response system 100, the server 90 identifies the communication partner based on the input character information and connects the preset communication destination and the communication partner for each communication partner. According to such a voice response system 100, it is possible to assist reception work and telephone response. Also, according to such a voice response system 100, it is possible to exclude those who may interfere with the user's work without the user having to handle them.
[0194] Furthermore, in the above voice response system 100, the server 90 extracts the keywords included in the input character information (especially voice), and connects to the connection destination corresponding to the keyword. Note that, for example, keywords such as the name of the other party and its connection destination are pre-associated.
[0195] According to such a voice response system 100, it is possible to assist operations such as telephone transfer and calling to the reception. [Modification Example of the Fourth Embodiment] In the above embodiment, the connection destination is configured to be set according to the other party, but by applying this technology, for example, in telephone reception such as telephone banking and telephone shopping, the requirements (keywords included in the character information) are recognized, and the connection destination may be changed according to the requirements.
[0196] Also, in the above voice response system 100, the server 90 may recognize the requirements for the other party to speak based on the keyword, and convey the outline of what the other party said to the user. According to such a voice response system 100, it is possible to assist the intermediary business with the customer.
[0197] [Fifth Embodiment] [Processing of the Fifth Embodiment] Next, the terminal device 1 may receive a request from another terminal device 1 and provide the information required by the other terminal device 1.
[0198] When configured in this way, in the process of S56, the server 90 requests the necessary information from another terminal device 1, obtains the necessary information from the other terminal device 1, and then generates a response. Then, in the terminal device 1 that provides the necessary information, the information providing terminal process shown in FIG. 13 is executed. The information providing terminal process is, for example, a process that starts when there is a request from the server 90.
[0199] As shown in FIG. 13, the information providing terminal process first extracts the information destination (S222). This information destination indicates another terminal device 1 that requests information, and the ID for specifying this other terminal device 1 is included in the request from the server 90.
[0200] Subsequently, it is determined whether the other party permits the provision of information (S224). Here, in the terminal information DB 113, the IDs of the parties who permit the provision of information, such as family members and friends, are recorded in advance. In this process, the determination is made by referring to this terminal information DB 113.
[0201] If the other party permits the provision of information (S224: YES), the information requested from its own memory 39, various sensors, etc. is obtained (S226), and this data is transmitted to the server 90 (S228). If the other party does not permit the provision of information (S224: NO), a message indicating the rejection of the information provision is transmitted to the server 90 (S230).
[0202] When such a process ends, the information providing terminal process ends. In this configuration, for example, as shown in the fourth row of FIG. 5, in response to the question "What is ○○ doing?", the server 90 requests the location information from ○○'s terminal device 1, and this terminal device 1 returns the location information.
[0203] Then, the server 90 recognizes the actions of Mr. ○○ based on the location information. For example, if the movement speed on the road is faster than the running speed of a human, it is determined that the person is moving on a train, and a response such as "Mr. ○○ is on the train. It seems that he is on his way home." will be generated.
[0204] [Effect according to the fifth embodiment] In the voice response system 100, the server 90 acquires information recorded in another terminal device 1 different from the requesting terminal device 1 from another terminal device 1 and provides it to another terminal device 1. That is, in the voice response system 100, the server 90 acquires information for generating a response to character information from another terminal device 1.
[0205] According to such a voice response system 100, a response can be generated based on the information recorded in another terminal device 1. Also, in the voice response system 100, when the terminal device 1 is requested for information for generating a response to character information from another terminal device 1, the terminal device 1 returns the information corresponding to this request.
[0206] In this configuration, the terminal device 1 is provided with sensors for detecting position information, temperature, humidity, illuminance, noise level, etc., and a database such as dictionary information, and extracts necessary information according to the request.
[0207] According to such a voice response system 100, it is possible to acquire information specific to another terminal device 1, such as the position of another terminal device 1. Also, it is possible to transmit its own specific information to another terminal device 1.
[0208] [Sixth embodiment] [Processing of the sixth embodiment] Next, in the voice response system of the sixth embodiment, a personality DB 106 in which personality information associating the personality of a user or a person related to the user according to a preset category is recorded is prepared. The personality DB 106 records, for example, as shown in FIG. 14, the names of the user and related persons and the personality categories of these persons in association with each other.
[0209] In addition, in the personality DB 106 shown in FIG. 14, a personality test is conducted on the user and related persons, and the test results are also recorded. Here, when generating personality information, well-known personality analysis techniques (such as the Rorschach test, the Thematic Apperception Test, etc.) may be used. In addition, when generating personality information, the technology of aptitude tests used by companies and the like in employment tests may also be used.
[0210] When generating personality information, for example, the personality information generation process shown in FIG. 15 is performed. The personality information generation process is, for example, a process that starts when an instruction to generate personality information is input using the operation unit 70 or the like in the terminal device 1.
[0211] In the personality information generation process, as shown in FIG. 15, first, the microphone 37 is turned on (S242), and one of the predetermined four-choice questions is output as voice (S244). At this time, for the four-choice questions, they may be obtained from the server 90, or the questions recorded in the memory 39 in advance may be presented.
[0212] Subsequently, it is determined whether there is a voice response from the subject (user or related person) (S246). If there is no response (S246: NO), the process of S246 is repeated. If there is a response (S246: YES), conversation parameters such as the ending of words and conversation speed are extracted (S248), and it is determined whether the current question is the last question (S250). If it is not the last question (S250: NO), the next question is selected (S252), and the process returns to S242.
[0213] If it is the last question (S250: YES), personality analysis is performed based on the answers to the four-choice questions (S254), and personality analysis is performed based on the conversation parameters (S256). Here, in the personality analysis based on conversation parameters, it is possible to capture the tendency that people with confidence in themselves have a strong ending of words, while those without confidence have a weak ending of words, and the tendency that hasty people have a fast conversation speed, while calm people have a slow conversation speed.
[0214] Subsequently, these personality analysis results are comprehensively analyzed, such as by weighted averaging (S258), and classified into personality categories (S260). Specifically, the personality of the subject obtained through the test is scored, and classified into personality categories for each score.
[0215] Subsequently, the subject is associated with the personality category (S262) and recorded in the personality DB 106 (S264). That is, the relationship between the subject and the personality category is transmitted to the server 90. At this time, the test results are also transmitted to the server 90, and the server 90 constructs a personality DB 106 as shown in FIG. 14. When such processing is completed, the personality information generation process is completed.
[0216] When using the personality DB 106 generated in this way, prepare in the response candidate DB 105 a correspondence between different responses and personality categories. Then, in the process of S56, the server 90 acquires response candidates representing a plurality of different responses to the character information, selects a response to be output from the response candidates according to the personality information, and outputs the selected response in the processes of S60 and S64.
[0217] [Effect according to the sixth embodiment] In the voice response system 100, the terminal device 1 generates personality information of the user or the person concerned based on the answers to a plurality of preset questions, and acquires the generated personality information.
[0218] According to such a voice response system 100, personality information can be generated in the server 90 or the terminal device 1. Furthermore, in the voice response system 100, the calculation unit 101 generates personality information of the user or the person concerned based on the character string included in the input character information.
[0219] According to such a voice response system 100, personality information can be generated in the process of the user using the voice response system 100. In addition, according to such a voice response system 100, different responses can be made according to the personalities of the user and those related to the user (related persons). Therefore, it is possible to improve the usability for the user.
[0220] [Modification Example of the Sixth Embodiment] In the above sixth embodiment, the response may be narrowed down to one according to the personality and then output, or different voice tones may be associated with and output for a plurality of responses.
[0221] Also, among the above personality information generation processes, the processes of S248 and S254 to S264 may be performed on the server 90. In this case, similar to the first embodiment and the like, while specifying the terminal device 1 to the server 90, voice and questions may be exchanged between the terminal device 1 and the server 90.
[0222] Furthermore, in the above voice response system 100, the server 90 may detect any of the actions and operations of the user, and generate learning information or personality information based on these.
[0223] According to such a voice response system 100, for example, when it is detected that the user jumps on the train continuously for several days, the user may be prompted to leave home a few minutes earlier from the next day, or when it is detected from the conversation that the user has a tendency to get angry easily, voice or music that calms the mood may be output.
[0224] [Seventh Embodiment] [Processing of the Seventh Embodiment] Next, in the voice response system of the seventh embodiment, a preference DB 108 in which preference information associated according to a preset category of the preferences of the user and related persons is recorded is prepared. The preference DB 108 records, for example, as shown in FIG. 16, the names of the user and related persons, and their preferences, associated with each of the types of preferences such as food preferences (food), color preferences (color), hobbies, etc.
[0225] In particular, regarding food preferences, they are classified into sweet lovers (sweet), spicy lovers (spicy), and those in between (neutral); regarding color preferences, they are classified into warm color systems (warm), cool color systems (cool), and those in between (neutral); regarding hobbies, they are classified into indoor hobbies (indoor), outdoor hobbies (outdoor), and hobbies that involve both indoor and outdoor activities (both indoor and outdoor).
[0226] When constructing such a preference DB 108, for example, the preference information generation process shown in FIG. 17 is executed. The preference information generation process is carried out, for example, between S48 and S54. Specifically, as shown in FIG. 17, keywords related to preferences are extracted from character information (S282), and among the objects identified by image processing, those related to preferences are extracted (S284). Note that the keywords related to preferences are associated in the preference DB 108 with the type of preference and the classification within that type (such as sweet, neutral, spicy for food preferences). In these processes, when the extracted keywords or objects are included in the preference DB 108, they are extracted as those related to preferences.
[0227] Subsequently, a counter is incremented for each group of keywords related to preferences (S288). For example, when an item like kimchi is extracted where the type of preference is "food preference" and the classification is "spicy", the counters corresponding to "food preference" and "spicy" are incremented.
[0228] Then, based on the counter values, the preference information (preference DB 108) is updated (S290). That is, for each "type of preference", the "classification" with the largest counter value is recorded in the preference DB 108 as the one that best matches the preference, as the characteristic of the user's or related person's preference. When such a process ends, the preference information generation process ends.
[0229] When using the preference DB108 generated in this way, in the response candidate DB105, prepare by associating different responses for each preference. In the process of S56, the server 90 acquires response candidates representing a plurality of different responses to the character information, selects a response to be output from the response candidates according to the preference information, and outputs the selected response in the processes of S60 and S64.
[0230] [Effect according to the seventh embodiment] In the voice response system 100, the server 90 generates preference information indicating the preference tendency of the user or the related person based on the character string included in the character information. Then, based on the preference information, a response to be output from the response candidates is selected, and the selected response is output.
[0231] According to such a voice response system 100, a response can be made according to the preference of the user or the related person. For example, when the user buys a present for a related person and asks the terminal device 1, "What does ○○-san want?", a response according to the preference information can be obtained.
[0232] [Modification example of the seventh embodiment] In the response candidate DB105, as shown in FIG. 18, a table associating the personality classification and the preference information may be provided.
[0233] For example, in the example shown in FIG. 18, the personality classification is associated with the preference for colors, and products that can be estimated to make a woman happy when received as a present are arranged in a matrix. In the process of S56, responses can also be generated taking into account both personality and preference in this way.
[0234] [Eighth embodiment] [Processing of the eighth embodiment] In the above embodiment, the voice is converted into character information, but the operation by the user may be converted into character information.
[0235] Specifically, the terminal device 1 captures the user's actions as captured images and transmits them to the server 90. The server 90 may perform, for example, the action text input process shown in FIG. 19. The action text input process is a process that starts when the user's body part appears in the captured image in the process of S48.
[0236] In the action text input process, as shown in FIG. 19, first, a captured image is acquired (S302). Then, it is determined whether the user is trying to input text by handwriting or by sign language (S304, S308).
[0237] In these processes, for example, when the upper body of the user appears in the captured image together with the face, it is determined that the user is trying to input text by sign language. When only the user's hand appears in the captured image without the user's face, it is determined that the user is trying to input text by handwriting.
[0238] If the user is trying to input text by handwriting (S304: YES), the behavior of the fingertip or pen tip is recorded (S306), and based on this behavior, the behavior is converted into character information (S312). Here, in the handwritten character / sign language DB 112, the behavior when writing characters is associated with the characters, and also the hand movement is associated with the characters expressed by sign language. In the process of S312, character information is generated by referring to the handwritten character / sign language DB 112.
[0239] If the user is trying to input text by sign language (S304: NO, S308: YES), the sign language content is recognized by referring to the handwritten character / sign language DB 112, and the above-described process of S312 is performed. If the user is not trying to input text by handwriting or sign language (S308: NO), the input process by other methods is performed (S314).
[0240] Subsequently, the characters input by operation are associated with the characters input by voice, and it is determined whether there is voice with similarity (whether the degree of coincidence between the reference waveform based on the characters and the pronunciation waveform is equal to or higher than the reference value) (S316). If there is such voice input (S316: YES), the accent and pronunciation characteristics when this user inputs this character are recorded in the learning DB107 in association with the character (S318), and the operation character input process is terminated.
[0241] If there is no such voice input (S316: NO), the operation character input process is terminated. to do. [Effect according to the eighth embodiment] In the voice response system 100, since the operation by the user is converted into character information, the user can input character information without making a sound.
[0242] [Modification example of the eighth embodiment] As the operation of this embodiment, it is sufficient if it is caused by not only handwriting of characters or body language (for example, sign language) but also muscle movement.
[0243] [Ninth embodiment] [Processing of the ninth embodiment] The content of the learning DB107 may be made available in another terminal device 1 different from the terminal device 1 normally used by the user when using this other terminal device 1. In this case, from the other terminal device 1, the ID and password of the terminal device 1 normally used are transmitted to the server 90 together with the usage request.
[0244] Then, the server 90 executes the other terminal usage process shown in FIG. 20. The other terminal usage process is a process that starts when a usage request is received. In the other terminal usage process, as shown in FIG. 20, first, it is determined whether an ID and a password have been input (S332). If the ID and password have not been input (S332: NO), the process of S332 is repeated.
[0245] Also, if an ID and a password are entered (S332: YES), it is determined whether authentication using the ID and the password has been completed (S334). If the authentication has been completed (S334: YES), a message indicating that the authentication has been completed is transmitted to another terminal device 1 (S336), and the other terminal device 1 is set to use the learning DB107 of the terminal device 1 corresponding to the ID and the password (S338).
[0246] If the authentication has not been completed (S334: NO), a message indicating an error is transmitted to another terminal device 1 (S340), and the other terminal usage process is terminated. [Effect according to the ninth embodiment] In addition, in the voice response system 100, the server 90 transfers the learning information of a certain terminal device 1 to another terminal device 1.
[0247] According to such a voice response system 100, even when a user using a certain terminal device 1 uses another terminal device 1, the learning information (learning information recorded in the server 90) recorded in the certain terminal device 1 can be used. Therefore, the generation accuracy of character information can be improved even when using another terminal device 1. In particular, it is effective when the user has a plurality of terminal devices 1.
[0248] Furthermore, in the voice response system 100, the server 90 outputs information about the user in response to an inquiry from a person other than the user. According to such a voice response system 100, for example, if the user's meal content or walking distance is detected, it is possible to answer questions at a hospital or the like on behalf of the user. Also, it may be learned about the health condition or self-introduction.
[0249] [Modification example of the ninth embodiment] Similar to the configuration of the ninth embodiment, when a request to end the use and an ID and a password are received, the use of the learning DB107 for the terminal device 1 corresponding to the ID and the password may be ended (prohibited).
[0250] [Tenth embodiment] [Processing of the 10th Embodiment] In the voice response system of the 10th embodiment, the server 90 stores the conversation content and asks questions to obtain the same content about what has been heard. Specifically, in S100 of the automatic conversation server processing shown in FIG. 7, the memory confirmation process shown in FIG. 21 is executed.
[0251] In the memory confirmation process, as shown in FIG. 21, the past conversation content is extracted from the learning DB107 (S352), and a question is generated using a keyword included in any of the conversation contents as an answer (S353). When such processing is completed, the memory confirmation process ends.
[0252] In the memory confirmation process, for example, questions such as "What was the menu for dinner yesterday?" or "Where did you go three days ago?" may be asked. [Effect of the 10th Embodiment] According to such a voice response system 100, it is possible to confirm the user's memory and promote the fixation of memory. It is considered to be effective also in suppressing the progression of dementia in the elderly.
[0253] [11th Embodiment] [Processing of the 11th Embodiment] Next, in the voice response system of the 11th embodiment, the terminal device 1 and the server 90 are used to configure the system so that the user can practice a foreign language.
[0254] Specifically, the pronunciation determination process 1 shown in FIG. 22, the pronunciation determination process 2 shown in FIG. 23, and the pronunciation determination process 3 shown in FIG. 24 are executed in order. However, the server 90 executes one of the processes of the pronunciation determination processes 1 to 3 every time the voice response server process (FIG. 2) is executed. Also, each of the pronunciation determination processes 1 to 3 is executed as the process of S56 described above.
[0255] First, in pronunciation determination process 1, as shown in FIG. 22, a response instructing to input a predetermined sentence by voice is generated (S362). In this process, for example, a sentence serving as a model for a foreign language is generated, and the user is prompted to imitate and speak following this sentence. When this process ends, pronunciation determination process 1 ends.
[0256] Next, when voice is input in accordance with pronunciation determination process 1, pronunciation determination process 2 is performed. In pronunciation determination process 2, as shown in FIG. 23, the accuracy of pronunciation and accent is scored (S372). In this process, the voice is captured as a waveform, and the degree of coincidence of the waveform with the waveform when the model sentence is used as a waveform is scored.
[0257] Then, this score is recorded in the memory (S374), and pronunciation determination process 2 ends. Subsequently, pronunciation determination process 3 is performed. In pronunciation determination process 3, as shown in FIG. 24, first, it is determined whether the score is less than the threshold value (S382).
[0258] If the score is less than the threshold value (S382: YES), a response instructing to input the same sentence again is generated (S384). In this process, for example, a response for prompting to imitate and speak following the model again is generated.
[0259] If the score is greater than or equal to the threshold value (S382: NO), a response indicating that the pronunciation was good and prompting to input the next sentence is generated (S386). For example, a response such as "Good pronunciation. Let's move on." is generated.
[0260] When such a process ends, pronunciation determination process 3 ends. [Effect according to the 11th Embodiment] In the voice response system 100, the server 90 detects the accuracy of the pronunciation and accent of the voice input by the user, and outputs the detected accuracy.
[0261] According to such a voice response system 100, it is possible to check the pronunciation and accent accuracy. For example, it is effective when practicing a foreign language. Furthermore, in the voice response system 100, when the accuracy level is below a certain value, the server 90 causes the same question to be output again.
[0262] According to such a voice response system 100, an accurate answer can be obtained by outputting the same question. [Modification Example of the 11th Embodiment] In the voice response system 100, when the accuracy level is below a certain value, for confirmation, the server 90 may output a voice including words closest to the pronunciation made by the user.
[0263] According to such a voice response system 100, the user can check the pronunciation and accent accuracy. [12th Embodiment] [Processing of the 12th Embodiment] Next, the voice response system of the 12th embodiment will be described. In the voice response system of the 12th embodiment, the emotion of the user is detected from the voice input by the user, and a response for healing the user is generated according to the emotion.
[0264] Specifically, the emotion determination process shown in FIG. 25 and the emotion response generation process shown in FIG. 26 are executed. The emotion determination process is implemented as the details of the process of S50 described above. As shown in FIG. 25, first, the emotion is scored from the tone color, the strength of the end of the sentence, the length of a sentence, the conversation speed, the words that come out unexpectedly, etc. (S392). Subsequently, the emotion is classified according to the score and recorded in the memory (S394).
[0265] When such a process ends, the emotion determination process ends. Subsequently, in the process of S56 described above, the emotion response generation process is executed. Specifically, as shown in FIG. 26, first, the emotion category set in the emotion determination process is determined (S412). If the emotion category is normal (S412: normal), an ordinary greeting sentence such as "Hello" is generated as a response (message) (S414).
[0266] If the emotion category is anger (S412: anger), a sentence for calming the other person's emotion, such as "Did I do something to offend you?", is generated as a response (S416). Further, if the emotion category is joy (S412: joy), a greeting sentence with a brighter nuance compared to an ordinary greeting sentence, such as "It's a fun day today", is generated as a response (S418).
[0267] If the emotion category is confusion (S412: confusion), a greeting sentence for showing concern for the other person, such as "Is something wrong?", is generated as a response (S420). When such processing is completed, the emotion response generation process is completed.
[0268] [Effect according to the 12th Embodiment] In the voice response system 100, the server 90 detects the user's irritation and agitation by detecting the voice uttered unexpectedly, and generates a message for suppressing the irritation and agitation.
[0269] According to such a voice response system 100, when the user has irritation or agitation, these can be suppressed. Therefore, the occurrence of troubles between the user and the surroundings can be suppressed.
[0270] [13th Embodiment] [Processing of the 13th Embodiment] Next, the voice response system of the 13th embodiment will be described. In the voice response system of the 13th embodiment, processing for guiding the user to an object in the captured image is performed. This processing is implemented as the details of the above-described S56 processing in the server 90.
[0271] When an input such as "Please guide me to the visible tower" is made by voice in the terminal device 1, guidance processing is performed in the process of S56. In the guidance processing, as shown in FIG. 27, first, terminal position information is acquired from the GPS receiver 27 or the like of the terminal device 1 (S432).
[0272] Then, based on the voice (character information) and image processing, a target object is specified from among the objects in the captured image, and this position is specified (S434). In this process, the position of the object is specified in the map information (which may be acquired from the outside or may be held by the server 90) based on the shape, relative position, etc. of the object. For example, when a tower is shown in the captured image, the tower is specified on the map from the position of the terminal device 1 and the shape of the tower.
[0273] Subsequently, a route to this object is searched (S436), and route information is acquired (S438). This process can be realized using the same process as that in a well-known cloud-based navigation device.
[0274] Then, a response for guiding the route is generated (S440). Also in this process, a response similar to the guidance by the navigation device may be generated. When such processing is completed, the guidance processing is terminated. When the guidance processing is performed while the user is moving, the automatic conversation server processing may be used to play a message on the condition that the user reaches the point to be guided.
[0275] [Effect according to the 13th Embodiment] In the voice response system 100, when character information is input, the server 90 generates a response corresponding to the captured image of the periphery of the voice response system 100 and outputs this response by voice.
[0276] According to such a voice response system 100, a response can be output by voice according to the captured image. Therefore, the usability can be improved as compared with a configuration that generates a response only from character information.
[0277] Also, in the voice response system 100, the server 90 searches for an object included in the character information from the captured image by image processing, specifies the position of the searched object, and guides the user to the position of this object.
[0278] According to such a voice response system 100, the user can be guided to an object in the captured image. Furthermore, in the voice response system 100, when the server 90 guides to a destination, it acquires route information such as weather, temperature, humidity, traffic information, road surface condition, etc. to the destination, and outputs the route information by voice.
[0279] According to such a voice response system 100, the situation (route information) to the destination can be notified to the user by voice. [Modification of the 13th Embodiment] In addition to the above configuration, character information may be input so as to respond what the recognized object is, and what (who) the recognized object is from the captured image may be output by voice.
[0280] Furthermore, in the voice response system 100, the server 90 may acquire a moving image capturing the shape of the user's mouth when inputting character information by voice instead of the process of S48. In this case, instead of the process of S52, the voice may be converted into character information, and based on the moving image, the unclear part of the voice may be estimated and the character information may be corrected.
[0281] According to such a voice response system 100, the pronunciation content can be estimated from the shape of the mouth, so the unclear part of the voice can be estimated well. [14th Embodiment] [Processing of the 14th Embodiment] Next, the voice response system of the 14th embodiment will be described. In the voice response system of the 14th embodiment, the user is required to perform a predetermined operation, and it is determined whether the user has performed the operation as required. In this configuration, in the automatic conversation terminal process shown in FIG. 6 and the automatic conversation server process shown in FIG. 7, the movement request process 1 shown in FIG. 28 and the movement request process 2 shown in FIG. 29 are sequentially executed as the details of the process of S56 described above.
[0282] First, when the process of S54 ends, the movement request process 1 is started. In the movement request process 1, as shown in FIG. 28, a response (message) instructing to move the line of sight or the head to a predetermined position is output (S452). When this process ends, the movement request process 1 ends.
[0283] Subsequently, when the process of S54 ends next, the movement request process 2 is started. In the movement request process 2, as shown in FIG. 29, it is determined whether the position of the line of sight or the head has moved as instructed (S462). In this process, the captured image by the camera is subjected to image processing, and the operation of the user is detected using the detection results of various sensors of the terminal device 1. When detecting the line of sight by image processing, a well-known line-of-sight recognition technique may be adopted.
[0284] If the line of sight or the head has not moved as instructed (S462: NO), the response generated in S452 is output again (S464). If the line of sight or the head has moved as instructed (S462: YES), another arbitrary response is generated (S466).
[0285] When such a process ends, the movement request process 2 ends. [Effect according to the 14th embodiment] In the voice response system 100 described above, the line of sight of the user is detected, and when the line of sight of the user does not move to a predetermined position in response to a call, a voice requesting to move the line of sight to the predetermined position is output.
[0286] According to such a voice response system 100, the user can be made to look at a specific position. Therefore, safety checks during vehicle driving and the like can be reliably performed. In addition, in the voice response system 100, the server 90 observes the position of the body part and the facial expression, and when there are few changes in response to the call, outputs a voice requesting to change the position of the body part and the facial expression.
[0287] According to such a voice response system 100, the position of the user's body part can be moved to a specific position, or the user can be induced to make a specific expression. The present invention can be used during vehicle driving, physical examinations, and the like.
[0288] [15th Embodiment] [Processing of the 15th Embodiment] Next, the voice response system of the 15th embodiment will be described. In the voice response system of the 15th embodiment, when the user inputs a broadcast program or a piece of music as voice, a process of complementing when the broadcast program or the piece of music is interrupted is performed.
[0289] In this configuration, as the details of S56 described above, the broadcast music complementing process shown in FIG. 30 is performed. In the broadcast music complementing process, as shown in FIG. 30, first, it is determined whether a broadcast program or a piece of music (the song when the user sings) has been interrupted (S482).
[0290] If it has been interrupted (S482: YES), a synchronized broadcast program or piece of music is set as the response content in the process of S492 described later (S484), and the broadcast music complementing process is terminated. If it has not been interrupted (S482: NO), if a broadcast program is being viewed, the broadcast program is acquired (S486), and if a piece of music is being played, the corresponding piece of music is acquired (S488).
[0291] Here, in the karaoke DB 116, a piece of music and lyrics are recorded in association with each other, and when acquiring a piece of music in this process, a piece of music with lyrics is acquired. Subsequently, a broadcast program or a piece of music being viewed by the user is identified (S490). Then, this broadcast program or piece of music is acquired and prepared for playback in synchronization with the broadcast program or piece of music being viewed by the user (S492), and the broadcast music completion process ends.
[0292] [Effect according to the 15th Embodiment] In the voice response system 100, the server 90 acquires a broadcast program similar to the broadcast program being viewed by the user, and when the broadcast program is interrupted, the interrupted broadcast program is complemented by outputting the broadcast program acquired by itself.
[0293] According to such a voice response system 100, it is possible to complement so that the broadcast program being viewed by the user is not interrupted. Also, in the voice response system 100, when the user sings a song by attaching lyrics to a song without lyrics, the server 90 compares the song with lyrics and the lyrics attached by the user, and outputs the lyrics in voice at the part where only the user's lyrics are missing.
[0294] According to such a voice response system 100, it is possible to complement the part that a user using a so-called karaoke device cannot sing (the part where the lyrics are interrupted). [16th Embodiment] [Processing of the 16th Embodiment] Next, the voice response system of the 16th embodiment will be described. In the voice response system of the 16th embodiment, when characters are included in the captured image and the terminal device 1 receives a question from the user about the reading of these characters, the information of these characters is acquired from the outside, and the reading of the characters included in this information is output in voice.
[0295] In this configuration, as details of S56 described above, the character explanation process shown in Fig. 31 is performed. In the character explanation process, as shown in Fig. 31, first, it is determined whether or not a question about the reading, such as "how to read", has been received (S502). If a question about the reading has been received (S502: YES), the reading of the image-recognized character is searched for from other servers connected via the Internet network 85 (S504), the obtained reading is set in the response (S506), and the character explanation process ends.
[0296] If it is not a question about reading (S502: NO), then it is a question about "words" such as those in a Japanese dictionary. It is determined whether or not a question about the meaning of "is received (S508)." If a question about the meaning is received, the meaning of the image-recognized characters (words) is searched for from other servers etc. connected via the Internet network 85 (S510), the obtained meaning is set as a response (S512), and the character explanation process is terminated.
[0297] [Effects of the 16th embodiment] According to such a voice response system 100, the reading of characters recognized by image recognition is searched for from other servers, etc., and the obtained reading is set in the response, so that the user can be informed of how to read characters, the meaning of words, etc.
[0298] [Seventeenth embodiment] [Processing of the 17th embodiment] Next, a voice response system according to a seventeenth embodiment will be described. In the voice response system according to the seventeenth embodiment, a server 90 detects abnormal behavior or a state of a user of the terminal device 1 based on a sensor value detected by the terminal device 1, and performs a process of reporting an abnormality when the abnormality is detected.
[0299] Specifically, in the terminal device 1, the action response terminal process shown in FIG. 32 is performed, and in the server 90, the action response server process is performed. In the action response terminal process, as shown in FIG. 32, first, the outputs from various sensors mounted on the terminal device 1 are acquired (S522), and the captured image by the camera 41 is acquired (S524). Then, the outputs from the various sensors and the captured image are packet-transmitted to the server 90 (S526), and the action response terminal process is terminated.
[0300] Next, in the action response server process, as shown in FIG. 33, first, the processes of S42 to S44 described above are performed. Subsequently, based on the position information of the terminal device 1 (the detection result by the GPS receiver 27), actions such as wandering are specified (S532), and the environment of the user is detected based on the detection results by the temperature sensors 15, 19, etc. (S534). Then, an abnormality is detected (S536).
[0301] In this process, an abnormality is detected based on the change in position information and the environment. For example, when the user does not move in a high-temperature or low-temperature place or when the user is present in a place where they usually do not go, it is detected that there is an abnormality (S536). Alternatively, the position information and the environment are scored, and when this score is below the reference value (when it is outside the reference range), it is determined that there is an abnormality.
[0302] Subsequently, it is determined whether an abnormality has been detected (S538). If no abnormality has been detected (S538: NO), the action response server process is terminated. If an abnormality has been detected (S538: YES), a message indicating that there is an abnormality is generated (S540) and reported to a predetermined contact (S542). Then, the processes of S62 to S68 (excluding S66) described above are performed, and the action response server process is terminated.
[0303] [Effect according to the 17th Embodiment] In the voice response system 100 described above, the server 90 detects the actions of the user and the surrounding environment of the user, and generates a message according to the detected actions and the surrounding environment.
[0304] According to such a voice response system 100, it is possible to notify dangerous places, restricted areas, etc. Further, it is possible to detect that the user has abnormal behavior or the like.
[0305] Furthermore, in the voice response system 100, the server 90 determines the health state based on the captured image of the user, and generates a message according to this health state. According to such a voice response system 100, the health state of the user can be managed.
[0306] Also, in the voice response system 100, when the health state is below the reference value, the server 90 makes a notification to a predetermined contact. According to such a voice response system 100, when the health state of the user is below the reference value, a notification can be made. Therefore, an abnormality can be notified to others earlier.
[0307] [Other Embodiments] The embodiments of the present invention are not limited to the above embodiments at all, and various forms can be adopted as long as they belong to the technical scope of the present invention.
[0308] For example, the voice response system 100 may mediate communication between two parties and among multiple parties. Specifically, when it is necessary to yield the road at an intersection or the like, the terminal devices 1 may negotiate which vehicle enters the intersection first. In this case, each terminal device 1 transmits information on the moving direction and the approaching speed to the intersection to the server 90, and the server 90 sets a priority for each terminal device 1 according to the moving direction and the approaching speed, and generates and outputs voices such as "stop" or "entry permitted" according to the priority.
[0309] Also, when the terminal device 1 receives an incoming call (incoming call) for communication that requires a real-time response, such as voice communication, it may be configured to receive the incoming call only when it is convenient for the user. Specifically, when the camera 41 can capture the user's face, it may be regarded as a convenient time for the user, and the incoming call may be received accordingly.
[0310] Furthermore, during voice communication or the like, there are people who become unhappy if they call someone and the other person does not respond. To suppress such feelings, it may be possible to inform the user waiting for a response from the other party about the situation of the other party. For example, the terminal device 1 manages the user's schedule. If the user does not respond to an incoming call, it is conceivable to search for what the user is doing or the free time in the user's schedule and inform the other party when the user can respond.
[0311] Also, if the user does not respond to an incoming call, the user's location may be notified to the calling party. For example, if the user is connected to the Internet or the like via a smartphone or a personal computer, it is possible to know which terminal is being operated. From this information, it is conceivable to identify the user's location and notify the calling party.
[0312] Furthermore, it may be possible to determine whether the user can respond to an incoming call by using location information such as GPS. Based on the location information, it is possible to determine whether the user is in a car, at home, etc. For example, if the user is in motion or on the bed, it may be determined that the user has a high level of public exposure or is sleeping and cannot respond to the incoming call. In such a case where the user cannot respond to the incoming call, it is conceivable to inform the calling party about what the user is doing as described above.
[0313] In addition, in order to obtain location information, a configuration using a security camera can also be considered. In recent years, security cameras have been installed in various places. Therefore, by using these security cameras and a configuration for identifying an individual such as face authentication, the location of the user can be recognized. Further, it is also possible to make a situation determination such as what the user is doing (whether the user is on the phone) using the security camera. Also, regarding whether or not an incoming call can be answered, it can be determined based on conditions such as whether another landline is being used (an incoming call cannot be answered while a landline is in use). It is possible.
[0314] Furthermore, when the user of the terminal device 1 wants to talk to someone, the user's personality learning results may be used to call a terminal device that is estimated to have good compatibility among a large number of unspecified users. Also, in such a case, it may be possible to talk to the user about a topic that is likely to liven things up (a topic that both users are interested in (extracted using the learning results)).
[0315] Also, when the use of the voice response device has not occurred for a long time (when the user has not spoken for more than a reference time), the voice response device may say something to the user. At this time, the words to be spoken may be selected using location information such as GPS.
[0316] [Relationship between each means described in the claims or means for solving the problem (the present invention) and the configuration in the embodiment] The terminal device 1 and the server 90 in the above embodiment correspond to an example of the voice response device of the present invention. Also, the processes of 22 and S56 in the above embodiment correspond to an example of the response acquisition means of the present invention.
[0317] Furthermore, the processes of S28, S60, and S64 in the above embodiment correspond to an example of the voice output means of the present invention. Also, the processes of S2 and S6 in the above embodiment correspond to an example of the voice input means of the present invention.
[0318] Furthermore, the process of S14 in the above embodiment corresponds to an example of the voice transmission means of the present invention. Also, the response candidate DB105 in the above embodiment corresponds to an example of the response recording means of the present invention.
[0319] Furthermore, the process of S56 in the above embodiment corresponds to an example of the personality information acquisition means of the present invention. Also, the processes of S22 and S56 in the above embodiment correspond to an example of the response acquisition means of the present invention.
[0320] Furthermore, the processes of S28, S60, and S64 in the above embodiment correspond to an example of the voice output means of the present invention. Also, the processes of S254, S258, and S260 in the above embodiment correspond to an example of the first personality information generation means and the second personality information generation means of the present invention. Also, the process of S56 in the above embodiment corresponds to an example of the personality information acquisition means of the present invention.
[0321] Furthermore, the processes of S22 and S56 in the above embodiment correspond to the response acquisition means of the present invention. Also, the processes of S28, S60, and S64 in the above embodiment correspond to an example of the voice output means of the present invention.
[0322] Furthermore, the processes of S254, S258, and S260 in the above embodiment correspond to an example of the first personality information generation means and the second personality information generation means of the present invention. Furthermore, the processes of S48 and S56 in the above embodiment correspond to an example of the response generation means of the present invention. Also, the processes of S28, S60, and S64 in the above embodiment correspond to an example of the voice output means of the present invention.
[0323] Furthermore, the modified example in the above embodiment: the process of S48 corresponds to an example of the voice input video acquisition means of the present invention. Also, the process of S52 in the above embodiment corresponds to an example of the character information conversion means of the present invention.
[0324] Furthermore, the preference information generation process in the above embodiment is an example of the preference information generation means of the present invention corresponds. Also, the process of S56 in the above embodiment corresponds to an example of the response candidate acquisition means of the present invention.
[0325] Furthermore, the action character input process in the above embodiment corresponds to an example of the character information generation means of the present invention. Also, the other terminal utilization process in the above embodiment in the other device information acquisition means corresponds to an example of the transfer means of the present invention.
[0326] Furthermore, the process of S98 in the above embodiment corresponds to an example of the reproduction condition determination means of the present invention. Also, the process of S100 in the above embodiment corresponds to an example of the message reproduction means of the present invention.
[0327] Furthermore, the process of S116 in the above embodiment corresponds to an example of the non-response time transmission means of the present invention. Also, the process of S372 in the above embodiment corresponds to an example of the speech accuracy detection means of the present invention.
[0328] Furthermore, the process of S374 in the above embodiment corresponds to an example of the accuracy output means of the present invention. Also, the process of S204 in the above embodiment corresponds to an example of the connection control means of the present invention.
[0329] Furthermore, the process of S50 in the above embodiment corresponds to an example of the emotion determination means of the present invention. Also, the process of S438 in the above embodiment corresponds to an example of the route information acquisition means of the present invention.
[0330] Furthermore, the process of S462 in the above embodiment corresponds to an example of the line-of-sight detection means of the present invention. Also, the process of S464 in the above embodiment corresponds to an example of the line-of-sight movement request transmission means of the present invention.
[0331] Furthermore, the process of S464 in the above embodiment corresponds to an example of the change request transmission means of the present invention. Also, the process of S486 in the above embodiment corresponds to an example of the broadcast program acquisition means of the present invention.
[0332] Furthermore, the process of S484 in the above embodiment corresponds to an example of the broadcast program complementing means and the lyrics adding means of the present invention. Also, the processes of S504 and S506 in the above embodiment correspond to an example of the reading output means of the present invention. Also, the processes of S522 and S524 in the above embodiment correspond to an example of the action environment detecting means of the present invention.
[0333] Furthermore, the process of S538 in the above embodiment corresponds to an example of the health state determination means of the present invention. Also, the process of S540 in the above embodiment corresponds to an example of the health message generation means of the present invention.
[0334] Furthermore, the process of S542 in the above embodiment corresponds to an example of the notification means of the present invention.
Explanation of Signs
[0335] 1... Terminal device, 10... Action sensor unit, 11... Three-dimensional acceleration sensor, 13... Axis gyro sensor, 15... Temperature sensor, 17... Humidity sensor, 19... Temperature sensor, 21... Humidity sensor, 23... Illuminance sensor, 25... Wetness sensor, 27... GPS receiver, 29... Wind speed sensor, 33... Electrocardiogram sensor, 35... Heart sound sensor, 37... Microphone, 39... Memory, 41... Camera, 50... Communication unit, 53... Radio phone unit, 55... Contact memory, 60... Notification unit, 61... Display, 63... Electric decoration, 65... Speaker, 70... Operation unit, 71... Touch pad, 73... Confirmation button, 75... Fingerprint sensor, 77... Rescue request lever, 80... Communication base station, 85... Internet network, 90... Server, 100... Voice response system, 101... Arithmetic unit, 102... Voice recognition DB, 103... Prediction conversion DB, 104... Voice DB, 105... Response candidate DB, 106... Personality DB, 107... Learning DB, 108... Hobby DB, 109... News DB, 110... Weather DB, 111... Reproduction condition DB, 112... Handwritten character / sign language DB, 113... Terminal information DB, 114... Emotion determination DB, 115... Health determination DB, 116... Karaoke DB, 117... Notification destination DB, 118... Sales DB, 119... Client DB.
Claims
【Claim 1】 An audio response system comprising a claimant device that is the source of a claim, a provider device that is an information provider, and a server capable of communicating with the claimant device and the provider device, wherein the provider device is configured to be able to set whether to permit information provision, the server obtains a request from the claimant device, and when the request includes a request designating the provider device, transmits the request to the designated provider device, and is configured to obtain provided information provided from the provider device when the information provision is permitted, an information acquisition unit; a provision unit configured to generate an audio response based on the provided information as a response to the request from the claimant device and provide the response to the claimant device; An audio response system comprising the above.
Citation Information
Patent Citations
Notifying device for position of automobile telephone set
JP1991120995A
Method for transferring information with equipment, equipment with interactive function applying the same and life support system formed by combining equipment
JP2001256036A
System, method and program for providing information
JP2002342356A
Response message generation apparatus, and terminal device thereof
JP2003108376A
Life adviser support system, adviser side terminal system, authentication server, server, support method and program
JP2009151766A
Cited By
Game machine
JP2025129264A
Game machine
JP2025129265A
Game machine
JP2025129266A