Information processing device

The system allows voice recognition devices to collaborate and coordinate responses across multiple units by using a private network and voice authentication, addressing the limitations of individual device operation and ensuring secure, coordinated communication.

JP7719926B2Active Publication Date: 2025-08-06MAXELL LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024139640
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2025-08-06
Estimated Expiration
2038-04-09

AI Technical Summary

Technical Problem

Existing voice recognition devices lack the capability to collaborate with each other, limiting their use to individual units and preventing coordinated responses to voice commands across multiple devices in different rooms.

Method used

A system that enables multiple voice recognition devices to communicate and collaborate through a private network, allowing them to switch between normal and conversation modes based on hot words and voice authentication, facilitating coordinated responses and communication within designated areas.

Benefits of technology

Enables seamless communication and coordinated responses across multiple voice recognition devices, enhancing usability and security by allowing devices to work together and ensuring only authorized individuals can participate in conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007719926000001
    Figure 0007719926000001
  • Figure 0007719926000002
    Figure 0007719926000002
  • Figure 0007719926000003
    Figure 0007719926000003
Patent Text Reader

Abstract

To utilize multiple speech recognition devices linked with each other.SOLUTION: A speech recognition device consists of: a speech input unit; a speech output unit; a communication control unit that transmits and receives data to and from other speech recognition devices; a conversation mode execution unit that transmits speech data input from the speech input unit to the other speech recognition devices, and outputs the speech data received from the other speech recognition devices from the speech output unit; a speech recognition unit that converts speech input from the speech input unit into text data; a hot word detection unit that detects conversation activation hot words that instruct the conversation mode execution unit to activate from the text data; and a command transmission unit that transmits control commands to the other voice recognition devices. When the hot word detection unit detects the conversation activation hot words, the command transmission unit transmits the control commands to the other speech recognition devices to activate the conversation mode execution unit provided in the other speech recognition devices.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention provides Information processing device Regarding. [Background technology]

[0002] In recent years, speech recognition has been applied to speech content, and by analyzing the content of speech, A so-called smart speaker that outputs a response based on the user's voice from an attached speaker, or A voice recognition device called an AI speaker is being produced that uses voice recognition technology and artificial intelligence technology. For example, Patent Document 1 describes an example of speech recognition technology as The interpretation candidates obtained by applying the recognition algorithm are integrated by reducing redundancy. Present the user with an opportunity to choose from alternative interpretations to provide the method presented and is structured to present these options without overlapping elements (summary excerpt).” In addition, Patent Document 2 states, "Speech recognition model and data in automatic speech recognition system" As a feature to manage data usage, when users may use the system It states that "these data can be tracked to predict future outcomes (summary excerpt)." [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-68952 [Patent Document 2] Special Publication No. 2015-537258 Summary of the Invention [Problem to be solved by the invention]

[0004] The voice recognition devices such as the smart speakers mentioned above are A person in the vicinity of the device can give voice instructions, which are then processed by the device and a response is obtained. However, it does not disclose how multiple devices can be used in conjunction with each other. In other words, an example of use at home is when family members share the device in the living room to check the weather forecast, When you request information on the Internet, such as news and music, the corresponding content will be displayed. By simply outputting the voice command, you can have the device recognized by another voice recognition device in a room other than the living room, for example, a child's room. It is not assumed that any communication or collaborative actions will be carried out between the device and the chair. Therefore, multiple voice recognition devices cannot be used in conjunction with each other, and There is room for innovation in new ways of using it.

[0005] The present invention has been made in consideration of the above-mentioned circumstances, and is a system that allows a plurality of voice recognition devices to be used in cooperation with each other. Information processing device The purpose is to provide. [Means for solving the problem]

[0006] In order to achieve the above object, the present invention has the configurations set forth in the claims. [Effects of the Invention]

[0007] According to the present invention, it is possible to use a plurality of voice recognition devices in cooperation with each other. Information processing device Objects, configurations, and effects other than those described above will be apparent from the following embodiments. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating a hardware configuration of a voice recognition device according to an embodiment of the present invention. [Figure 2] Functional block diagram of the voice recognition device [Figure 3A] FIG. 10 is a diagram showing an example of hot word data stored in a reference memory. [Figure 3B] FIG. 10 is a diagram showing an example of voice authentication data stored in a reference memory. [Figure 4] A diagram showing a collaboration system using a voice recognition device [Figure 5] 10 is a flowchart showing a first call operation in a linkage system. [Figure 6] FIG. 10 is a diagram showing an example of location estimation data. [Figure 7] A flowchart showing the flow of a first switching process between normal mode and conversation mode. [Figure 8] Flowchart showing the flow of a second switching process between normal mode and conversation mode [Figure 9] Conceptual diagram showing the first setting process at the time of new installation [Figure 10] Conceptual diagram showing the second setting process during new installation [Figure 11] Flowchart showing the flow of the second setting process at the time of new installation [Figure 12] FIG. 10 is a diagram showing a schematic configuration of a collaboration system for voice recognition devices in a house according to a second embodiment; [Figure 13] FIG. 10 is a diagram showing an example of usage status data in which usage status for each time period for a person is recorded. [Figure 14] FIG. 10 is a diagram showing an example of call priority data that records call destination priorities for each person by time period. [Figure 15] FIG. 1 is a diagram showing an example of connecting a voice recognition device and a mobile communication terminal via a dock. [Figure 16] Flowchart showing the process of switching to answering machine mode DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, examples of embodiments of the present invention will be described with reference to the drawings. The same functions are denoted by the same reference numerals, and duplicate explanations will be omitted.

[0010] First Embodiment 1 is a hardware configuration diagram of a voice recognition device 1 according to this embodiment. Device 1 can be a device dedicated to voice recognition, a mobile phone terminal, a smartphone, a personal computer, or any other device. Electronic devices with existing communication functions, such as personal computers, game consoles, etc. The voice recognition device 1 may also have a communication function such as a wired LAN or a wireless LAN. , wireless communication over mobile phone lines, Bluetooth (registered trademark), proximity wireless such as RFID , etc., and a communication interface corresponding to the communication function may be used. It has one or more of the following.

[0011] Specifically, the voice recognition device 1 includes a CPU 101, a memory 103, an external interface, and a Wired LAN I / F 104, wireless LAN I / F 105, and wireless communication I / F 1 06, an audio input unit 107 (for example, a microphone), an audio output unit 108 (for example, a speaker), a display The output units 109 (for example, liquid crystal displays) are connected to each other via the bus 102. 2 includes a person detection sensor I / F 110, a timer 111, an RTC 112, and a camera 113. may be connected.

[0012] The memory 103 includes an internal memory 1031 made of a volatile memory and a non-volatile memory. It includes a reference memory 1032.

[0013] The human detection sensor I / F110 can be used with any type of sensor, such as a human motion sensor or a sound collection sensor. This is an interface for attaching an external human detection sensor.

[0014] FIG. 2 is a functional block diagram of the voice recognition device 1. As shown in FIG.

[0015] The voice recognition device 1 includes a voice processing engine 120. The voice processing engine 120 , mainly includes a voice processing unit 1201, a voice recognition unit 1202, a hot word detection unit 1203, a voice analysis unit 1204, and a voice recognition unit 1205. The image processing unit 1202 includes an analysis unit 1204 and a feature extraction unit 1205 .

[0016] The audio processing engine 120 is a CPU 101 that processes audio data stored in the reference memory 1032. The program is read and loaded into the internal memory 1031, and the program is processed according to the audio processing program. By executing the above process, the functions of the voice processing engine 120 are realized.

[0017] Specifically, when a person speaks something to the voice recognition device 1, the voice is input to the voice input unit. 107, and the voice (analog data) is converted into digital voice data. will be done.

[0018] The audio processing unit 1201 performs adjustments such as removing ambient noise contained in the audio data.

[0019] The voice recognition unit 1202 performs voice recognition processing to convert voice data into character string data.

[0020] The hot word detection unit 1203 detects whether the character string data is a start of an operation for the voice recognition device 1. A specific phrase that requests startup, such as restarting the device or returning it from standby mode (hereinafter referred to as "hot word") Determine whether the string data contains

[0021] FIG. 3A is a diagram showing an example of hot word data 150 stored in reference memory 1032. The hot word data 150 includes a registered hot word 1501 and the registered hot word A type 1502 that specifies the operation of the voice recognition device 1 of the hotkey 1501 and a registered hotkey Device-specific information for identifying the voice recognition device 1 that is called using the code 1501 is specified. The hot word detection unit 1203 is data associated with the call destination 1503. , hotword detection based on whether the character string data is listed in the hotword data 150 Make an exit.

[0022] The voice analysis unit 1204 interprets what is spoken to the voice recognition device 1, that is, converts the text into a character string. response data and a control command corresponding to the data, and outputs the response data from the voice output unit 108. and outputs a control command to the CPU 101, and executes the processing indicated by the audio data. 1. For example, a control command may be a command to play a specific piece of music.

[0023] The voice recognition unit 1202 is not the voice recognition device 1 but is connected to the voice recognition device 1. The external server 201 (see FIG. 4) may be provided with the voice recognition processing. This reduces the load on the voice recognition device 1. When performing speech recognition processing in 1, the amount of data communication with the external server 201 can be reduced. can.

[0024] Furthermore, as a characteristic function of the voice recognition device 1 according to the first embodiment, In addition to operating independently like a smartphone, it can also detect voices spoken to it in designated areas such as inside a house or building. A private communication network (hereinafter referred to as the home LAN 210: Fig. 4) is installed in the The voice recognition device 1 can be transferred to another device (see reference) and conversation can be conducted. Therefore, the hot word detection unit 1203 switches to a conversation mode for conversation. The voice recognition device 1 detects a hot word that is the starting point of the transition. Extract voice and image characteristics from people who can participate in conversations via 210 and match them with registered data It has a feature extraction unit 1205 that performs the determination.

[0025] FIG. 3B is a diagram showing an example of voice authentication data 160 stored in reference memory 1032.

[0026] The voice authentication data 160 uniquely identifies people who can participate in a conversation over the home LAN 210. The speaker 1601 as the information to be used, and the speaker's attributes, for example, the speech recognition device 1 and the communication network The master has the authority to set up the network, or the master does not have the authority to set up the network and is connected to the home LAN 210. The speaker type 1602 indicates whether the person is a "general" who only has the right to participate in the conversation, and the voice of each individual. This is data associated with a speaker template 1603 that indicates the characteristics.

[0027] FIG. 4 is a diagram showing a linkage system 100 using the voice recognition device 1, which is used in a home. This example shows an example of a case where the device is installed inside a house. This embodiment can be applied to spaces where specific people gather, such as offices and classrooms. is.

[0028] (First calling action) In Figure 4, the first voice recognition system is installed in each room 1, room 2, room 3, and room 4 of the house. Device 1A, second voice recognition device 1B, third voice recognition device 1C, fourth voice recognition device The voice recognition device 1D has the same functions as the voice recognition device 1 shown in FIG. The first to fourth voice recognition devices 1A to 1D in the house are installed in rooms 2 to 4, respectively. The second AP 2032 to the fourth AP 2034 are access points or wireless repeaters. The first to fourth audio signals are then transmitted to the router 202 in the room 1 via the respective audio signals. Each of the recognition devices 1A to 1D is connected to the external Internet 200 via a router 202. is connected to.

[0029] The first call operation in the linkage system 100 will be described below in accordance with the steps in FIG. 5. When starting the process of this step, the first to fourth voice recognition devices 1A to 1D are all powered on, and the audio input unit 107 and the audio processing engine 120 This state is called standby mode.

[0030] When person A in room 1 calls out to first voice recognition device 1A, person A's The voice is captured by the voice input unit 107 of the first voice recognition device 1A, and then the voice is subjected to hot word detection. The output unit 1203 determines whether the first hot word indicates a start request. When the detection unit 1203 determines that the call voice is the first hot word (S101 / Ye s), and the voice response data of predetermined fixed phrases such as "Is there anything I can help you with?" is sent to the voice output unit. The hot word detection unit 1203 detects the first hot word and outputs it to the voice analysis unit 108. 204, and the voice analysis unit 1204 issues an execution command to the normal mode execution unit 1406. As a result, the first voice recognition device 1A switches to the normal mode. The first mode is an operation mode in which all of the functions of the first voice recognition device 1A can be executed.

[0031] If the hot word detection unit 1203 determines that the word is not a hot word (S101 / No), , maintain standby mode.

[0032] Next, Person A calls out to Person B in another room, saying, "Mr. B!" The voice is input as voice data via the voice input unit 107, and then the voice is input to the voice processing engine After the sound processing unit 1201 in 120 performs adjustments such as removing ambient noise, the hot word detection unit 1203 whether it is a second hot word (a hot word instructing a switch request to conversation mode) is determined (S102).

[0033] The voice data of the call "Mr. B!" is used as a hot word and the hot word data is 150. The hot word detection unit 1203 detects that the word "Mr. B!" is registered in advance in the second hot word detection unit 1203. If it is determined that the word is a hot word (S102 / Yes), the voice analysis unit 1204 receives the second hot word. In this embodiment, a hot word that triggers the transition to conversation mode is output in the home. The name of a living person, but not limited to this, can also be another set phrase, for example, It can be anything, such as "talk" or "connect."

[0034] The voice analysis unit 1204 analyzes that the second hot word is a request to switch to the conversation mode, The necessary control commands are selected (S103). A mode switching command to switch chairs 1A to 1D to conversation mode, and a command to say "Mr. B!" The voice data is transmitted from the first voice recognition device 1A to the second to fourth voice recognition devices 1B to 1D. The audio transfer command and the audio playback frame are transmitted to the audio output unit 108. This applies to the following brands.

[0035] The conversation mode execution unit 1403 also starts up in the first voice recognition device 1A, and the device switches to the conversation mode. Replace (S104).

[0036] The command transmission unit 1402 of the first voice recognition device 1A receives the voice transfer command. "Mr. B!" is sent to the second to fourth voice recognition devices 1B to 1D via the communication control unit 1410. The voice data is transferred and the mode is switched to conversation mode for each person. The first voice recognition device 1A also transmits a command and a voice playback command. The transmitting unit 1402 transmits the voice data to the second to fourth voice recognition devices 1B to 1D. Then, measurement of the elapsed time from the transfer of the data is started (S105).

[0037] Each of the second to fourth voice recognition devices 1B to 1D has call voice data. The voice output unit 108 reproduces the voice "Mr. B!", and the conversation mode execution unit 1403 is activated. The voice data is played in the order of the second to fourth voice recognition. The devices 1B to 1D may play audio simultaneously, or in a predetermined order. The voice may be output from the first to fourth voice recognition devices 1B to 1D. For example, the order in which the voice recognition devices were installed or the priority of each room can be used. The order may be the same as the order shown in FIG. 14.

[0038] Person B in Room 2 answers and responds with a response such as "Yes!" When the voice is received, the second voice recognition device 1B receives the voice as voice data via the voice input unit 107. The response voice data is returned to the first voice recognition device 1A that originated the call. Then, the second voice recognition device 1B is associated with the person B. The linking information that the person in room 2 where room 1B is present is person B is the first, third, and fourth It is also registered as a shared voice recognition device 1A, 1C, and 1D.

[0039] FIG. 6 shows an example of location estimation data.

[0040] As an example of linking and registering person B and second voice recognition device 1B, When A receives the above response voice data, it will respond as "Person B = second voice recognition device 1B." The location estimation data may be added to the location estimation data (see FIG. 6) registered in advance in the reference memory 1032. The first to fourth voice recognition devices 1A to 1D in the home are identified by the MAC address of the device and the A specific identifier such as an IP address assigned to a device in the LAN 210 may be used.

[0041] Furthermore, the location estimation data is transmitted from the first voice recognition device 1A to other voice recognition devices in the house. The registration request is sent to the second to fourth voice recognition devices 1B to 1D. The above-mentioned information is stored in the reference memory 1032 of the second to fourth speech recognition devices 1B to 1D. The current estimation data is already stored in the reference memory 1032 in each of the first to fourth speech recognition devices 1A to 1D. If the person is already registered, the linking operation is not performed. The name part included in the voice of the call "You!" is processed by the voice processing engine 120 or an external When extracted on the server 201 and used as name data, "Person B's name = second speech recognition device This will allow for matching with "Chair 1B."

[0042] Please note that we will explain how to register hot words to determine whether to switch to conversation mode later. This is done when installing the voice recognition device during initial setup or when setting up a new registration, using the method described below.

[0043] The command transmission unit 1402 of the first voice recognition device 1A receives the time measured by the timer 111. If the elapsed time is equal to or greater than the waiting time threshold for determining whether or not a response is received (S106 / Yes ), and for the device that has not responded among the second to fourth voice recognition devices 1B to 1D, A command to return from the talk mode to the standby mode is sent (S107).

[0044] A response of voice data from at least one of the second to fourth voice recognition devices 1B to 1D. If there is a response within the time threshold (S106 / No), the first voice recognition device The second voice recognition device 1B is kept in conversation mode with the second voice recognition device 1A. Hold (S108).

[0045] The conversation mode execution unit 1403 switches to the conversation mode in step S104. After that, the voice input to the voice input unit 107 of the first voice recognition device 1A and the step In step S105, the voice input to the voice input unit 107 of the second voice recognition device 1B is Speech is transmitted and received between a first speech recognition device 1A and a second speech recognition device 1B.

[0046] Either the first voice recognition device 1A or the second voice recognition device 1B is in a conversation mode. When the third hot word that ends the process is detected (S109 / Yes), The device that received the voice command, for example, the second voice recognition device 1B, switches to standby mode (S110 ), together with a command to switch the first voice recognition device 1A to standby mode. In response to this, the first voice recognition device 1A also enters the standby mode. and conversation mode ends.

[0047] Also, if the second hot word is not detected in step S102 (S102 / No ), the first voice recognition device 1A maintains the normal mode without switching to the conversation mode. The process ends as it is (S111).

[0048] In the above example, the first voice recognition device 1A of the person who first called out is the master. The second to fourth voice recognition devices 1B to 1C of the person in the house 2 who is called become star devices. Although control instructions are given to each of the 1D, they are not limited to this. The second voice recognition device 1B that responded may become the master.

[0049] Another method of linking is, for example, when the fourth voice recognition device 1D is installed in a home for the first time. , and registers that the main user or operator of the fourth voice recognition device 1D is person D. By doing so, the first to third voice recognition devices 1A can be connected to the home LAN 210 immediately after installation. 1C, and the linked data is stored in the reference memory in the first to fourth speech recognition devices 1A to 1D. It is also possible to register "Person D = fourth voice recognition device 1D" in 1032.

[0050] The above example is a call from the first voice recognition device 1A to the second voice recognition device 1B. This is an example, but this procedure can be applied to all voice recognition devices in the home. , even if a voice recognition device other than the above examples calls to another voice recognition device You can do the same.

[0051] In the above example, the first to fourth voice recognition devices 1A to 1D use wireless LAN for communication with each room. Although it uses AN, it can be connected via wired LAN or via a mobile phone line. Furthermore, it is possible to change the communication mode to a different interface only when in conversation mode. For example, in normal mode, wireless LAN is used, and in conversation mode, Bluetooth is used. It would also be possible to use a different wireless system such as h.

[0052] (Second calling action) As in the above embodiment, once the voice recognition device 1 and the person are linked, When an object A and a person B have a conversation, a first voice recognition device 1A used by the person A and a second voice recognition device 1B used by the person B are used. The communication between the second voice recognition device 1B used by the user is opened, and the communication between the second voice recognition device 1B and the other voice recognition devices is opened. By closing communication between the two, the confidentiality of the conversation content can be increased.

[0053] For example, if person A calls person B for the second time or later, the command of the first voice recognition device 1A is The transmitting unit 1402 transmits a second location estimation data associated with the person B registered in the location estimation data (FIG. 6). Establish communication with the voice recognition device 1B (by sending a mode switching command and responding to it). Communication is established by the second voice recognition device 1B, and the voice data is sent to the second voice recognition device 1B.

[0054] The conversation mode execution unit 1403 subsequently controls the first voice recognition device 1A and the second voice recognition device 1B. The voice data is transmitted only between person A and person B, and person A and person B are not in a direct conversation. This allows you to send voice commands to all voice recognition devices in your home. There is no need to send data.

[0055] During this conversation mode, for example, if person B responds within a predetermined time after the call, If no voice data is received, the conversation mode execution unit 1403 of the first voice recognition device 1A It is determined that the conversation partner is absent. Then, the conversation mode execution unit 1403 sends a command. The communication unit 1402 can now communicate with other voice recognition devices in the house that had previously been closed to communication. Restart the channel, i.e., other voice recognition devices (third and fourth voice recognition devices 1C and 1D) Sends voice data to and waits for a response.

[0056] Here, for example, if there is a response from the fourth voice recognition device 1D in room 4, the first voice The voice recognition device 1A starts communication with the fourth voice recognition device 1D and resumes the conversation. In this case, the first voice recognition device 1A uses the location estimation data of the internal reference memory 1032. The data 190 does not store information linking person B with the fourth voice recognition device 1D. It is also possible to consider that person B has temporarily moved to another location, or to consider that person B and the fourth voice Create link information for recognition device 1D, assign a priority, and link person B and the second voice recognition device. The association between person B and the fourth voice recognition device 1D is given a higher priority than the association between person B and the fourth voice recognition device 1D. It is also possible to set this, and connections will be made in order of priority and await a response.

[0057] Furthermore, if there is no response from the fourth voice recognition device 1D, the other voice recognition devices in the house are sequentially searched. The voice data is sent to the recognition device (in this example, the third voice recognition device 1C) and a response is made. Then, the voice data is sent to the third voice recognition device 1C and a response is waited for. If there is no response from any of the voice recognition devices in the rooms within the specified time, The first voice recognition device 1A determines that there is no such voice recognition, and issues a message to the operator, person A, such as " There was no response. Or, as mentioned above, Instead of determining that one voice recognition device 1A is not responding to another voice recognition device, Then, the other voice recognition device determines that there is no response voice from person B within a predetermined time, By returning no-response information to the first voice recognition device 1A, the first voice recognition device 1A becomes no-response. When the call is received, a predetermined response voice such as "There was no response" is output. But that's okay.

[0058] The predetermined response voice data may be stored in advance in the memory 103 or may be stored via the internet. It is also possible to use one stored on an external server 201 on the network 200.

[0059] (Third calling action) In this example, the person detection sensor I / F 110 of the voice recognition device 1 in FIG. It is equipped with an image sensor that can capture images, or a human detection sensor that can detect the presence or absence of people. The device will then connect to the speaker and respond based on the results of the detection. Alternatively, a camera 113 may be used.

[0060] For example, in response to a call from person A, the first voice recognition device 1A recognizes the voices of the people in each room. When transmitting the call voice data of person A to the second to fourth voice recognition devices 1B to 1D, The second to fourth voice recognition devices 1B to 1D each have a person detection sensor to detect the presence or absence of a person. If the room is determined to be empty, a voice recognition device installed in that location will notify the user of the absence. is returned to the first voice recognition device 1A and received by the command transmission unit 1402.

[0061] Then, the command transmission unit 1402 of the first voice recognition device 1A transmits the absence notification in the form of the first voice recognition device 1A. The voice of the call is not output to the voice recognition device that sent the signal to the recognition device 1A. .

[0062] On the other hand, the command transmission unit 1402 of the first voice recognition device 1A transmits the absence notification to the first voice recognition device 1B. The voice data is transmitted to the voice recognition device 1A, and the voice data is The voice recognition device receives the data and plays it back to make a call. This can be done in the same way as in the embodiment.

[0063] The above person recognition and determination may be performed by a commonly used method, such as using an infrared sensor. It is possible to detect the presence or absence of a person from the movement of the person detected by the human presence sensor.

[0064] Furthermore, a camera 113 may be used as a person detection sensor. 04 extracts human features (for example, facial images) from images taken by the camera 113 and determines whether or not a person is present. Furthermore, face recognition data in which face images are associated with people may be stored in advance in a reference memory 1. Check the person's correspondence information stored in 032 to see if the person being called is present in the room. If the person being called is captured by the camera 113 and is determined to be in the room, The first voice recognition device 1A and the second voice recognition device 1B are connected to each other through a communication connection. You will be able to talk to B.

[0065] (Fourth calling action) In another embodiment, person A in FIG. 4 calls person B, and someone other than person B If the user replies, the communication connection between the first voice recognition device 1A and the second voice recognition device 1B is established. A second call will be made without establishing a relationship.

[0066] First, the first to fourth voice recognition devices 1A to 1D installed in each room are provided with the following voice recognition functions: Voice authentication data 160 (FIG. 3B) of people residing in The voice authentication data 160 is extracted by the feature extraction unit 1205 of the voice processing engine 120 from the person's voiceprint, Create a speaker template 1603 using voice intonation or frequency characteristics of the voice. The voice authentication data 160 is stored in the reference memory 1032 of the first to fourth voice recognition devices 1A to 1D. It is generated by storing it in advance as

[0067] The voice authentication data 160 is generated at the time of initial setting of each of the first to fourth voice recognition devices 1A to 1D. Then, the voice recognition unit 1405 of the first voice recognition device 1A can recognize the voice. The voice characteristics of the person who responded are compared with the voice characteristics data of person B registered in the authentication data 160. If they match, the person is determined to be person B, and the determination result is sent to the command transmission unit 1402. In response to this, the command transmission unit 1402 sends the meeting command to the second voice recognition device 1B. Send a command to switch to talk mode.

[0068] If the voice recognition unit 1405 determines that the voice characteristics do not match, it is not person B. and process it.

[0069] The voice authentication data 160 is collected from all the first to fourth voice recognition devices in each room as described above. The voices of the people in each room responding to the call from person A are stored in advance in each of 1A to 1D. It is also possible to compare the voice of person B with the voice of person B that is expected to be called, but instead In addition, the voice recognition data is stored only in the reference memory 1032 of the first voice recognition device 1A, which is a specific master. The authentication data 160 is stored, and whether the voice characteristics match or not is determined on the first voice recognition device 1A. You may make a judgment.

[0070] Alternatively, the voice authentication data 160 may be stored in a device such as a server installed in the home. ~ Voice data and voice authentication data 1 sent from each of the fourth voice recognition devices 1A to 1D 60 to determine whether it matches or does not match.

[0071] Furthermore, voice authentication data 160 is stored on an external server installed outside the home, and voice characteristics are recorded. It's also good to compare.

[0072] As in the example above, by determining the voice characteristics of the speaker 1601, the family members residing in the house can be identified. Prevent outsiders from joining the conversation, other than those authorized by family members to join the conversation. This can prevent this and improve security.

[0073] In addition, if the voice characteristics of a speaker other than the registered speaker are detected, for example, The display output unit 109 of the first voice recognition device 1A on the other side displays an alarm. At this time, the voice output unit 108 outputs a voice such as "There is a response from an outsider." You may want to issue a warning like this.

[0074] (Switch between normal mode and conversation mode 1) FIG. 7 is a flowchart showing the flow of a first switching process between the normal mode and the conversation mode. Contrary to the example in Figure 5, conversation mode is set as the default.

[0075] In this example, after the first to fourth voice recognition devices 1A to 1D are installed, the main power is turned on. In this state, the conversation mode execution unit 1403 is activated and a conversation mode is started in which the operator waits for a call to another person. At this time, the operator can enter a hot word to switch modes, such as "mode change". When a predetermined word such as "ji" is uttered (S201 / Yes), the mode switching unit 14 01 switches to the normal mode (S202), and the normal mode execution unit 1406 starts (S2 03).

[0076] In step S201, if the mode switching hotword is not detected (S20 1 / No), the conversation mode is maintained by the conversation mode execution unit 1403.

[0077] The normal mode execution unit 1406 continues to execute the normal mode execution while the condition for returning to the conversation mode is not satisfied (S204 / No), maintain normal mode.

[0078] When the conditions for returning from normal mode to conversation mode are met (S204 / Yes), the conversation starts again. The mode execution unit 1403 is activated and the mode is returned to the conversation mode. You can also set a hot word for the system to automatically return if there is no response from the operator within a specified time. It may be possible to do so.

[0079] (Switch between normal mode and conversation mode 2) FIG. 8 is a flowchart showing the flow of the second switching process between the normal mode and the conversation mode. be.

[0080] In this example, the voice recognition device 1 starts up the voice input unit 107 and detects only the presence or absence of voice. When the voice input unit 107 detects a voice (S301 / No), the process continues to monitor the voice (S301 / Yes), the hotword detection unit 1203 detects that the detected voice is The first hotword to request activation (request activation of normal mode) or activation of conversation mode It is determined whether the second hot word is the requested hot word (S302). If the voice does not match any of the two hot words (S302 / No), the process returns to the voice detection process.

[0081] If the first hotword is detected (S302 / first hotword), execute normal mode For example, the voice recognition device 1 is started up by the setting unit 1406 (S303). If the specified nickname is called, the process will proceed as normal mode.

[0082] Also, if the second hot word is detected (S302 / second hot word), the conversation mode The execution unit 1403 is started (S304). For example, when calling the name of a family member or someone in the house, If so, it is determined to be in conversation mode and the following processing is carried out.

[0083] The first and second hot words can be set in advance, or can be set by the operator after installation. You can also change the hotword to suit your needs. To use this function, you need to install the dedicated application software for setting up on your smartphone or computer. Alternatively, a master operator 501 may be determined in advance. Star Operator 501's voice can be used to change hot words on the first to fourth voice recognition devices 1A to 1D. Alternatively, the master speech recognition device, e.g., the first speech recognition device 1 A to slave devices existing in other homes, for example, second to fourth voice recognition devices 1B to 1 At this time, the master operator 501 may issue a change instruction to D. Hotwords can be easily changed by only allowing them to be changed if they are recognized as such. The voice setting of the master operator 501 is set to a voice setting dedicated to the smartphone or PC. The voice of the master operator 501 is registered as the first voice. The voice feature data is extracted by the feature extraction unit 1205 in the voice processing engine 120 of the voice recognition device 1A. You can create it and register it in the reference memory 1032, or you can use it on a smartphone or a PC. It is also possible to create the voice authentication data 160 and register it in the reference memory 1032. If a home server 601 connected to a home network environment is provided, voice feature data is stored. The home server 601 stores the information and the server and the voice recognition device cooperate to call the information. It is also possible to compare and confirm the characteristics of the voice spoken with the characteristics of the voice in the stored data.

[0084] (New installation setting 1) FIG. 9 is a conceptual diagram showing the first setting process at the time of new installation.

[0085] When installing a new voice recognition device1 in your home, it is important to The electronic device 401 is provided with dedicated application software (initial setting application software). Then, the network connection with the router 202 in the home is established. Continued setup, device registration if there is an existing voice recognition device1, voice recognition with family members and other people in the home data relating to the correspondence with the recognition device 1, voice authentication data 160 of family members, etc. In this example, a newly installed voice recognition device and a wireless device in the home are configured. To set up a connection with a communication device such as a wired router, use the above smartphone or computer. How to set it up or use automatic methods like WPS (WiFi Protected Setup) The connection is set up using a dynamic setting method.

[0086] In addition, the correspondence between the above persons and the voice recognition device 1 is that person A mainly uses The device that person B mainly uses is the first voice recognition device 1A, and the device that person B mainly uses is the second voice recognition device 1B. The voice recognition device 1B is used to set up the association with the smartphone or PC. This is done on the application software.

[0087] Furthermore, the person's voice authentication data 160 can be acquired from an existing device, such as a third voice recognition device 1C. It is also possible to read and share the data stored on the electronic device. It is also possible to understand the data in advance and set it.

[0088] (New installation setting 2) FIG. 10 is a conceptual diagram showing the second setting process at the time of new installation, and FIG. 11 is a conceptual diagram showing the second setting process at the time of new installation. 10 is a flowchart showing the flow of a second setting process when the camera is placed.

[0089] In this example, a master operator 501 who has the authority to set up communication devices in the home can Set up the connection for the new voice recognition device 1S.

[0090] First, the master operator 501 starts speaking to the new voice recognition device 1S. When the input unit 107 receives a voice input, the connection setting process is started. Measurement begins (S401).

[0091] The hot word detection unit 1203 of the new voice recognition device 1S detects the fourth hot word for initial setting. When the device detects the W501 (S402 / OK), the initial settings for the new voice recognition device 1S are Specifically, the initial setting unit 1408 starts the initial setting process. - The voice data and initial setting request data of the operator 501 are transmitted to an existing device in the home, for example, The first to fourth voice recognition devices 1A to 1D and the home server 601 are transmitted (S403 The above transmission process is carried out within a predetermined time (S404).

[0092] The reason for limiting the transmission process to a predetermined time is to prevent the initial setting request data from spreading outside the home during transmission. This is to reduce the possibility that the voice data of the master operator 501 will be intercepted.

[0093] The transmitted request data and voice data are transmitted to the first to fourth voice recognition devices already installed in the home. The signal is received by the stations 1A to 1D (S403).

[0094] The feature extraction unit 1205 of each of the first to fourth speech recognition devices 1A to 1D extracts the broad The voice data transmitted based on the casted initial setting request data is sent to the master operator 50. The feature extraction unit 1205 verifies whether the character is one of the first to fourth speech recognition functions (S405). The sound of the master operator 501 stored in the reference memory 1032 of each of the devices 1A to 1D Voice features are extracted from the speaker template and the broadcasted speech data. The data is extracted and compared, and if it matches (S405 / OK), it is applied to the new voice recognition device 1S. Then, the initial setting is performed (S406), and the connection setting process is completed.

[0095] If the determination results in steps S402 and S405 are NG (S402 / NG), 405 / NG) also ends this process.

[0096] The initial setting is performed by setting all the voices in the house among the first to fourth voice recognition devices 1A to 1D. This is done by a master voice recognition device that controls all voice recognition devices. For example, a device installed in the living room of a house (used relatively frequently by family members) This role will be played by the first device installed in the home, or by the voice recognition device installed in the home. may be set by the master operator 501 as the master voice recognition device.

[0097] Alternatively, the home server 601 in the home can be used as a master device to verify the above voice data. Or, the home server 601 may execute the initial setting of the newly connected device. When this is implemented, the voice data and initial setting request data of the master operator 501 are input to the existing first The fourth voice recognition device 1A to 1D receives the voice signal and transfers it to the home server 601, or The home server 601 receives the data. The home server 601 then receives the data. The voice template, which is the voice feature data of 01, is stored and the voice feature is checked to see if it matches. If the verification is successful, the home server 601 sends the home ID to the new voice recognition device 1S. Various communication settings are instructed to enable connection to the internal LAN 210 .

[0098] Second Embodiment In this embodiment, a conversation system in a home using a voice recognition device 1 is FIG. 12 shows the voice recognition device in the home in the second embodiment. 1 is a diagram showing a schematic configuration of a linkage system 100a of a chair 1. FIG.

[0099] The difference from FIG. 4 is that a home server 601 is provided on the home LAN 210. The server 601 stores voice authentication data including voice data of people present in the house and voice features. The home server 601 stores the first to fourth voice recognition devices 160. The system detects the presence or absence of people in the home using voice data and data sent from each of A to D. It is determined which of the first to fourth voice recognition devices 1A to 1D a person present in the vicinity of is present. Always monitor.

[0100] As a result, even when person A calls person D in FIG. 12, the call of person A The first voice recognition device 1A that receives the voice recognition signal 1B is determined to be closer to the person D from the home server 601. Information on the voice recognition device to be determined (fourth voice recognition device 1D) is obtained.

[0101] Then, voice data is transmitted only to the first voice recognition device 1A to the fourth voice recognition device 1D. The data is sent to the person you want to call, and you can call them without having to check their location each time. It will be possible to have a conversation only with the voice recognition device.

[0102] In FIG. 12, the people present are person B in room 2 and person D in room 4. Each of the first to fourth voice recognition devices 1A to 1D is connected to a person detection sensor I / F provided in the device itself. The presence status is grasped at 110, and the result is transmitted to the home server 601. The voice recognition device 1A notifies the home server 601 of which voice recognition device is detecting a person. By inquiring whether the second voice recognition device in room 2 and room 4 is the Priority is given to communication connections with 1B and the fourth voice recognition device 1D.

[0103] Furthermore, the home server 601 collects information such as whether there is a person or not and whether there is a voice or not, It is possible to always know which person is in which room. In this way, the first voice recognition The home server 601 checks the destination of the call data from the device 1A to the person D. The call data is sent to a fourth voice recognition device 1D located in a room 4 where the person D is present. This can be done.

[0104] Next, person D responds to the call played on the fourth voice recognition device 1D, A communication connection is established between the voice recognition device 1A and the fourth voice recognition device 1D, and person A A conversation between person A and person D becomes possible.

[0105] In the above example, the home server 601 controls the first to fourth voice recognition devices 1A to 1D. The presence status of each room is monitored by the person detection sensors installed in each room. Or, in addition, data on the usage status of each voice recognition device may be used.

[0106] FIG. 13 shows the usage status data 170 that records the usage status of person A for each time period. 14 shows an example of the time for person A determined based on the usage status data 170. An example of call priority data 180 for each time zone is shown. Similar data will be created for the above.

[0107] The reference memory 1032 of each of the first to fourth speech recognition devices 1A to 1D stores a usage status data. For example, for person A, When a call is made and the first voice recognition device 1A responds, the first to fourth voice recognition devices Each of the devices 1A to 1D has its own usage status data 170 and call priority. The response record and call priority are written into the data 180 and updated.

[0108] Furthermore, the first voice recognition device 1A receives the updated usage data 170 and the call priority. The ranking data 180 is broadcast to the home LAN 210. Second to fourth speech recognition Each of the devices 1B to 1D receives the updated usage status data 170 and the call priority. The usage status data 170 stored in the reference memory 1032 of the own device using the ranking data 180 and updating the call priority data 180.

[0109] In this state, Person B calls Person A at 7:00 on Monday from the fourth voice recognition device 1D. The command transmission unit 1402 of the fourth voice recognition device 1D sends the command in the order of call priority. The priority data 180 is referred to, and the priority data 180 is selected from the first to third voice recognition devices 1A to 1C other than the own device. In descending order of priority, the second voice recognition device 1B, the first voice recognition device 1A, the third voice recognition device 1B, and so on. The data is read out in order of the recognition device 1C.

[0110] The call priority data 180 is not based only on the usage status data 170 but also on the user For example, if you know that you are near a specific device during a certain time period, If so, temporarily change the call priority data 180 for that device. The priority may be set to 1.

[0111] Also, for example, in the first usage status data 170, if person A uses the second voice recognition device 1 It can be seen that B is frequently used between 8 PM and 8 AM on Saturdays and Sundays. If the time is within that time, it is determined that person A is likely to be in the room, and a connection in conversation mode is established. During times other than those mentioned above, the home server 601 treats person A as absent. It can also be understood.

[0112] (Conversation with someone outside the home) In the above example, if the called person D is determined to be absent, It is also possible to transfer audio data to communication devices such as smartphones owned by users.

[0113] In this case, the internet address, line information, and device ID of Person D and their communication device By registering such device information in the reference memory 1032 of the home server 601, Call data is transferred to the communication device along with the information.

[0114] When the call information arrives at the communication device owned by person D, the device notifies the person by displaying the screen, outputting a sound, vibrating, etc. When person D responds, the first voice recognition device in the home A call can also be initiated between 1A and a communication device owned by person B outside the home.

[0115] If person D does not respond to the call at this stage, the home server 601 will play the first sound. The first voice recognition device 1A issues absence notification data to the second voice recognition device 1A. If there is no response from the object D, a predetermined voice output will be generated, such as "There is no response at the moment." conduct.

[0116] In the above-described embodiment, the first voice recognition device 1A is used to recognize the other second to fourth voice recognition devices. Although an example of calling out to the voice recognition devices 1B to 1D is shown, the present invention is not limited to this. Instead, the call can be made from any of the second to fourth voice recognition devices 1B to 1D. Therefore, it is possible to call other voice recognition devices from any voice recognition device in the home. In addition, multiple voice recognition devices according to this embodiment can be installed. When installing a new system, the system can be expanded using the above installation method.

[0117] (Notification from outside the home) FIG. 15 shows an example in which the voice recognition device 1 and the portable communication terminal 71 are connected by a dock 701. vinegar.

[0118] The voice recognition device 1 further includes a dock 701. The dock 701 is connected to a mobile communication terminal 7 7. A charging control interface 711 that charges the battery 1 and a communication interface 712 that communicates via a connection terminal. The control interface 712 includes: Serial Bus) or certain mobile communication terminal interfaces, or This can be achieved through wireless charging and wireless communication functions.

[0119] When a call is received by the mobile communication terminal 71, the mobile communication terminal 71 sends a communication control input The voice output unit 108 outputs the voice to the voice recognition device 1 via the interface 712. It will output an incoming call notification sound such as "This is a call" or "This is an email."

[0120] The owner of the mobile communication terminal 71 may ask, "Please play it," "Who is it from?" or "What is the content?" When the user responds, the mobile communication terminal 71 is instructed to receive the call or transfer the contents of the mail. , announce the name of the person you're notifying, act as a speakerphone for calls, or send emails. If you do this, you can have the email contents output as audio.

[0121] Furthermore, if the owner of the mobile communication terminal 71 is not present, whether or not there is a response within a predetermined time 12. The mobile communication terminal 601 in FIG. 12 is assumed to be in another location in the house. The voice recognition device 1 that is determined to be closest to the current location of the last 71 owners is identified, The incoming call notification is transferred to the voice recognition device 1. Based on the transferred incoming call notification, The same operation as the above call operation is performed.

[0122] As a method for recognizing the current location of the owner of the mobile communication terminal 71, a home service The server 601 detects the usage status of each voice recognition device 1 in the house, The extracted data of the voice spoken to the device 1, the sounds picked up by each voice recognition device 1, 701 and the mobile communication terminal 71, the person in the house is using which voice recognition device It may also be possible to determine whether the user is near S1.

[0123] Furthermore, devices such as the mobile communication terminal 71 that are usually worn by their owners (wearable devices) (or any other device) is close enough to the voice recognition device 1, such as the connection with the dock 701. Can the camera of the voice recognition device 1 communicate using short-range communication that can determine whether the person is in the vicinity? If you can confirm that the device is in the same room by calling 113, the device owner will be in that room. It is assumed that the person in the room is a voice recognition device 1, and the priority of the call to that device is increased. It is also acceptable to take measures such as

[0124] <Third embodiment> (Monitoring the home while the person is away) The third embodiment is an application of the configuration of the second embodiment shown in FIG. 12 to a different usage mode. FIG. 16 shows a flowchart of the process of switching to the answering mode. This is a flow chart.

[0125] Anyone in the home who is authorized to access each voice recognition device1 can use the answering machine. The fifth hot word to switch to the mode, for example, "Please stay home!" When any one of the voice recognition devices 1A to 1D detects the fifth hot word (S501 / Yes), the feature extraction unit 1205 of the voice recognition device that detected the fifth hot word The voice of the person registered in the authentication data 160 as a person authorized to access the home LAN The characteristics are compared with the voice characteristic data extracted by the voice recognition device 1 for judgment.

[0126] This determination can be made by using the methods described in the above embodiments. In addition, the procedure for checking the hotword and determining whether to allow access can be performed either first or second. stomach.

[0127] If the feature extraction unit 1205 determines that they match (S502 / Yes), the command transmission unit 1 The mobile phone 402 sends an instruction to the home server 601 to switch to the answering machine mode (S503).

[0128] When the home server 601 receives the instruction to switch to the answering machine mode, the home server 601 waits a predetermined time after the reception. After a certain time has passed, the first to fourth voice recognition devices 1A to 1D in the house will all be able to recognize the voices at a volume above a predetermined volume. When the above sound is detected, an abnormal sound notification is sent to the home server 601. In order to shift to the answering machine mode, an instruction to shift to the answering machine mode is issued (S504).

[0129] Each of the first to fourth voice recognition devices 1A to 1D has a mode switching unit 14 01 switches to the answering machine mode, and the answering machine mode execution unit 1407 processes the answering machine mode. (S505).

[0130] At least one of the first to fourth voice recognition devices 1A to 1D receives voice messages during an answering machine mode. When the first to fourth voice recognition devices 1A to 1F detect the voice (S506 / Yes), The 1D feature extraction unit 1205 extracts voice feature data of the person (family member) registered in the voice authentication data 160. If the detected voice matches the voice feature data registered in the voice authentication data 160, the voice is compared with the voice feature data. If so (S507 / Yes), the mode switching unit 1401 switches from the answering machine mode to the normal mode. The system is restored (S508).

[0131] In step S507, the feature extraction unit 1205 extracts the detected voice data as voice authentication data. If it is determined that the voice characteristic data does not match that of the person (family member) registered in the database 160 (S507 / No), the mode switching unit 1401 switches to the alert mode (S509), and the alert mode execution unit 1409 starts up.

[0132] The alert mode execution unit 1409 is configured to execute the alert mode, for example, the first to fourth voice recognition devices 1A to The camera 113 installed in each 1D can be activated to record image data and audio data. The input unit 107 executes recording processing of the detected voice. When the home server 601 receives the abnormality information, the home server 601 may transmit the abnormality information. Sending emails and other messages to mobile communication terminals 71 such as mobile phones and smartphones owned by registered family members Send an alarm.

[0133] The family member who receives the alarm can also connect to the home server 601 for communication. The server 601 converts the sound picked up by the voice input unit 107 of the voice recognition device 1 that detected the abnormal sound into a voice signal. The data is received and sent directly to the family members to check the situation inside the home. It is also possible to do so.

[0134] While the security mode is active, if the conditions for canceling the security mode are met, for example, a call from a family member who has just returned home When the first hot word that commands the return to normal mode is detected (S510 / Yes), the mode The mode switching unit 1401 cancels the alert mode and returns to the normal mode (S508).

[0135] Although the above example shows the use of multiple voice recognition devices 1, the voice recognition device 1 can be used alone. When the voice recognition device 1 operates independently, the voice recognition device 1 is connected to the home server 6. Alternatively, instead of the home server 601, the home server 601 may be provided with a function to be executed by the Internet. It is also possible to use a cloud server or the like that exists on the internet.

[0136] According to this embodiment, the voice recognition device 1 in the home is linked to recognize the voices of the family members while they are away. The alarm mode can be activated based on whether there is an input and whether the voice characteristics are registered or not. This allows the voice recognition devices 1 installed in multiple rooms to be used to detect intruders into a home. If an abnormality is detected, the voice recognition device1 throughout the house will activate alarm mode and detect the intruder. It is possible to track visitors, record their movements and facial images, and notify family members.

[0137] In addition, if the above abnormal sounds are detected, a warning sound or voice will be emitted from the voice recognition device 1. It is also possible to output a warning sound such as a siren, or a voice message such as "Who is it?" If you register such sounds, they can be played back when an abnormal sound is detected, which can be expected to have a crime prevention effect. .

[0138] The above embodiment describes a one-to-one call between voice recognition devices. However, this is not limited to this, and one-to-many and many-to-many communication modes are also possible. In this case, the person's voice data in conversation mode is sent to multiple voice recognition devices.

[0139] According to this embodiment, by using a plurality of voice recognition devices in cooperation with each other, it is possible to In other words, communication within the home can be carried out via the network. It is possible to communicate with voice recognition devices installed in other locations as if they were in the same room or location. This allows users to communicate with each other like a human, and it is possible to use voice recognition devices in other locations. This allows for smooth communication with others. [Explanation of symbols]

[0140] 1: Voice recognition device 100: Collaboration system 100a: Collaboration system 101: CPU 102: Bus 103: Memory 104:Wired LANI / F 105: Wireless LAN I / F 106: Wireless communication I / F 107: Audio input section 108: Audio output section 109: Display output unit 111: Timer 113: Camera

Claims

1. An information processing device, an audio input unit; an audio output unit; a communication interface for transmitting and receiving data to and from a voice recognition device; a storage unit that stores voice authentication information linking a person who is permitted to have a conversation using the voice recognition device with voice feature information of the person, and first usage status information linking the person who is permitted to have a conversation with the number of responses from each voice recognition device of the person; a control unit, The communication interface includes: Communicating with a plurality of speech recognition devices; Transmitting voice information input from the voice input unit to the voice recognition device; receiving speech information from the speech recognition device; The audio output unit outputting the audio information received by the communication interface; The control unit If the voice input from the voice input unit includes a conversation activation hot word, a conversation mode is executed; The voice input unit Accepts input of a call voice to call a specific person, The control unit Detecting whether the voice feature information of the calling voice matches the voice authentication information stored in the storage unit; If a match is detected with the voice authentication information, controlling the plurality of speech recognition devices to transmit a control command for activating the conversation mode in descending order of the number of responses in the first usage status information; If the specific person does not respond, control is performed to transmit the voice information to a communication device of the specific person.

1. An information processing device comprising:

2. An information processing device, an audio input unit; an audio output unit; a communication interface for transmitting and receiving data to and from a voice recognition device; a storage unit that stores voice authentication information that links a person who is permitted to have a conversation using the voice recognition device with voice feature information of the person; A human sensor and a control unit, The communication interface includes: Communicating with a plurality of speech recognition devices; Transmitting voice information input from the voice input unit to the voice recognition device; receiving speech information from the speech recognition device; The audio output unit outputting the audio information received by the communication interface from the audio output unit; The voice input unit Accepts voice input to call a specific person, The control unit If the voice input from the voice input unit includes a conversation activation hot word, a conversation mode is executed; If the specific person does not respond, control is performed to transmit the voice information to a communication device of the specific person; Furthermore, if there is an answering machine hot word, the system controls to execute an answering machine mode that detects the intrusion of a person based on the detection of the human sensor, and controls to stop the answering machine mode when a match between the voice feature information and the voice authentication information is detected during the execution of the answering machine mode.

1. An information processing device comprising:

3. 3. The information processing device according to claim 1, When the control unit detects the conversation starting hotword, the control unit transmits voice information in which the conversation starting hotword is detected and a command to play the voice information to the voice recognition device.

1. An information processing device comprising:

4. 2. The information processing device according to claim 1, the control unit transmits the control command to the voice recognition device when a match between the voice feature information and the voice authentication information is detected.

1. An information processing device comprising:

5. 2. The information processing device according to claim 1, further comprising an imaging unit, the storage unit further stores face authentication information linking a person who is permitted to have a conversation using the voice recognition device with a captured image of the person; The control unit further detects whether the captured image captured by the imaging unit matches the face authentication information, When a match between the captured image and the face authentication information is detected, the control command is transmitted to the voice recognition device.

1. An information processing device comprising:

6. 2. The information processing device according to claim 1, the control unit transmits the control command to the voice recognition device that output the person detection information based on person detection information indicating that the voice recognition apparatus has detected the presence of a person.

1. An information processing device comprising:

7. 3. The information processing device according to claim 1, further comprising a timer; the voice input unit accepts a voice input for calling a specific person; the control unit transmits to the voice recognition device a voice that calls the specific person and a playback command that causes the voice recognition device to play back the voice; acquiring, from the timer, the voice calling the specific person and the time elapsed since the playback command was transmitted; outputting, from the audio output unit, a response message notifying that there is no response message from the specific person after the elapsed time has exceeded a predetermined waiting time; 1. An information processing device comprising:

8. 2. The information processing device according to claim 1, the control unit transmits to all of the plurality of voice recognition devices a voice calling the specific person and a playback command for causing each of the plurality of voice recognition devices to play back the voice; when the communication interface receives a response message from the specific person from one of the plurality of voice recognition devices, the communication interface maintains communication with the voice recognition device that transmitted the response message and disconnects communication with the remaining voice recognition devices; 1. An information processing device comprising:

9. 3. The information processing device according to claim 2, the voice input unit accepts a voice input for calling a specific person; the control unit transmits to all of the plurality of voice recognition devices a voice calling the specific person and a playback command for causing each of the plurality of voice recognition devices to play back the voice; when the communication interface receives a response message from the specific person from one of the plurality of voice recognition devices, the communication interface maintains communication with the voice recognition device that transmitted the response message and disconnects communication with the remaining voice recognition devices; 1. An information processing device comprising:

10. 2. The information processing device according to claim 1, Further comprising a time measurement unit, the storage unit further stores second usage status data that defines an order in which the voice recognition device calls the plurality of voice recognition devices according to a time period; the control unit acquires from the time measurement unit a time when the conversation activation hotword was detected, and transmits the control command to each of the plurality of voice recognition devices in accordance with a call priority order defined for a time period including the time in the second usage status data.

1. An information processing device comprising:

11. 3. The information processing device according to claim 2, Further comprising a time measurement unit, the storage unit further stores second usage status data that defines an order in which the voice recognition device calls the plurality of voice recognition devices according to a time period; the control unit acquires from the time measurement unit a time when the conversation activation hotword was detected, and transmits a control command to activate the conversation mode to each of the plurality of voice recognition devices in accordance with a call priority determined for a time period including the time in the second usage status data.

1. An information processing device comprising:

Citation Information

Patent Citations

  • Door phone system

    JP2008053989A

  • Security system for multiple dwelling building

    JP2008067310A

  • Intercom system including other person authentication function and other person judgment method

    JP2008306620A

  • Intercom apparatus, programs to be functioned as means in intercom apparatus, intercom system and information processing method of intercom apparatus

    JP2010098598A

  • Consolidating speech recognition results

    JP2013068952A