Voice control method
The voice control method coordinates voice outputs from multiple speakers based on their status and user ownership to deliver notifications at appropriate times, addressing the challenge of timely information delivery in multi-user environments.
Patent Information
- Application Number
- JP2024072276
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-25
- Filing Date
- 2024-04-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-07-15
AI Technical Summary
Existing voice notification systems in home appliances struggle to ensure that information is delivered at appropriate timings, especially when multiple devices owned by different users are involved, leading to potential confusion or difficulty in hearing for the users.
A voice control method that determines the timing for each speaker to output voice based on whether they are currently speaking, ownership, and user association, allowing for coordinated voice output to ensure appropriate timing and minimize overlap.
Ensures that voice notifications are delivered at optimal times, reducing confusion and improving user experience by avoiding simultaneous or overlapping voice outputs from multiple devices.
Smart Images

Figure 0007702684000001 
Figure 0007702684000002 
Figure 0007702684000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a voice control method and a server device.
Background Art
[0002] Conventionally, in electronic devices such as home appliances, there is a device that outputs (speaks) voice (see, for example, Patent Document 1).
[0003] Patent Document 1 discloses a server device that creates voice data for a given electronic device to speak based on characteristic information set based on at least one of the attribute information of a user of the electronic device and the attribute information of the electronic device.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] There is a system that notifies a user of information such as home appliances in voice, like the system including the server device disclosed in Patent Document 1. For this type of system, it is required that the information be easy for the user to hear. For this purpose, the speaker that notifies the user of the information in voice needs to notify the user of the information in voice at an appropriate timing.
[0006] The present disclosure provides a voice control method and the like in which each of a plurality of speakers can notify information in voice at an appropriate timing.
Means for Solving the Problems
[0007] A voice control method according to an aspect of the present disclosure is a voice control method executed by a server device, including: a determination step of determining whether each of a plurality of speakers capable of outputting voice is outputting voice; a timing determination step of determining, based on the determination result in the determination step, whether to immediately output voice to at least one of the plurality of speakers, or to wait until the speaker that is outputting voice finishes outputting the voice and then determine the timing to output voice to the at least one speaker; and an output step of outputting voice to the at least one speaker at the timing determined in the timing determination step. In the timing determination step, owner information indicating each owner of the plurality of speakers is acquired, and when a speaker owned by the same owner as the at least one speaker among the plurality of speakers is outputting voice, the timing to output voice to the at least one speaker after the speaker finishes outputting the voice is determined.
[0008] A voice control method according to an aspect of the present disclosure is a voice control method executed by a server device, including: a determination step of determining whether each of a plurality of speakers capable of outputting voice is outputting voice; a timing determination step of determining, based on the determination result in the determination step, whether to immediately output voice to at least one of the plurality of speakers, or to wait until the speaker that is outputting voice finishes outputting the voice and then determine the timing to output voice to the at least one speaker; and an output step of outputting voice to the at least one speaker at the timing determined in the timing determination step. In the timing determination step, owner information indicating each owner of the plurality of speakers is acquired, and when the owner of the at least one speaker is a first user and a second user, and a speaker owned by at least one of the first user and the second user among the plurality of speakers is outputting voice, the timing to output voice to the at least one speaker after the speaker finishes outputting the voice is determined.
[0009] A voice control method according to one aspect of the present disclosure is a voice control method executed by a server device, including a determination step of determining whether each of a plurality of speakers capable of outputting voice is outputting voice, and based on the determination result in the determination step, a timing determination step of determining a timing of causing at least one of the plurality of speakers to output voice immediately, or causing the at least one speaker to output voice after waiting until a speaker that is outputting voice finishes outputting the voice, and an output step of causing the at least one speaker to output voice at the timing determined in the timing determination step. In the timing determination step, owner information indicating each owner of the plurality of speakers is acquired, and when at least one of the speakers is owned by the first user among the first user and the second user, and among the plurality of speakers, at least one of the one or more speakers owned by the first user is owned by the second user, when a speaker owned by the second user is outputting voice, the timing of causing the at least one speaker to output voice after the speaker finishes outputting the voice is determined.
[0010] Further, a server device according to one aspect of the present disclosure includes a determination unit that determines whether each of a plurality of speakers capable of outputting voice is outputting voice, and an output unit that causes at least one of the plurality of speakers to output voice at a timing of causing at least one of the plurality of speakers to output voice immediately, or causing the at least one speaker to output voice after waiting until a speaker that is outputting voice finishes outputting the voice, based on the determination result of the determination unit and owner information indicating each owner of the plurality of speakers.
[0011] These general or specific aspects may be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. [Effect of the Invention]
[0012] According to the present disclosure, it is possible to provide a voice control method or the like in which each of a plurality of speakers can notify information by voice at an appropriate timing. [Brief Description of the Drawings]
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0014] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that each of the embodiments described below shows a specific example of the present disclosure. Therefore, numerical values, shapes, materials, components, arrangements and connection forms of the components, steps, and the order of steps shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Thus, among the components in the following embodiments, components not described in the independent claims indicating the highest concept of the present disclosure are described as arbitrary components.
[0015] Also, each figure is a schematic diagram and is not necessarily drawn precisely. Also, in each figure, the same reference numerals are given to the same constituent members.
[0016] (Embodiment) [Configuration] FIG. 1 is a schematic diagram showing a specific configuration of a voice utterance system 500 according to an embodiment.
[0017] The voice utterance system 500 is a device that, when information such as information indicating a change in the processing state, information notifying a failure, information prompting the user to replace parts such as a filter, and information notifying (recommended notification) the user of the functions of the device 600 is output in the device 600, notifies (outputs) the user of the information by voice (in other words, utters the information). For example, assume that the device 600 is a washing machine and the washing has ended. In this case, for example, the device 600 transmits operation information indicating that the washing has ended to the server device 100. When the server device 100 receives the operation information, it transmits a voice file (voice data) for outputting a voice utterance sentence such as "The washing is finished" to the uttering body 200, which is a device capable of outputting voice. The uttering body 200 has a device such as a speaker for outputting voice, and outputs (that is, utters) a voice utterance sentence such as "The washing is finished" based on the received voice file.
[0018] The voice conversation system 500 includes one or more devices 600, a server device 100, and one or more speakers 200.
[0019] The device 600 is, for example, an electrical appliance such as a refrigerator, a washing machine, a microwave oven, a lighting device, a doorbell, etc., and is a device (information source device) capable of outputting information of the device 600. More specifically, for example, the device 600 is a communicable electrical appliance (home appliance) in the user's home. The device 600 transmits, for example, identification information which is a unique identifier indicating the device 600, device information indicating the performance (specifications) etc. of the device 600, operation information indicating the content of processing (operation), and state information indicating the state of the device 600 such as a failure, etc. to the server device 100. Note that the operation information may include device information indicating the device 600 that executed the operation content indicated by the operation information.
[0020] Also, the device 600 transmits, for example, information indicating the user of the device 600 to the server device 100. The information indicating the user is received from the user, for example, via a reception unit that receives an input from the user such as a touch panel (not shown) that the device 600 has.
[0021] Note that the device 600 is a device different from a portable terminal such as a smartphone, for example. Specifically, the device 600 is a device that can be used by a plurality of users (for example, assumed to be used by a plurality of users), different from a portable terminal, for example.
[0022] The user of a portable terminal such as a smartphone is specified. Therefore, when notifying the user of information by the portable terminal, even if the portable terminal is notifying the user of another piece of information, the user of the portable terminal is only the user who is the target of the notification, that is, since it is assumed that the user possesses the portable terminal, the portable terminal may simply perform the plurality of notifications in order when making a plurality of notifications to the user.
[0023] On the one hand, home appliances are not necessarily in the possession of the user who is the recipient of the notification, such as being shared among family members, and may be in the possession of other users. Therefore, in order to notify a specific user of information regarding such a home appliance, device 600, there are issues such as the need to hold off on the notification when it is in the possession of someone other than the user.
[0024] Therefore, in the voice conversation system 500, in order to be able to appropriately notify the user of device 600 of information regarding device 600, for example, device 600 transmits information indicating the user of device 600 to the server device 100 together with the device information and operation information of device 600 and the like.
[0025] Device 600 is realized, for example, by a communication interface for communicating with the server device 100, an execution unit that executes processes such as refrigeration, washing, and heating, a detection unit realized by a sensor or the like for detecting the state of device 600, and a control unit realized by a processor, a memory, and the like that control various processes of device 600.
[0026] Based on the information received from device 600, the server device 100 determines the utterance sentence (scenario) to be output to the speaker 200 and outputs the created utterance sentence to the speaker 200 as voice. For example, when the server device 100 receives operation information from device 600, it selects a voice file (voice data) corresponding to the operation information and transmits the selected voice file to the speaker 200 as notification information (also referred to as voice information), so that the speaker 200 outputs voice based on the voice file.
[0027] Examples of the utterance sentence include a sentence indicating that device 600 has started operating, a sentence indicating that device 600 has finished operating, a sentence indicating that it has operated in conjunction with another device 600, a sentence prompting the user for a version update, a sentence recommending to the user the use of the functions that device 600 has, a sentence indicating that it has malfunctioned, and the like.
[0028] The server device 100 is realized by a computer including, for example, a communication interface for communicating with devices such as the device 600 and the speaker 200, a non-volatile memory storing a program, a volatile memory which is a temporary storage area for executing the program, an input / output port for transmitting and receiving signals, a processor for executing the program, and the like.
[0029] The speaker 200 is an electrical appliance such as an air conditioner, a television, or a self-propelled vacuum cleaner (so-called robot vacuum cleaner), and is a device (home appliance equipped with a speaker) having a component capable of outputting sound such as a speaker.
[0030] For example, when the speaker 200 receives voice information such as a voice file from the server device 100, it outputs a voice based on the received voice information.
[0031] In FIG. 1, three devices 600 are illustrated, but the number of devices 600 included in the voice speaking system 500 may be one or a plurality, and is not particularly limited.
[0032] Also, in FIG. 1, three speakers 200 are illustrated, but the number of speakers 200 included in the voice speaking system 500 may be one or a plurality, and is not particularly limited.
[0033] The server device 100 is communicably connected to each of the three devices 600 and the three speakers 200 via a network such as the Internet.
[0034] The server device 100 and each of the three devices 600 and the three speakers 200 may be communicably connected via a LAN (Local Area Network) or the like, or may be communicably connected via wireless communication.
[0035] In addition, the communication standards used for communication between the server device 100 and each of the three devices 600 and the three speakers 200 are not particularly limited. Examples of the communication standards include Wi-Fi (registered trademark), Bluetooth (registered trademark), or ZigBee (registered trademark).
[0036] Each of the three devices 600 and the three speakers 200 is arranged, for example, inside the house where the user lives. Also, the server device 100 is arranged, for example, outside the house.
[0037] FIG. 2 is a block diagram showing the server device 100 according to the embodiment. In FIG. 2, only one device 600 is shown representatively. Also, in FIG. 2, three speakers 200 are shown. To distinguish the three speakers 200, they are labeled as speaker 201, speaker 202, and speaker 203.
[0038] The server device 100 includes an acquisition unit 110, a scenario determination unit 120, a speaker determination unit 130, a determination unit 140, a timing determination unit 150, an output unit 160, and a storage unit 170.
[0039] The acquisition unit 110 is a processing unit that acquires information related to the device 600, such as device information including the performance, type, model number, etc. of the device 600, and operation information indicating the operation history (contents of operations) of the device 600. The acquisition unit 110 acquires the device information and / or the operation information, for example, by communicating with the device 600 via a communication unit such as a communication interface (not shown) provided in the server device 100. The communication unit is, for example, a communication interface for communicating with the device 600 and the speaker 200. The communication unit is realized, for example, by a connector or the like to which a communication line is connected when communicating with the speaker 200 and the device 600 by wire, and is realized by an antenna and a wireless communication circuit or the like when communicating wirelessly.
[0040] Note that when the server device 100 includes a reception device such as a mouse or a keyboard for receiving an input from the user, the device information and / or the operation information may be acquired via the reception device.
[0041] The acquisition unit 110 stores the acquired device information and operation information in the storage unit 170 or outputs it to the scenario determination unit 120.
[0042] The scenario determination unit 120 is a processing unit that determines whether the operation information acquired by the acquisition unit 110 satisfies a predetermined condition and determines a speech sentence to be spoken to the speaker 200. Specifically, the scenario determination unit 120 determines whether an event that causes the speaker 200 to output voice has occurred based on the operation information acquired by the acquisition unit 110. For example, in the storage unit 170, the operation content corresponding to the type of device 600 that determines that an event has occurred (that is, a predetermined condition is satisfied) is stored. For example, the scenario determination unit 120 determines whether an event that causes the speaker 200 to output voice has occurred by determining whether the operation content indicated by the operation information acquired by the acquisition unit 110 matches the operation content corresponding to the type of device 600 that determines that an event has occurred and is stored in the storage unit 170.
[0043] Examples of the predetermined conditions include that the device 600 has started operating, the device 600 has ended operating, has operated in association with another device 600, can be upgraded, has failed, etc.
[0044] Note that the predetermined condition may be arbitrarily determined in advance.
[0045] For example, when the scenario determination unit 120 determines that the operation content indicated by the operation information acquired by the acquisition unit 110 satisfies a predetermined condition, the scenario determination unit 120 determines a speech sentence corresponding to the operation information. For example, in the storage unit 170, speech sentences associated with the operation content are stored, and the scenario determination unit 120 determines the speech sentence to be output as voice to the speaker 200 by selecting the speech sentence associated with the operation content indicated by the operation information.
[0046] The speaker determination unit 130 is a processing unit that determines which of the plurality of speakers 200 will output the speech sentence determined by the scenario determination unit 120 as voice. For example, in the storage unit 170, device information indicating the device 600 and speaker information indicating the speaker 200 are pre-associated and stored. For example, when the acquisition unit 110 acquires the operation information of the first device which is an example of the device 600, and the device information of the first device is associated with the speaker information of the speakers 201 and 202, the speakers 201 and 202 output the speech sentence corresponding to the operation information as voice. Also, for example, when the acquisition unit 110 acquires the operation information of the second device which is another example of the device 600, and the device information of the second device is associated with the speaker information of the speaker 201, the speaker 201 outputs the speech sentence corresponding to the operation information as voice.
[0047] Also, for example, in the storage unit 170, owner information indicating the owner of the device 600 and the speaker 200 is stored in association with the device information and the speaker information. In this case, for example, assuming that the acquisition unit 110 has acquired the operation information of the device 600, the speaker determination unit 130 determines the speaker 200 so that the speaker 200 having the same owner as the device 600 outputs the speech sentence corresponding to the operation information as voice. In this way, for example, the speaker determination unit 130 determines which of the plurality of speakers 200 included in the voice speech system 500 will output the speech sentence determined by the scenario determination unit 120 as voice, based on the device information, the speaker information, and the owner information.
[0048] Note that the owner information may be pre-stored in the storage unit 170. Alternatively, for example, the acquisition unit 110 may acquire the owner information received from the user by a reception device such as a smartphone (not shown) via the communication unit (not shown) described above, and store the acquired owner information in the storage unit 170.
[0049] The determination unit 140 is a processing unit that determines whether each of the plurality of speakers 200 is outputting voice. For example, the determination unit 140 determines whether each of the speakers 201, 202, and 203 is outputting voice.
[0050] Note that whether or not the plurality of speakers 200 are outputting voice here indicates, for example, whether or not the server device 100 is outputting a speech sentence to the speaker 200 as voice. For example, some speakers 200 may output voice to notify information of their own devices, or when the speaker 200 is a television, it may output voice in accordance with the video. Thus, the voice output by the speaker 200 determined by the determination unit 140 may or may not include voice other than the voice (voice based on the speech sentence) output by the server device 100 to the speaker 200.
[0051] For example, the determination unit 140 determines whether or not each of the speaker 201, the speaker 202, and the speaker 203 is outputting as voice the speech sentence determined by the scenario determination unit 120. For example, the determination unit 140 determines whether or not each of the speaker 201, the speaker 202, and the speaker 203 is outputting voice based on the timing determined by the timing determination unit 150 described later and the length of the speech sentence determined by the scenario determination unit 120. The output time of the voice corresponding to the length of the speech sentence may be stored in the storage unit 170 in advance, or information indicating the time required to output one sound or the like is stored in the storage unit 170 in advance, and the time required to output the speech sentence as voice may be calculated from the information and the speech sentence. Alternatively, the determination unit 140 may communicate with each of the speaker 201, the speaker 202, and the speaker 203 via the communication unit (not shown) provided in the server device 100 described above to obtain information (voice output information) indicating whether or not each of the speaker 201, the speaker 202, and the speaker 203 is speaking.
[0052] The timing determination unit 150 is a processing unit that determines, based on the determination result of the determination unit 140, the timing of causing at least one of the plurality of speakers 200 to output voice immediately, or causing at least one of the speakers 200 that is outputting voice to output voice after waiting until the output of the voice ends.
[0053] For example, when the timing determination unit 150 determines that the speaker determination unit 130 causes a plurality of speakers 200 to output a speech sentence (more specifically, the same speech sentence) as voice, for a first speaker among the plurality of speakers 200 that is not outputting voice, the timing for immediately outputting voice to the first speaker is determined, and for a second speaker among the plurality of speakers 200 that is outputting voice, the timing for outputting voice to the second speaker is determined after waiting until the output of the voice ends.
[0054] Alternatively, for example, when the timing determination unit 150 determines that the speaker determination unit 130 causes a plurality of speakers 200 to output a speech sentence (more specifically, the same speech sentence) as voice, when at least one of the plurality of speakers 200 is outputting voice, the timing for outputting voice to the at least one speaker 200 is determined after the at least one speaker 200 ends the output of the voice.
[0055] Alternatively, for example, the timing determination unit 150 acquires owner information indicating the owner of each of the plurality of speakers 200, and when a speaker 200 owned by the same owner as at least one speaker 200 for which voice is to be output among the plurality of speakers 200 is outputting voice, the timing for outputting voice to the at least one speaker 200 is determined after the speaker 200 ends the output of the voice.
[0056] In this case, for example, when at least one of the plurality of speakers 200 that is owned by the user who is the target of the utterance sentence to be output as voice among the plurality of speakers 200 is outputting voice, the timing determination unit 150 determines the timing to output voice to the at least one speaker 200 after the speaker that is outputting voice has finished outputting voice. For example, when the server device 100 acquires operation information from the device 600, in order to notify the user who is the owner of the device 600 of the utterance sentence based on the operation information, the server device 100 outputs the utterance sentence as voice to the speaker 200 that is owned by the user who is the target (notification target) of the utterance sentence, that is, the speaker 200 that has the same owner as the owner of the device 600. For example, in such a case, the timing determination unit 150 determines the timing to output voice to the at least one speaker 200 (for example, the speaker 201) based on whether or not a speaker 200 (for example, the speaker 202) that has the same owner as the at least one speaker 200 (for example, the speaker 201) that outputs the utterance sentence as voice is outputting voice.
[0057] Alternatively, for example, the timing determination unit 150 acquires owner information indicating the owner of each of the plurality of speakers 200, and when the owner of at least one speaker 200 that outputs voice is the first user and the second user, when a speaker 200 that is owned by at least one of the first user and the second user among the plurality of speakers 200 is outputting voice, the timing determination unit 150 determines the timing to output voice to at least one speaker 200 that is owned by at least one of them after the speaker 200 that is owned by at least one of them has finished outputting voice.
[0058] Alternatively, for example, the timing determination unit 150 acquires owner information indicating the owner of each of the plurality of speakers 200, and at least one speaker 200 that outputs voice is owned by the first user among the first user and the second user. Among the plurality of speakers 200, when at least one of the one or more speakers 200 owned by the first user is owned by the second user, when the speaker 200 owned by the second user is outputting voice, after the speaker 200 owned by the second user finishes outputting voice, the timing for outputting voice to at least one speaker 200 that outputs voice is determined.
[0059] Note that the timing determination unit 150 may output, as timing information together with the voice information, information indicating to output voice immediately, or information indicating an instruction to wait until the speaker 200 finishes outputting voice and then output voice, to the output unit 160 described later. Alternatively, for example, the timing determination unit 150 may output, as timing information together with the voice information, information indicating the time to output voice, or information indicating the time from receiving the voice information until outputting voice, to the output unit 160.
[0060] A specific example of the processing method for the timing determination unit 150 to determine the timing for outputting the uttered sentence as voice to the speaker 200 will be described later.
[0061] The output unit 160 is a processing unit that controls the output of the voice of the speaker 200. Specifically, based on the determination result of the determination unit 140, the output unit 160 causes at least one of the plurality of speakers 200 to output voice immediately, or waits until the speaker 200 that is outputting voice finishes the output of the voice, and then causes the at least one speaker 200 to output voice at a timing when the voice is output. More specifically, the output unit 160 causes the speech sentence determined by the scenario determination unit 120 to be output as voice at the timing determined by the timing determination unit 150 to at least one speaker 200 determined by the speaker determination unit 130. For example, the output unit 160 transmits voice information, which is information for outputting the speech sentence as voice to one or more speakers 200, and timing information indicating the timing determined by the timing determination unit 150, to one or more speakers 200 determined by the speaker determination unit 130 via the above-described communication unit (not shown) provided in the server device 100.
[0062] The voice information is information for outputting a speech sentence corresponding to the operation information of the device 600 as voice to the speaker 200. For example, the voice information is a voice file (voice data) corresponding to the operation information of the device 600. The voice file is stored in the storage unit 170, for example, in association with the operation content.
[0063] For example, the output unit 160 acquires a voice file corresponding to the speech sentence determined by the scenario determination unit 120 based on the operation information acquired by the acquisition unit 110 from the storage unit 170, and outputs (transmits) the acquired voice file as voice information to the speaker 200.
[0064] Thereby, when the speech sentence set (selected) by the user satisfies a predetermined condition (for example, the device 600 has executed a predetermined operation, has entered a predetermined state, etc.), the speech sentence is output as voice from one or more speakers 200 determined by the speaker determination unit 130 at the timing determined by the timing determination unit 150.
[0065] Note that the server device 100 may receive voice information from a computer such as another server device different from the server device 100. For example, the storage unit 170 may store information indicating a URL (Uniform Resource Locator) corresponding to a voice file. For example, after determining the utterance sentence, the scenario determination unit 120 may obtain the voice information by transmitting information indicating a URL corresponding to the voice information corresponding to the determined utterance sentence to the other server device.
[0066] Each processing unit of the acquisition unit 110, the scenario determination unit 120, the speaker determination unit 130, the determination unit 140, the timing determination unit 150, and the output unit 160 is realized from a memory, a control program stored in the memory, and a processor such as a CPU (Central Processing Unit) that executes the control program. Further, these processing units may be realized from one memory and one processor, or may be realized by a plurality of memories and a plurality of processors that are different from each other or in an arbitrary combination. Further, these processing units may be realized by, for example, a dedicated electronic circuit or the like.
[0067] The storage unit 170 is a storage device that stores device information indicating the device 600, speaker information indicating the speaker 200, owner information indicating the owner of the device 600 and the speaker 200, and information indicating a plurality of utterance sentences (scenario information). Further, the storage unit 170 may store a voice file corresponding to the utterance sentence.
[0068] The storage unit 170 is realized by, for example, an HDD (Hard Disk Drive) or a flash memory or the like.
[0069] Note that, for example, the storage unit 170 may store setting information indicating a speech sentence to be output in voice. The setting information is information indicating a speech sentence (more specifically, information indicating a speech sentence) among one or more speech sentences stored in the storage unit 170 that is set to be output in voice by the user. Depending on the user, there may be information that the user wants to be notified of in voice and information that does not need to be notified in voice. Therefore, for example, the acquisition unit 110 acquires, as setting information, information indicating whether to output in voice a speech sentence received from the user by a reception device such as a smartphone (not shown) via the communication unit (not shown) described above, and stores the acquired setting information in the storage unit 170. For example, when the acquisition unit 110 acquires operation information, the scenario determination unit 120 may determine whether to output in voice a speech sentence related to the operation information to the speech body 200 based on the setting information stored in the storage unit 170. The setting information may be set for each user.
[0070] As described above, the speech body 200 is, for example, an electrical appliance such as an air conditioner, a television, or a self-propelled vacuum cleaner, and is a device provided with a component capable of outputting voice such as a speaker. The speech body 200 outputs voice based on voice information such as a voice file received from the server device 100.
[0071] Note that the speech sentence and the voice file corresponding to the speech sentence are stored in a storage unit (not shown) such as an HDD, and the speech body 200 may be provided with the storage unit. In this case, for example, the output unit 160 may transmit, as voice information, information indicating a speech sentence to be output in voice to the speech body 200, or information indicating a voice file associated with the speech sentence. In this case, for example, the speech body 200 selects, from among one or more voice files stored in the storage unit, a voice file for outputting voice based on the received voice information, and outputs voice based on the selected voice file.
[0072] The voice output device 200 includes, for example, a memory storing a control program for causing a speaker to output voice based on voice information received from the server device 100, a processor that executes the control program, and a communication interface for communicating with the server device 100. The communication interface is realized, for example, by a connector or the like to which a communication line is connected when the voice output device 200 performs wired communication with the server device 100, and is realized by an antenna and a wireless communication circuit or the like when performing wireless communication.
[0073] The voice output device 200 includes, for example, a communication unit 210, a voice control unit 220, and a voice output unit 230.
[0074] The communication unit 210 is a communication interface for communicating with the server device 100.
[0075] The voice control unit 220 is a processing unit that causes the voice output unit 230 to output voice based on voice information received (acquired) from the server device 100 (more specifically, the output unit 160) via the communication unit 210. Specifically, the voice control unit 220 transmits voice output information indicating whether or not the voice output unit 230 is outputting voice to the server device 100 via the communication unit 210, and receives voice information and timing information indicating the timing at which to output voice from the server device 100 via the communication unit 210, and causes the voice output unit 230 to output voice based on the received voice information at the timing based on the received timing information.
[0076] The voice control unit 220 is realized by a memory, a control program stored in the memory, and a processor such as a CPU that executes the control program. Further, the voice control unit 220 may be realized by, for example, a dedicated electronic circuit or the like.
[0077] The voice output unit 230 is a device that outputs voice under the control of the voice control unit 220. The voice output unit 230 is realized by, for example, a speaker or the like.
[0078] [Specific Example] Next, a specific example of a processing method for the timing determination unit 150 to determine the timing for the speaker 200 to output a spoken sentence in voice will be described. In the first to fifth examples described below, the speaker 201 and the speaker 202 will be described assuming that the user A is the owner. Also, in the first to fifth examples described below, the speaker 202 and the speaker 203 will be described assuming that the user B is the owner. That is, the speaker 202 is shared by the user A and the user B. Further, in the first to fifth examples described below, a case where information is output in voice to the user B is shown.
[0079] <Example 1> FIG. 3 is a diagram for explaining a first example of a processing method for the server device 100 according to the embodiment to determine the timing for the speaker 200 to output a spoken sentence in voice.
[0080] In this example, it is assumed that a spoken sentence will be output in voice to the speaker 202 and the speaker 203, and the speaker 202 is outputting voice. That is, in this example, the speaker 202 and the speaker 203 are candidates for speaking, and the speaker 202 is speaking.
[0081] In this case, the timing determination unit 150 determines the timing so that the speaker 202 that is speaking outputs voice after waiting until the speaking ends. On the other hand, the timing determination unit 150 determines the timing so that the speaker 203 that is not speaking immediately speaks the spoken sentence. Therefore, in this example, the speaker 202 and the speaker 203 that speak the same spoken sentence speak the spoken sentence at different timings.
[0082] In this way, in the first example, the timing determination unit 150 determines the timing so that, among two or more speakers 200, for the first speaker that is not outputting voice, voice is immediately output to the first speaker, and for the second speaker that is outputting voice among the two or more speakers 200, the timing is determined so that voice is output to the second speaker after waiting until the output of the voice ends.
[0083] Note that the speaker 200 that is a candidate for speech can be either user A or user B as the owner, and the owner is not particularly limited. For example, when outputting information for user B in voice, the speaker 200 may be at least one of the speakers 202 and 203 owned by user B.
[0084] <Second Example> FIG. 4 is a diagram for explaining a second example of a processing method for determining the timing at which the server device 100 according to the embodiment outputs a speech sentence to the speaker 200 in voice.
[0085] In this example, it is assumed that a speech sentence is to be output in voice to the speakers 202 and 203, and the speaker 202 is outputting voice. That is, in this example, the speakers 202 and 203 are candidates for speech, and the speaker 202 is in the middle of speaking.
[0086] In this case, the timing determination unit 150 determines the timing so that the speaking speaker 202 speaks after waiting until the speech ends. Also, for the speaker 203 that is not speaking, the timing determination unit 150 determines the timing so that it speaks after waiting until the speech of the speaker 202 ends. Therefore, in this example, the speakers 202 and 203 that speak the same speech sentence speak the speech sentence at the same timing.
[0087] As described above, in the second example, when at least one of the two or more speakers 200 that are all candidates for speech outputs voice, the timing determination unit 150 determines the timing so that the two or more speakers 200 output voice after the at least one of the speakers 200 finishes outputting voice (for example, so that the timings at which the same speech sentence is output in voice are the same).
[0088] <Third Example> FIG. 5 is a diagram for explaining a third example of a processing method for determining the timing at which the server device 100 according to the embodiment outputs a speech sentence to the speaker 200 in voice.
[0089] In this example, it is assumed that the utterance body 203 will output an utterance sentence in voice, and the utterance body 202 is outputting voice. That is, in this example, the utterance body 203 is a candidate for utterance, and the utterance body 202 is in the middle of uttering.
[0090] In this example, the timing determination unit 150 identifies the utterance body 200 whose owner is the same user B as that of the utterance body 203 by acquiring the respective owner information of the utterance bodies 201, 202, and 203. In this example, the timing determination unit 150 identifies the utterance body 202 whose owner is the same user B as that of the utterance body 203. Also, for example, when the utterance body 202 whose owner is the same as that of the candidate utterance body 203 is uttering, the timing determination unit 150 determines the timing so that the utterance body 203 starts uttering after the utterance body 202 finishes uttering. On the other hand, for example, even if the utterance body 202 whose owner is the same as that of the candidate utterance body 203 is not uttering and the utterance body 201 whose owner is different from that of the candidate utterance body 203 is uttering, the timing determination unit 150 determines the timing so that the utterance body 203 starts uttering immediately.
[0091] As described above, in the third example, the timing determination unit 150 acquires the owner information indicating the respective owners of the plurality of utterance bodies 200, and when an utterance body 200 whose owner is the same as that of at least one utterance body 200 that outputs voice among the plurality of utterance bodies 200 is outputting voice, the timing determination unit 150 determines the timing so that the at least one utterance body 200 starts outputting voice after the utterance body 200 finishes outputting voice.
[0092] Note that, for example, the determination unit 140 may acquire the respective owner information of the utterance bodies 201, 202, and 203, and determine whether each of the utterance body 203 and the utterance body 202 whose owner is the same user B as that of the utterance body 203 is in the middle of uttering, or may determine whether each of the utterance bodies 201, 202, and 203, which are all the utterance bodies included in the voice utterance system 500, is in the middle of uttering.
[0093] <Fourth Example> FIG. 6 is a diagram for explaining a fourth example of a processing method for determining the timing at which the server device 100 according to the embodiment causes the speaker 200 to output a speech sentence as voice.
[0094] In this example, it is assumed that the speaker 202 is about to output a speech sentence as voice and the speaker 201 is outputting voice. That is, in this example, the speaker 202 is a speech candidate and the speaker 201 is in the middle of speaking.
[0095] In this example, the timing determination unit 150 acquires the respective owner information of the speakers 201, 202, and 203, and thereby identifies the speaker 200 in which at least one of the speakers 202 and the owner is the same user A and user B. In this example, the timing determination unit 150 identifies the speaker 201 in which the speaker 202 and the owner are the same user A, and the speaker 203 in which the speaker 202 and the owner are the same user B. Further, for example, when at least one of the speakers 201 and 203 in which at least one of the speakers 202 and the owner of the speech candidate is the same is speaking, the timing determination unit 150 determines the timing so that the speaker 202 is caused to speak after both the speakers 201 and 203 have finished speaking. In this example, since the speaker 201 in which at least one of the speakers 202 and the owner of the speech candidate is the same is speaking, the timing determination unit 150 determines the timing so that the speaker 202 is caused to speak after the speaker 201 has finished speaking. Therefore, in this example, for example, when the speaker 201 in which at least one of the speakers 202 and the owner of the speech candidate is the same is not speaking and the speaker 203 in which at least one of the speakers 202 and the owner of the speech candidate is the same is speaking, the timing determination unit 150 determines the timing so that the speaker 202 is caused to speak after the speaker 203 has finished speaking.
[0096] As described above, in the fourth example, the timing determination unit 150 acquires owner information indicating the owner of each of the plurality of speakers 200, and when the owners of at least one speaker 200 that outputs voice are the first user and the second user, among the plurality of speakers 200, when the speaker 200 of which at least one of the first user and the second user is the owner is outputting voice, after the speaker 200 of which at least one of them is the owner finishes outputting voice, the timing is determined so that voice is output to at least one speaker 200 of which at least one of them is the owner.
[0097] <Example 5> FIG. 7 is a diagram for explaining a fifth example of a method for determining the timing at which the server device 100 according to the embodiment causes the speaker 200 to output a speech sentence as voice.
[0098] In this example, it is assumed that a speech sentence is to be output as voice from the speaker 203 and the speaker 201 is outputting voice. That is, in this example, the speaker 203 is a speech candidate and the speaker 201 is in the middle of speaking.
[0099] In this example, the timing determination unit 150 obtains the owner information of each of the speakers 201, 202, and 203, and determines whether there is an owner other than user B for the speakers 202 and 203 that have the same owner as the speaker 203, which is owned by user B. In this example, since the speaker 202 owned by user B is also owned by user A, it is determined that there is an owner other than user B for the speakers 202 and 203 owned by user B. Further, when the timing determination unit 150 determines that there is an owner other than user B for the speakers 202 and 203 owned by user B, the timing determination unit 150 identifies the speaker 200 owned by the owner other than user B. In this example, the timing determination unit 150 identifies the speaker 201 owned by user A, who is an owner other than user B, for the speakers 202 and 203 owned by user B. Also, for example, when the identified speaker 200 is speaking, the timing determination unit 150 determines the timing so that the speaker 203 starts speaking after the identified speaker 200 finishes speaking. In this example, since the identified speaker 201 is speaking, the timing determination unit 150 determines the timing so that the speaker 203 starts speaking after the identified speaker 201 finishes speaking.
[0100] As described above, in the fifth example, the timing determination unit 150 obtains the owner information indicating the owner of each of the plurality of speakers 200, and when at least one speaker 200 that outputs voice is owned by the first user (for example, user B) and the second user (for example, user A), and the second user owns at least one of the one or more speakers 200 owned by the first user among the plurality of speakers 200, when the speaker 200 owned by the second user is outputting voice, the timing determination unit 150 determines the timing so that the at least one speaker 200 that outputs voice starts outputting voice after the speaker 200 owned by the second user finishes outputting voice.
[0101] Note that the above-described first example, second example, third example, fourth example, and fifth example may be arbitrarily combined and implemented within a possible range.
[0102] For example, in the above-described fifth example, when causing a voice to be output from one speech body 200 owned by the first user, it may be determined whether or not another speech body 200 owned by the first user is in the middle of a speech. For example, when the other speech body 200 is in the middle of a speech, wait until the other speech body 200 finishes outputting the voice, and then cause the one speech body 200 to output the voice. Here, when the owner of the one speech body 200 includes not only the first user but also the second user, when the other speech body 200 owned by the first user is not in the middle of a speech, it may further be determined whether or not a speech body 200 owned by the second user is in the middle of a speech. In this case, for example, when the other speech body 200 owned by the first user is not in the middle of a speech and the speech body 200 owned by the second user is not in the middle of a speech, cause the one speech body 200 to output the voice. On the other hand, when the speech body 200 owned by the second user is in the middle of a speech, wait until the speech body 200 owned by the second user finishes outputting the voice, and then cause the one speech body 200 to output the voice.
[0103] [Processing Procedure] Subsequently, the processing procedure of the processing executed by the server device 100 will be described.
[0104] FIG. 8 is a flowchart showing the processing procedure of the server device 100 according to the embodiment.
[0105] First, the scenario determination unit 120 determines whether or not the acquisition unit 110 has acquired the operation information of the device 600 from the device 600 (S101).
[0106] When the scenario determination unit 120 determines that the acquisition unit 110 has not acquired the operation information (No in S101), the process returns to step S101.
[0107] On the other hand, when the scenario determination unit 120 determines that the acquisition unit 110 has acquired the operation information (Yes in S101), based on the operation information, a speech sentence is determined (S102).
[0108] Next, the speaker determination unit 130 determines at least one speaker 200 that outputs the utterance sentence determined by the scenario determination unit 120 in voice, based on, for example, device information indicating the device 600 that has executed the operation indicated by the operation information (S103).
[0109] Next, the determination unit 140 determines whether or not a plurality of speakers 200 provided in the voice utterance system 500 (more specifically, the speakers 200 in which speaker information indicating the speakers 200 is stored in the storage unit 170) are outputting voice (S104).
[0110] Next, based on the determination result of the determination unit 140, the timing determination unit 150 determines the timing of causing at least one of the plurality of speakers 200 to output voice immediately, or causing at least one of the plurality of speakers 200 that are outputting voice to output voice after waiting until the output of the voice ends (S105). The timing determination unit 150 determines the timing of causing at least one speaker 200 determined by the speaker determination unit 130 to output voice, using, for example, any of the determination methods of the first to fifth examples described above.
[0111] Next, the output unit 160 causes at least one speaker 200 determined by the speaker determination unit 130 to output the utterance sentence determined by the scenario determination unit 120 in voice at the timing determined by the timing determination unit 150 (S106).
[0112] Note that the information handled in step S101 may be any information as long as it is information for notifying the user, such as not only the operation information of the device 600 but also information indicating an upgrade of the device 600, information indicating that a failure has occurred, etc. Regarding the processing after step S102 as well, based on information for notifying the user, such as information indicating an upgrade of the device 600, information indicating that a failure has occurred, etc., an utterance sentence may be determined and the utterance sentence may be output in voice from the speaker 200.
[0113] Subsequently, the processing procedure of the processing executed by the speaker 200 will be described.
[0114] FIG. 9 is a flowchart showing the processing procedure of the speaker 200 according to the embodiment.
[0115] First, the voice control unit 220 transmits voice output information indicating whether or not the voice is being output from the voice output unit 230 to the server device 100 via the communication unit 210 (S201). The timing at which the voice control unit 220 executes step S201 is not particularly limited. The voice control unit 220 may repeatedly execute step S201 at a predetermined period arbitrarily determined in advance, or may execute step S201 when receiving information requesting voice output information from the server device 100.
[0116] Note that the voice control unit 220 may transmit information indicating that the utterance has ended (that is, that the voice output from the voice output unit 230 has ended) to the server device 100 via the communication unit 210 as the voice output information.
[0117] According to this, since the server device 100 can also grasp that the utterance has been started for the speaker 200, if it is known when the utterance has ended, the server device 100 can appropriately determine whether or not each speaker 200 is in the middle of an utterance.
[0118] In addition, the server device 100 may determine that the utterance of the speaker 200 has ended when voice output information indicating that the utterance has ended is not received for a predetermined time.
[0119] The server device 100 executes step S104 shown in FIG. 8 based on the received voice output information, for example, and further transmits voice information such as a voice file and timing information.
[0120] Next, the voice control unit 220 receives voice information and timing information indicating the timing at which the voice is to be output from the server device 100 via the communication unit 210 (S202).
[0121] Next, the voice control unit 220 causes the voice output unit 230 to output a voice based on the voice information at a timing based on the timing information received in step S202 (S203).
[0122] [Effects, etc.] As described above, the voice control method according to the embodiment includes a determination step (S104) of determining whether or not a plurality of speakers 200 capable of outputting voices are outputting voices, and based on the determination result in the determination step, among the plurality of speakers 200, at least one speaker 200 is caused to output a voice immediately, or the speaker 200 that is outputting a voice waits until the output of the voice ends and then outputs a voice to the at least one speaker 200 at a timing to output a voice to the at least one speaker 200, and an output step (S106).
[0123] According to this, for example, by outputting voices from a plurality of speakers 200 simultaneously, it is possible to avoid a timing at which it is difficult for the user to hear the voices and output voices from the speakers 200. Thus, according to the voice control method according to the embodiment, the speaker 200 can notify information by voice at an appropriate timing.
[0124] Further, for example, the voice control method according to the embodiment further includes a timing determination step (S105) of determining, based on the determination result in the determination step, whether to immediately output a voice to at least one of the plurality of speakers 200 or to wait until the speaker 200 that is outputting a voice finishes outputting the voice and then output a voice to the at least one speaker 200 at a timing to output a voice to the at least one speaker 200. In this case, for example, in the output step, a voice is output to the at least one speaker 200 at the timing determined in the timing determination step.
[0125] As a result, in the output step, based on the determination result in the determination step, at the timing of causing at least one of the plurality of speakers 200 to output voice immediately, or causing the at least one speaker 200 to output voice after waiting until the speaker 200 that is outputting voice finishes the output of the voice, the at least one speaker 200 can be caused to output voice.
[0126] Also, for example, in the timing determination step, for a first speaker among the plurality of speakers 200 that is not outputting voice, the timing for causing the first speaker to output voice immediately is determined, and for a second speaker among the plurality of speakers 200 that is outputting voice, the timing for causing the second speaker to output voice after waiting until the output of the voice ends is determined.
[0127] According to this, when causing a spoken sentence to be output as voice, whether the speaker 200 outputs voice is determined based on whether voice is currently being output. Therefore, the process of timing determination is simplified.
[0128] Also, for example, in the timing determination step, when at least any one of the plurality of speakers 200 is outputting voice, the timing for causing the at least one speaker 200 to output voice after at least any one of the speakers 200 finishes the output of the voice is determined.
[0129] According to this, the user can hear the same information at the same timing. Therefore, it is suppressed that the user has a misunderstanding or feels uncomfortable due to hearing the same information at the same timing.
[0130] Also, for example, in the timing determination step, ownership information indicating the owner of each of the plurality of speakers 200 is acquired. When, among the plurality of speakers 200, a speaker 200 owned by the same owner as at least one speaker 200 for which voice output is to be performed is outputting voice, the timing for outputting voice to the at least one speaker 200 is determined after the speaker 200 that is outputting the voice has finished outputting the voice.
[0131] Among the plurality of speakers 200, it is highly likely that information for the user is output in voice from the speakers 200 owned by the same user. Therefore, if different utterance sentences are output in voice at the same timing from each of the plurality of speakers 200 owned by the same user, the user has to listen to a plurality of pieces of information simultaneously, and there is a possibility that the information cannot be heard correctly. Thus, when, among the plurality of speakers 200, a speaker 200 owned by the same owner as at least one speaker 200 for which voice output is to be performed is outputting voice, by determining the timing so that voice is output to the at least one speaker 200 after the speaker 200 that is outputting the voice has finished outputting the voice, it is possible to suppress notifying the same user of different pieces of information at the same timing.
[0132] Also, for example, in the timing determination step, when, among the plurality of speakers 200, a speaker 200 owned by the same owner as at least one speaker 200 that is the target of the utterance sentence to be output in voice is outputting voice, the timing for outputting voice to the at least one speaker 200 is determined after the speaker 200 that is outputting the voice has finished outputting the voice.
[0133] According to this, it is further possible to suppress notifying the same user of different pieces of information at the same timing.
[0134] Also, for example, in the timing determination step, owner information indicating each owner of the plurality of speakers 200 is acquired. When at least one speaker 200 that outputs voice has owners who are the first user and the second user, among the plurality of speakers 200, when a speaker 200 of which at least one of the first user and the second user is the owner is outputting voice, the timing for outputting voice to the at least one speaker 200 is determined after the speaker 200 that is outputting voice finishes outputting the voice.
[0135] For example, as shown in FIG. 6, when the speaker 201 whose owner is user A is outputting voice, if voice is further output from the speaker 202 also owned by user A, user A may have a concern that the voice will be difficult to hear even if the information of the voice output from the speaker 202 is information for user B. Therefore, when a speaker 200 of which at least one of the first user and the second user is the owner among the plurality of speakers 200 is outputting voice, by determining the timing so that voice is output to at least one speaker 200 of which at least one of them is the owner after the speaker 200 of which at least one of them is the owner finishes outputting the voice, it is possible to suppress the situation where information cannot be correctly heard by either the first user or the second user.
[0136] Also, for example, in the timing determination step, owner information indicating each owner of the plurality of speakers 200 is acquired. When at least one speaker 200 that outputs voice has the first user as the owner among the first user and the second user, and among the plurality of speakers 200, if the second user owns at least one of the one or more speakers 200 owned by the first user, when the speaker 200 owned by the second user is outputting voice, the timing for outputting voice to the at least one speaker 200 is determined after the speaker 200 that is outputting voice finishes outputting the voice.
[0137] For example, as shown in FIG. 7, when user A and user B share the same speaker 202, it is highly likely that user A and user B are often in the same space. That is, the speakers 200 owned by user A and the speakers 200 owned by user B are likely to be arranged in the same space. Therefore, if the speakers 200 owned by user A and the speakers 200 owned by user B are simultaneously made to output sound, it may be difficult to hear the information for either user A or user B. Thus, at least one of the speakers 200 for outputting sound is owned by the first user among the first user and the second user, and among the plurality of speakers 200, when at least one of the one or more speakers 200 owned by the first user is owned by the second user, when the speaker 200 owned by the second user is outputting sound, the timing is determined so that the at least one speaker 200 for outputting sound outputs sound after the speaker 200 owned by the second user finishes outputting sound, thereby suppressing the simultaneous output of sound from the speakers 200 located in the same space.
[0138] Further, the server device 100 according to the embodiment includes a determination unit 140 that determines whether each of the plurality of speakers 200 capable of outputting sound is outputting sound, and based on the determination result of the determination unit 140, among the plurality of speakers 200, whether to immediately output sound to at least one of the speakers 200, or to output sound to the at least one speaker 200 after waiting until the speaker 200 that is outputting sound finishes outputting the sound, and an output unit 160 that outputs sound to the at least one speaker 200 at the timing.
[0139] According to this, the same effect as the above-described voice control method according to the embodiment is achieved.
[0140] In addition, the speaker 200 according to the embodiment includes a voice output unit 230 that outputs voice, a communication unit 210 for communicating with the server device 100, and a voice control unit 220 that causes the voice output unit 230 to output voice based on the voice information received from the server device 100 via the communication unit 210. The voice control unit 220 transmits voice output information indicating whether or not the voice output unit 230 is outputting voice to the server device 100 via the communication unit 210, and receives voice information and timing information indicating the timing for outputting voice from the server device 100 via the communication unit 210, and causes the voice output unit 230 to output voice based on the received voice information at the timing based on the received timing information.
[0141] According to this, the speaker 200 can suppress the output of the voice based on the voice information received from the server device 100 together with other voices, making it difficult for the user to hear.
[0142] (Other embodiments) As described above, the voice control method and the like according to the present disclosure have been described based on the embodiments. However, the present disclosure is not limited to the above embodiments.
[0143] For example, the device 600 and the speaker 200 may be the same device or different devices. That is, the device that transmits device information, operation information, etc. to the server device 100 and the device that is controlled by the server device 100 to output a spoken sentence as voice may be the same device or different devices.
[0144] Further, for example, the server device 100 may obtain device information and operation information regarding the device 600 from a server device other than the device 600 or the like. Further, the server device 100 may obtain information such as transportation services, weather information, or disaster prevention information used by the user using the device 600 from the other server device, and cause the speech body 200 to speak this information. Further, for example, the server device 100 may cause the speech body 200 owned by the user to speak service information such as the above-described transportation service used by the user. For example, when the server device 100 receives the above-described service information from another server device or the like, it may cause the speech body 200 owned by the user to speak a voice such as "There is one piece of luggage scheduled to be delivered tomorrow morning". The server device 100 may receive information regarding the service used by the user from a smartphone, tablet terminal, personal computer, or the like owned by the user. In this case, the voice speech system may not include the device 600.
[0145] Further, for example, the server device 100 may determine a speech sentence based on the device information and operation information obtained from the device 600 and the information obtained from the other server device. For example, when the device 600 is a washing machine, the server device 100 may cause the speech body 200 to speak a speech sentence recommending the drying operation of the washing machine to the user based on the information indicating that the selection by the washing machine has ended obtained from the washing machine and the weather information obtained from the other server device.
[0146] Further, for example, the plurality of speech bodies 200 determined by the determination unit 140 may be all the speech bodies 200 included in the voice speech system 500, or may be a plurality of speech bodies 200 necessary for the timing determination unit 150 to determine the timing among all the speech bodies 200 included in the voice speech system 500.
[0147] Also, for example, in FIGS. 3 to 8, an example was described in which user A and user B are each the owner of two speakers 200, and users A and B share the speaker 202 among the plurality of speakers 200. The number of speakers 200 owned by each of user A and user B, and the number of speakers 200 shared by user A and user B may each be one, or may be plural, may be the same, or may be different, and may be arbitrary.
[0148] Also, for example, in the above embodiment, the speaker during speech waiting starts a new speech after the speech of the currently speaking speaker ends. However, depending on the speech content, one speaker may interrupt and start speaking during the speech of another speaker. The speech content may be arbitrarily determined in advance and is not particularly limited.
[0149] Also, for example, in the above embodiment, all or part of the components of the processing units such as the acquisition unit 110, the scenario determination unit 120, and the speaker determination unit 130 provided in the server device 100 may be configured by dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as an HDD or a semiconductor memory.
[0150] Also, for example, the components of the above processing unit may be configured by one or more electronic circuits. Each of the one or more electronic circuits may be a general-purpose circuit or a dedicated circuit.
[0151] One or more electronic circuits may include, for example, semiconductor devices, ICs (Integrated Circuits), or LSIs (Large Scale Integrations). The IC or LSI may be integrated on one chip or on multiple chips. Here, we call it an IC or LSI, but the name may change depending on the degree of integration, and it may be called a system LSI, VLSI (Very Large Scale Integration), or ULSI (Ultra Large Scale Integration). Also, an FPGA (Field Programmable Gate Array) programmed after the manufacture of the LSI can be used for the same purpose.
[0152] Also, all or part of the components of a processing unit such as the voice control unit 220 included in the speaker 200 may be configured by dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as an HDD or a semiconductor memory.
[0153] Also, for example, the components of the above processing unit may be composed of one or more electronic circuits.
[0154] Also, the general or specific aspects of the present disclosure may be realized by a system, a device, a method, an integrated circuit, or a computer program. Alternatively, it may be realized by a computer-readable non-transitory recording medium such as an optical disk, an HDD, or a semiconductor memory storing the computer program. Also, it may be realized by any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
[0155] In addition, as long as it does not depart from the spirit of the present disclosure, various modifications conceived by those skilled in the art applied to this embodiment, or forms constructed by combining components in different embodiments are also included in the scope of the present disclosure.
Industrial Applicability
[0156] The present disclosure can be applied to a device that controls a device capable of outputting sound.
Explanation of Signs
[0157] 100 Server device 110 Acquisition unit 120 Scenario determination unit 130 Speaker determination unit 140 Determination unit 150 Timing determination unit 160 Output unit 170 Storage unit 200, 201, 202, 203 Speaker 210 Communication unit 220 Voice control unit 230 Voice output unit 500 Voice speaking system 600 Device
Claims
1. A voice control method executed by a server device, comprising: a determination step of determining whether each of a plurality of speakers capable of outputting voice is outputting voice; a timing determination step of determining, based on the determination result in the determination step, whether to immediately output voice to at least one of the plurality of speakers, or to wait until the speaker that is outputting voice finishes outputting the voice and then determine the timing to output voice to the at least one speaker; an output step of outputting voice to the at least one speaker at the timing determined in the timing determination step, wherein in the timing determination step, owner information indicating the owner of each of the plurality of speakers is acquired, when a speaker owned by the same owner as the at least one speaker among the plurality of speakers is outputting voice, determining the timing to output voice to the at least one speaker after the speaker finishes outputting voice Voice control method.
2. In the timing determination step, when a speaker owned by the same owner as the at least one speaker that is the target of the speech sentence to be output by voice among the plurality of speakers is outputting voice, determining the timing to output voice to the at least one speaker after the speaker finishes outputting voice The voice control method according to Claim 1.
3. A voice control method executed by a server device, comprising: a determination step of determining whether each of a plurality of speakers capable of outputting voice is outputting voice; a timing determination step of determining, based on the determination result in the determination step, whether to immediately output voice to at least one of the plurality of speakers, or to wait until the speaker that is outputting voice finishes outputting the voice and then determine the timing to output voice to the at least one speaker; an output step of outputting voice to the at least one speaker at the timing determined in the timing determination step, wherein in the timing determination step, owner information indicating the owner of each of the plurality of speakers is acquired, When the owners of the at least one speaker are the first user and the second user, when a speaker of which at least one of the first user and the second user is the owner among the plurality of speakers outputs voice, the timing for causing the at least one speaker to output voice after the speaker finishes outputting voice is determined. Voice control method.
4. A voice control method executed by a server device, a determination step of determining whether each of a plurality of speakers capable of outputting voice is outputting voice; a timing determination step of determining, based on the determination result in the determination step, whether to immediately cause at least one of the plurality of speakers to output voice, or to cause the at least one speaker to output voice after waiting until the speaker that is outputting voice finishes outputting the voice; an output step of causing the at least one speaker to output voice at the timing determined in the timing determination step, and including, in the timing determination step, owner information indicating the owner of each of the plurality of speakers is acquired, when the at least one speaker is owned by the first user among the first user and the second user, and among the plurality of speakers, when the second user owns at least one of the one or more speakers owned by the first user, when the speaker owned by the second user outputs voice, the timing for causing the at least one speaker to output voice after the speaker finishes outputting voice is determined. Voice control method.
Citation Information
Patent Citations
Voice output control system, and voice output device
JP2009265278A
Navigation device, navigation system, terminal device, navigation server, navigation method, and program
JP2011163778A
Voice server
JP2015164251A
A unified framework for device configuration, interaction and control, and related methods, devices and systems
JP2016502137A
Information processing device and information processing method
WO2019087546A1