Audio output device, speech recognition device, and speech recognition system
The voice recognition system accurately differentiates between audio signals from audio output devices and user voices by identifying speech portions and generating a distinct signal, enhancing recognition accuracy.
Patent Information
- Application Number
- PCT/JP2024/019887
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
Existing voice recognition systems fail to distinguish between audio signals generated by audio output devices and real user voices, leading to potential misinterpretation and erroneous recognition.
A voice signal acquisition unit identifies speech portions with meaning, generates an identification signal for user-spoken voice, and an output unit outputs this signal to differentiate between device-generated and user-spoken audio.
Effectively distinguishes between audio signals from audio output devices and user voices, ensuring accurate voice recognition and preventing misinterpretation.
Smart Images

Figure JP2024019887_04122025_PF_FP_ABST
Abstract
Description
Audio output device, audio recognition device, and audio recognition system
[0001] The present invention relates to a voice output device, a voice recognition device, and a voice recognition system using the same, which identify an acquired voice signal and determine whether the voice signal is valid or invalid in a voice recognition device that recognizes and interprets voice.
[0002] 2. Description of the Related Art In recent years, voice recognition technology has improved, and it has become common to recognize voices and control home appliances and the like based on the results of interpreting the voice content.
[0003] For example, it has become possible for a smartphone to recognize the voice of the user operating the smartphone and operate home appliances such as air conditioners and lighting equipment based on the results of interpreting the voice content.
[0004] Special table number 2016-524193
[0005] In the above-mentioned Patent Document 1, an audio device that receives an audio signal and performs voice recognition detects a wake expression (a wake-up word that is a word used to activate voice recognition operation), and if the wake expression is received from many directions, it is determined to have been generated by the audio device, and if the wake expression is received from a single direction or a limited number of directions, it is determined to have been spoken by the user.
[0006] As a result, the system distinguishes whether the wake expression is generated by the audio device or spoken by the user, and if it is determined that the wake expression was generated by the audio device, the wake expression is invalidated, and if it is determined that the wake expression was spoken by the user, the wake expression is valid.
[0007] However, this is based on the premise that the audio signal has a wake expression, and no consideration has been given to audio signals that do not have a wake expression.
[0008] An object of the present invention is to provide a voice output device and a voice recognition device that are capable of performing appropriate voice recognition, and a voice recognition system that uses them.
[0009] In order to solve the above problem, the present invention comprises a voice signal acquisition unit that acquires a voice signal, a voice signal judgment unit that determines a speech portion that can be recognized as having a meaning as language from the voice signal acquired by the voice signal acquisition unit, a voice signal processing unit that generates an identification signal for the speech portion of the voice signal acquired by the voice signal acquisition unit to identify the voice signal as a voice spoken by a user, and an output unit that outputs the identification signal.
[0010] By using the technology of the present invention, it is possible to appropriately distinguish between an audio signal generated by an audio output device such as an audio device and an audio signal (real voice) uttered by a user.
[0011] FIG. 10 is a schematic diagram illustrating a system overview of Example 1. FIG. 11 is a system configuration diagram illustrating an example of the internal configuration of a television of Example 1. FIG. 12 is a flowchart illustrating a main processing procedure in the television of Example 1. FIG. 13 is a schematic diagram illustrating an identification signal in Example 1. FIG. 14 is a system configuration diagram illustrating an example of the internal configuration of a smartphone of Example 1. FIG. 15 is a flowchart illustrating a main processing procedure in the smartphone of Example 1. FIG. 16 is a flowchart illustrating a processing procedure of a voice response function in the smartphone and network server of Example 1. FIG. 17 is a timing chart illustrating another example of the identification signal of Example 1. FIG. 18 is a schematic diagram illustrating a system overview of Example 2. FIG. 19 is a schematic diagram illustrating a system overview of Example 3. FIG. 19 is a flowchart illustrating a main processing procedure in a renderer of Example 3. FIG. 19 is a flowchart illustrating a main processing procedure in the television of Example 4. FIG. 19 is a flowchart illustrating a main processing procedure in the smartphone of Example 4. FIG. 19 is a flowchart illustrating a main processing procedure in the television of Example 5. FIG. 19 is a flowchart illustrating a main processing procedure in the smartphone of Example 5. FIG. 19 is a flowchart illustrating a main processing procedure in the television of Example 6. FIG. 19 is a flowchart illustrating a main processing procedure in the smartphone of Example 6. FIG. 19 is an external view showing an example of a smartphone used in Example 7. FIG. 19 is a schematic diagram illustrating a calculation principle for calculating the relative position between the smartphone and the television of Example 7. FIG. 10 is a schematic diagram illustrating a calculation principle for calculating the distance between a smartphone and a user in another embodiment.
[0012] Hereinafter, examples of embodiments of the present invention will be described with reference to the drawings.
[0013] FIG. 1 is a schematic diagram for explaining the basic system of this embodiment.
[0014] In the basic system of this embodiment, a user 10 speaks a voice signal (human voice) 11 to operate a remote-controlled home appliance 20 (hereinafter, the voice signal (human voice) 11 spoken by the user 10 will be referred to as human voice 11). The spoken human voice 11 as an instruction to operate is interpreted via a voice response function of a smartphone 1 (hereinafter, also referred to as a "smart phone"), which is an example of a smart device, and the smartphone 1 issues a remote control signal 21, thereby operating the remote-controlled home appliance 20 by voice.
[0015] However, the microphone of the smartphone 1 can also pick up sounds other than the human voice 11 emitted by the user 10. Therefore, there is a possibility that sounds other than the human voice 11 emitted by the user 10 may be erroneously recognized as the human voice 11 emitted by the user 10. In FIG. 1 , a television 30 is assumed as the audio output device, and a state in which the television 30 outputs an audio signal (speaker sound) 31 via a speaker (not shown) is shown (hereinafter, the audio signal output via the speaker will be referred to as speaker sound). While the present embodiment will be described using the television 30 as the audio output device, it goes without saying that the audio output device is not limited to the television 30. For example, any device capable of outputting audio acquired from an internal or external storage medium or broadcast waves into the surrounding space, such as a radio or audio device, may be used. Furthermore, the voice recognition device is not limited to the smartphone 1, and any device with a voice recognition function, such as a smart speaker, may be used.
[0016] In this embodiment, the smartphone 1 distinguishes between the human voice 11 spoken by the user 10 and the speaker sound 31 emitted from the television 30, thereby enabling the human voice 11 spoken by the user 10 to be valid and the speaker sound 31 emitted from the television 30 to be invalid.
[0017] Furthermore, the smartphone 1 in this embodiment is connected to an Internet network 17 via an access point 15, and transmits and receives data to and from a network server 16 on the Internet network 17. The smartphone 1 utilizes this mechanism to realize a voice response function, which will be described later.
[0018] The present embodiment will be described below with reference to the drawings. In this specification and the drawings, the same components are denoted by the same reference numerals.
[0019] Here, the television 30 used as the audio output device in this embodiment will be described with reference to FIGS.
[0020] [Example of Television System Configuration] The television 30 main body used in this embodiment is composed of various blocks described below. Fig. 2 is a system configuration diagram showing an example of the internal configuration of the television 30 of this embodiment. The television 30 is composed of a control unit 302, a system bus 303, a storage unit 304, a tuner 305, an illuminance sensor 355, a communication unit 306, a video unit 307, an audio unit 308, and an operation input unit 309.
[0021] The control unit 302 controls the entire television 30 in accordance with a predetermined operation program. The control unit 302 may be configured, for example, by a microprocessor unit. The system bus 303 is a data communication path for transmitting and receiving various commands, data, and the like between the control unit 302 and each unit within the television 30.
[0022] The storage unit 304 stores programs for controlling the operation of the television 30, settings for the television 30 (such as brightness and volume), object information including video and audio content, and the like. Examples of object information include recording files obtained by receiving and recording television broadcasts, and audio / video files downloaded from the Internet 17 or an external storage medium. The programs and object information stored in the storage unit 304 may be obtained or updated by downloading from a network server (not shown). The storage unit 304 also stores information that must be retained even when power is not supplied to the television 30 from an external source. The storage area that must retain this information may include devices such as semiconductor memory such as flash ROM and magnetic disk drives such as HDDs. Furthermore, the storage capacity can be increased by connecting an external HDD or, for example, a cassette HDD to the television 30.
[0023] The tuner 305 is connected to an antenna (not shown) and receives digital television broadcasts (terrestrial, BS, CS). The content of the received digital television broadcast is output to the outside via a video unit 307 and an audio unit 308 under the control of the control unit 302.
[0024] The illuminance sensor 355 is a sensor that detects the brightness around the television 30, and the result can be reflected in the brightness of the display 371. The communication unit 306 is composed of a LAN (Local Area Network) communication device 361, a short-range wireless communication device 363, and an infrared communication device 364. The LAN communication device 361 is connected to the Internet network 17 via the access point 15, and transmits and receives data to and from a network server 16 on the Internet network 17. Note that there may be multiple network servers 16 for each service or service content.
[0025] The short-range wireless communication device 363 performs short-range wireless communication with a device having a short-range wireless communication function. This short-range wireless communication allows the audio signal of the television 30 to be heard wirelessly using wireless earphones or speakers that support short-range wireless communication. In this embodiment, Bluetooth (registered trademark) communication is used as the short-range wireless communication. The infrared communication device 364 receives a remote control signal (infrared rays) emitted from a remote control that operates the television 30. As a result, the television 30 can be operated using the remote control.
[0026] The LAN communication device 361 and the short-range wireless communication device 363 each include an encoding circuit, a decoding circuit, an antenna, etc. Furthermore, the communication unit 306 may include other communication devices.
[0027] The video unit 307 is composed of a display 371 such as a liquid crystal display or organic electroluminescence display, and a video processor 372 that processes video signals. The display 371 displays and provides video and additional information to a user watching the television 30. The audio unit 308 is composed of a speaker 382 and an audio processor 383 that processes audio signals. The speaker 382 outputs sounds required by the user. External earphones or headphones may be used via an earphone microphone terminal. In this embodiment, the audio processor 383 generates an identification signal that identifies the audio as having been emitted from the television 30. The identification signal will be described later.
[0028] The operation input unit 309 is an operation input key such as a push button for inputting operation instructions to the television 30. Specifically, operation inputs such as power on / off, channel switching, and volume increase / decrease are performed.
[0029] Furthermore, the television 30 can be operated wirelessly using a remote control 369 paired with the television 30. The remote control 369 emits a remote control signal 21 to the television 30. In this embodiment, infrared rays are used as the remote control signal 21. The emitted remote control signal 21 is received by an infrared communication device 364 of the communication unit 306, and the television 30 is controlled by the control unit 302.
[0030] The hardware configuration example of the television 30 shown in FIG. 2 includes many components that are not essential to this embodiment, but the effects of this embodiment are not diminished even if these components are not included.
[0031] 3 is a flowchart showing a main processing procedure in the television 30 of this embodiment. The processing procedure in FIG. 3 will be explained with reference to the system configuration diagram in FIG.
[0032] The processing procedure in the television 30 of this embodiment is realized by the program of the control unit 302 stored in the memory unit 304, using the infrared communication device 364 of the communication unit 306, the audio processor 383 of the audio unit 308, etc.
[0033] When the process starts (S431), first, an audio information acquisition process (S432) is performed to acquire audio information including audio signals and additional information. That is, the control unit 302 functions as an audio signal acquisition unit that acquires audio signals. Specifically, playback of a digital television broadcast received via the tuner 305 is started. Alternatively, playback of an audio / video file stored in the storage unit 304 is started.
[0034] Next, the control unit 302 determines whether the audio information acquired in the audio information acquisition process (S432) contains subtitle information as additional information (S433). In this embodiment, the subtitle information is the text information (closed captions) in digital television broadcasts. Furthermore, the audio-video file contains a subtitle information file, which stores three types of information (subtitle number, start and end times, and subtitle content). In this embodiment, this subtitle information file is analyzed, and an identification signal is added to the period from the start time to the end time of each subtitle. Furthermore, when playing the audio-video file, the identification signal can be added by pre-reading. Pre-reading allows for a margin of error in the voice recognition process on the smartphone 1. If subtitle information is present, it is determined that the audio information explicitly contains voice, and the smartphone 1 can identify the audio information as a target for voice recognition. In other words, the control unit 302 functions as an audio signal determination unit that determines the section in which subtitle information is valid and the speech portion of the acquired audio information whose meaning can be recognized as language.
[0035] If it is determined in the determination process of S433 that the audio information acquired in the audio information acquisition process (S432) does not contain subtitle information, the process proceeds to the process of S435, which will be described later. If it is determined in the determination process of S433 that the audio information acquired in the audio information acquisition process (S432) contains subtitle information, the process proceeds to the next process, which is the process of adding an identification signal to the audio information (S434). In this embodiment, the control unit 302 creates an identification signal using the audio processor 383 and adds the identification signal to the audio information. The audio processor 383 functions as an audio signal processing unit that generates the identification signal.
[0036] The identification signal in this embodiment will now be described with reference to Fig. 4, which is a schematic diagram for explaining the identification signal in this embodiment.
[0037] 4 is an audio output characteristic diagram in which an audio signal is represented by the frequency of the audio signal on the horizontal axis and the level of the audio signal at each frequency on the vertical axis. In this embodiment, a suppression section is formed for a specific frequency band of the audio signal, and the presence of this suppression section is used as an identification signal 1001. In this embodiment, the speaker sound 31 (see FIG. 1) output from the speaker 382 in the section corresponding to the subtitle information is an audio signal in which the specific frequency band is suppressed, as shown in FIG.
[0038] In Fig. 4, the vertical axis represents signal level and the horizontal axis represents frequency. In Fig. 4, the solid line represents the audio signal before the suppression section SS is formed, and the dashed line represents the audio signal level after the identification signal 1001 is added in the suppression section SS. Note that the frequency bands other than the suppression section SS are the same as the solid line. As shown in Fig. 4, processing is performed in the suppression section SS to lower the audio signal level. In Fig. 4, the suppression section SS is shown to have a certain width, but it is also possible to lower the signal level of only specific frequencies. This is because human voice changes gradually as shown by the solid line, and does not experience abrupt changes like the suppression section SS. Therefore, by detecting such a suppression section SS, it can be recognized as the identification signal 1001.
[0039] 4, the suppression section SS is fixed to the same frequency band, but the frequency band may be changed over time. Also, the suppression section SS does not have to be formed continuously over time, but may be formed intermittently (discontinuously).
[0040] Next, the audio signal (output audio signal in the audio information) acquired in the audio information acquisition process of S432 is output from the speaker 382 as speaker sound 31 (S435). However, for earphones that do not emit audio to the outside, voice recognition by the smartphone 1 is not performed, so the addition of an identification signal is not applied.
[0041] Next, it is determined whether the audio information has ended (S436). If it is determined in the determination process of S436 that the audio information has not ended, the process returns to the audio information acquisition process of S432 (S432) and continues. If it is determined in the determination process of S436 that the audio information has ended, the main processing procedure in the television 30 is terminated (S437).
[0042] Here, the smartphone 1 used in this embodiment will be described with reference to FIGS.
[0043] [Example of system configuration of smartphone] The main body of the smartphone 1 used in this embodiment is composed of various blocks described below. Fig. 5 is a system configuration diagram showing an example of the internal configuration of the smartphone 1 of Example 1.
[0044] The smartphone 1 is composed of a control unit 2, a system bus 3, a memory unit 4, a sensor unit 5, a communication unit 6, a video unit 7, an audio unit 8, an operation input unit 9, etc.
[0045] The control unit 2 is a microprocessor unit that controls the entire smartphone 1 according to a predetermined operation program. The system bus 3 is a data communication path for transmitting and receiving various commands and data between the control unit 2 and each unit within the smartphone 1.
[0046] The memory unit 4 stores programs for controlling the operation of the smartphone 1, operation setting values, detection values from the sensor unit 5, object information including video and audio content, library information, etc. The programs, object information, and library information stored in the memory unit 4 may be obtained or updated by downloading from a network server (not shown). The object information may also be data such as videos and still images captured using the camera function within the smartphone 1. The memory unit 4 also stores information that must be retained even when power is not supplied to the smartphone 1 from an external source. The storage area that must be retained may include devices such as semiconductor memory such as flash ROM or SD memory, or magnetic disk drives such as HDDs.
[0047] The sensor unit 5 is a group of various sensors for detecting the state of the smartphone 1. The sensor unit 5 is composed of a positioning sensor 51, a geomagnetic sensor 52, an acceleration sensor 53, a gyro sensor 54, and an illuminance sensor 55 that detects the brightness of the surroundings of the smartphone 1. These sensors make it possible to detect the position, tilt, direction, movement, etc. of the smartphone 1. Furthermore, the smartphone 1 may be equipped with other sensors such as an altitude sensor.
[0048] The communication unit 6 is composed of a LAN communication device 61, a telephone network communication device 62, a short-range wireless communication device 63, and an infrared communication device 64. The LAN communication device 61 is connected to the Internet network 17 via an access point 15 and transmits and receives data to and from a network server 16 on the Internet network 17. Note that there may be multiple network servers 16 for each service or service content. The network connection is performed wirelessly via a Wi-Fi (registered trademark) or other standard. The telephone network communication device 62 performs telephone communication (calls) and data transmission and reception via wireless communication with a base station of a mobile telephone communication network. Communication with the base station or the like may be performed via the LTE (Long Term Evolution) system, the 5G system (a fifth-generation mobile communication system aiming for high speed, large capacity, low ground clearance, and multiple simultaneous connections), or other communication methods. The short-range wireless communication device 63 performs short-range wireless communication with a device having short-range wireless communication capabilities. In this embodiment, Bluetooth (registered trademark) communication is used as the short-range wireless communication. The infrared communication device 64 performs infrared communication with the remote-controlled home electric appliance 20 that operates by a remote control signal (infrared rays), and performs remote control operation of the remote-controlled home electric appliance 20. The LAN communication device 61, the telephone network communication device 62, and the short-range wireless communication device 63 each include an encoding circuit, a decoding circuit, an antenna, etc. Furthermore, the communication unit 6 may include other communication devices.
[0049] The video unit 7 is composed of a camera 71, a display 72, and a video processor 73 that processes video signals. The camera 71 converts light input through a lens into an electrical signal using electronic devices such as a CCD (Charge Coupled Device) or a CMOS (Complementary Metal Oxide Semiconductor) sensor. The display 72 is a display device such as a liquid crystal display, and provides video and additional information to the user 10 of the smartphone 1.
[0050] The audio unit 8 is composed of a microphone 81, a speaker 82, and an audio processor 83 that processes audio signals. The microphone 81 collects sounds in the real space and the user's 10 voice 11, converts them into audio signals, and inputs them. An external microphone may be used via the earphone microphone terminal. Needless to say, it is possible to use different types of microphones depending on the application, such as connecting an external microphone to a short-range wireless communication device 63. The speaker 82 outputs audio signals as physical sounds required by the user 10. An external earphone or headphone may be used via the earphone microphone terminal. Needless to say, it is possible to use different types of microphones depending on the application, such as connecting an earphone or headphone to a short-range wireless communication device 63 and outputting audio to the user 10. In this embodiment, the audio processor 83 detects an identification signal, which will be described later.
[0051] The operation input unit 9 is an operation input unit such as a touch panel for inputting operation instructions and the like to the smartphone 1. The smartphone 1 of this embodiment also includes a vibrator 95 for vibrating the smartphone 1 and a light 96.
[0052] 5 includes many components that are not essential to this embodiment, but the effects of this embodiment are not impaired even if these components are not provided. Furthermore, components not shown, such as an electronic money payment function, may also be added.
[0053] [Smartphone Processing Procedure] Fig. 6 is a flowchart showing the main processing procedure in the smartphone 1 of this embodiment. The processing procedure in Fig. 6 will be explained with reference to the system configuration diagram in Fig. 5. The processing procedure in the smartphone 1 of this embodiment is realized by the program of the control unit 2 stored in the memory unit 4, using the infrared communication device 64 of the communication unit 6, the audio processor 83 of the audio unit 8, etc.
[0054] When the process starts (S401), first, an audio signal acquisition process (S402) is performed to acquire an audio signal from the microphone 81 of the audio unit 8. That is, a picked-up audio signal based on the sound picked up by the microphone 81 is acquired.
[0055] Next, the voice signal acquired in the voice signal acquisition process of S402 is converted into text data by the voice processor 83, and the converted text data is stored in the storage unit 4 (S403).
[0056] Next, it is determined whether the audio signal acquired in the audio signal acquisition process (S402) is the real voice 11 spoken by the user 10 (S404). The reason for determining whether the audio signal is the real voice 11 spoken by the user 10 is that, as shown in FIG. 1 , the audio signal acquired by the microphone 81 of the audio unit 8 may contain not only the real voice 11 of the user 10 but also the speaker sound 31 emitted by the television 30, and this must be distinguished. In this embodiment, an identification signal 1001 is added to the speaker sound 31 emitted by the television 30. By detecting the presence or absence of this identification signal 1001 using the audio processor 83, it is possible to distinguish between the speaker sound 31 emitted by the television 30 and the real voice 11 spoken by the user 10. That is, in step S404, the control unit 2 functions as an audio signal discrimination unit that determines whether the picked-up audio signal is a human voice 11 (first audio signal) spoken by the user 10 or a second audio signal output by the television 30 (audio output device).
[0057] If it is determined in the determination process of S404 that the voice signal acquired in the voice signal acquisition process (S402) is not the real voice 11 spoken by the user 10, the process proceeds to the determination process of S407, which will be described later. If it is determined in the determination process of S404 that the voice signal acquired in the voice signal acquisition process (S402) is the real voice 11 spoken by the user 10, it is determined whether the voice recognition content stored in the storage unit 4 in the process of S403 is a word (designated operation word) required to operate the remote-controlled home electric appliance 20 (S405).
[0058] If it is determined in the determination process of S405 that the voice-recognized words are not words (designated operation words) necessary for operating the remote-controlled home electric appliance 20, the process proceeds to the determination process of S407, which will be described later. If it is determined in the determination process of S405 that the voice-recognized words are words (designated operation words) necessary for operating the remote-controlled home electric appliance 20, the process proceeds to the next designated operation instruction (S406). The designated operation words are defined by the remote-controlled home electric appliance 20.
[0059] The designated operation instruction in S406 is transmitted to the remote-controlled home electric appliance 20 via the infrared communication device 64 as a remote control signal 21 corresponding to the designated operation word.
[0060] Next, it is determined whether the audio signal has ended (S407). If it is determined in the determination process of S407 that the audio signal has not ended, the process returns to the audio signal acquisition process of S402 and continues the main process. If it is determined in the determination process of S407 that the audio signal has ended, the main process procedure in the smartphone 1 is terminated (S408).
[0061] That is, if it is determined in the determination process of S404 that the picked-up voice signal is not the real voice 11 uttered by the user 10, S405 and S406 are not executed, and therefore the series of processes from S402 to S407 are not executed at all. In other words, if it is determined to be the second voice signal, the voice recognition unit does not perform voice recognition processing on the picked-up voice signal. On the other hand, if it is determined in the determination process of S404 that the picked-up voice signal is the real voice 11 uttered by the user 10, the series of processes from S402 to S407 are all executed. In other words, if it is determined to be the first voice signal, the control unit 2 (voice recognition unit) performs voice recognition processing on the picked-up voice signal.
[0062] Here, the process (S404) of determining whether the voice 11 is the real voice uttered by the user 10 in FIG. 6 will be described using the identification signal 1001.
[0063] The speaker sound 31 emitted from the television 30 is subjected to frequency processing only when the audio information includes audio, forming a suppression section SS and generating an identification signal 1001. Therefore, in the determination process of S404, the audio processor 83 checks for the presence of a suppression section, which is the identification signal 1001, and if a suppression section is detected, it is determined to be the speaker sound 31 from the television 30. As a result, for audio signals in which the smartphone 1 detects the identification signal 1001, the smartphone 1 does not perform audio operation processing based on the audio signal, thereby disabling the audio for the speaker sound 31 output from the television 30.
[0064] 6, the order of S403 and S404 may be reversed. The presence or absence of the identification signal 1001 in the voice signal acquired in S402 is detected, and if the identification signal 1001 is not detected, the voice signal is converted into text data. In other words, it may be determined first whether the voice signal is a human voice 11 before converting it into text data.
[0065] The processing procedure in the smartphone 1 of this embodiment is realized by the program of the control unit 2 stored in the memory unit 4, using the LAN communication device 61 of the communication unit 6, the voice processor 83 of the voice unit 8, etc. On the other hand, part of the processing procedure in the smartphone 1 may be realized by the network server 16. The processing procedure in this case will be explained using Figure 7. Figure 7 is a flowchart showing the processing procedure of the voice response function of the smartphone 1 and the network server 16.
[0066] 7 is a flowchart showing the processing steps of the voice response function of the smartphone 1 and the network server 16. In this embodiment, the network server 16 has an AI assistant function and is generally referred to as a cloud. In FIG. 7, the processing steps of the smartphone 1 are shown in the left half of the processing steps (S461 to S467), and the processing steps of the network server 16 (cloud) are shown in the right half of the processing steps (S471 to S476).
[0067] When the process starts (S471), the network server 16 (cloud) goes into standby mode to wait for the reception of text data.
[0068] When the smartphone 1 starts processing (S461), it first performs a voice signal acquisition process (S462) to acquire a voice signal from the microphone 81 of the audio unit 8. Next, the voice signal acquired in the voice signal acquisition process of S462 is speech-recognized (converted to text) by the voice processor 83, and the text data is stored in the storage unit 4 (S463). Next, it connects to the Internet network 17 via the access point 15, transmits the text data stored in S463 to the network server 16 (S464), and goes into standby mode to wait for a reply from the network server 16 (cloud).
[0069] The network server 16 (cloud), which is on standby waiting to receive text data, receives the text data sent from the smartphone 1 (S472). Next, it interprets the received text data (S473). Next, it creates a reply to the smartphone 1 based on the results of the interpretation in S473 (S474). Next, it transmits the reply to the smartphone 1 (S475). This series of processes from S472 to S475 is repeated each time text data is received, and then ends (S476).
[0070] The smartphone 1, which is in standby mode waiting to receive a reply message, receives the reply message sent from the network server 16 (cloud) (S465). Next, the smartphone 1 outputs the reply message from the network server 16 (cloud) as audio through the speaker 82 built into the smartphone 1 (S466), and ends the processing of the voice response function of the smartphone 1 (S467).
[0071] The operation of FIG. 7 described above is the operation of a general voice response function, but the smartphone 1 of this embodiment responds only to user utterances by replacing part of FIG. 7 with part of the operation of FIG. 6. Specifically, S462 corresponds to S402 in FIG. 6, and S463 corresponds to S403 and S404 in FIG. 6. Here, if S404 is "N", S464 and subsequent steps are not executed. In other words, if it is determined that the voice 11 is not the real voice of the user 10, it is not necessary to create a response sentence, and the processing from step S464 in FIG. 7 onwards is not executed. If it is determined that the voice 11 is not the real voice of the user 10, the network server 16 also does not execute any processing.
[0072] Through the above processing, when the user 10 speaks to the smartphone 1, the smartphone 1 appears to be able to understand the content and respond with voice.
[0073] Therefore, when the user 10 speaks a specified operation word to the smartphone 1 to operate the remote-controlled home appliance 20, the smartphone 1 transmits the specified operation word to the network server 16 (cloud) according to the above procedure, and receives from the network server 16 (cloud) the specified operation word to operate the remote-controlled home appliance 20. As a result, the smartphone 1 sends a remote control signal 21 to the remote-controlled home appliance 20 according to the specified operation word, and can operate the remote-controlled home appliance 20.
[0074] Of course, this voice response can be used for purposes other than operating remote-controlled home appliances. For example, it can be applied to checking emails and managing schedules on the smartphone 1.
[0075] In this embodiment, the voice response is realized using the network server 16 (cloud), but it can also be realized using the voice processor 83 of the smartphone 1. In this embodiment, the identification signal 1001 is generated by the voice processor 383, but it goes without saying that the generation of the identification signal 1001 can be realized by software alone or software in combination with hardware. Also, in this embodiment, the detection of the identification signal 1001 and voice recognition are performed by the voice processor 83, but it goes without saying that the detection of the identification signal 1001 and voice recognition can be realized by software alone or software in combination with hardware.
[0076] In this embodiment, the identification signal 1001 for identifying the speaker sound 31 is realized by frequency processing, but the identification signal can also be realized by other means than frequency processing.
[0077] For example, an audio digital watermark can be used as the identification signal. Audio digital watermarking methods include periodic phase modulation and echo diffusion, both of which are difficult for the human ear to distinguish. This audio digital watermark is added as an identification signal to the audio signal (speaker sound) 31 output from the television 30, which is an audio output device. The smartphone 1 can then detect the presence or absence of the audio digital watermark to distinguish between the human voice 11 emitted by the user 10 and the speaker sound 31. Therefore, even if an audio digital watermark is used as the identification signal, the effect of the above-described embodiment is the same, and the speaker sound 31 output from the television 30, which is an audio output device, can be disabled.
[0078] Furthermore, ultrasonic waves can also be used as the identification signal. Ultrasound is a high-frequency sound that exceeds the range of human hearing, and generally refers to sound waves with a vibration frequency of 20 kHz or higher. Using ultrasonic waves as the identification signal requires that the speaker 382 of the television 30, which is the audio output device, be a wideband speaker, and the microphone 81 of the smartphone 1, which is the voice recognition device, be a wideband microphone. However, this can be achieved by employing ultrasonic waves capable of transmitting and receiving signals. When using ultrasonic waves as the identification signal, the audio output device transmits ultrasonic waves only in sections where subtitle information is present, and the voice recognition device disables voice recognition during sections where ultrasonic waves are received. Therefore, even if ultrasonic waves are used as the identification signal, the effect of the above-described embodiment is the same, and the speaker sound 31 output from the television 30, which is the audio output device, can be disabled.
[0079] In the above description, an identification signal is added to the time from the start time to the end time of the subtitles, but an identification signal indicating the start of a sound that is not to be recognized by the smartphone 1 may be added just before the start time of the subtitles, prior to the subtitle information. Then, just before the end of the subtitles, an identification signal indicating the end of the sound that is not to be recognized by the smartphone 1 may be added just before the end of the subtitles, prior to the end of the subtitle information.
[0080] An example is shown in Figure 8. Figure 8 is a timing chart showing the relationship between closed caption information and identification signals. In Figure 8, the section containing closed caption information is between times t2 and t4. Therefore, an identification signal indicating the start described above is added between times t1 and t2, just before time t2, and an identification signal indicating the end described above is added between times t3 and t4, just before time t4. By configuring as shown in Figure 8, it is possible to suppress the increase in the amount of information due to the addition of identification signals and reduce the deterioration of sound quality.
[0081] The following describes a second embodiment of the present invention. The basic hardware and software configurations of the second embodiment are the same as those of the previous embodiment, and the following mainly describes the differences between this embodiment (second embodiment) and the previous embodiment, and omits explanations of common parts as much as possible to avoid duplication.
[0082] In the above-described embodiment, the remote control signal 21 was sent from the smartphone 1, but in this embodiment, the remote control signal is sent using a smart remote control.
[0083] A smart remote control is a remote control that can simultaneously control multiple remote control devices, instead of the multiple remote controls that were previously required.
[0084] 9 is a schematic diagram for explaining an overview of a system when using the smart remote control in this embodiment. For simplicity of the diagram, the television 30, which is an audio output device, is omitted.
[0085] The smartphone 1 is connected to an internet network 17, to which a network server 16 is connected, via an access point 15. The network server 16 includes a network server that performs various types of calculation processing and a network server that stores various types of data, and the smartphone 1 can utilize these servers as needed. A smart remote control 19 is connected to the internet network 17, and the smartphone 1 controls the smart remote control 19 via the internet network 17. Under the control of the smartphone 1, the smart remote control 19 sends a remote control signal 21 to a remote-controlled home appliance 20. As a result, the user 10 can remotely operate the remote-controlled home appliance 20 via the smartphone 1 using their voice 11.
[0086] In this embodiment, voice recognition and identification signal detection are performed by the smartphone 1, as in the previous embodiment. The addition of an identification signal for identifying speaker sound is performed by the television 30, as in the previous embodiment. This embodiment also has the same effect as the first embodiment, and the speaker sound 31 output from the television 30, which is an audio output device, can be disabled. It goes without saying that the functions of the smartphone 1 in this embodiment can also be realized by a smart speaker.
[0087] The following describes a third embodiment of the present invention. The basic hardware and software configurations of the third embodiment are the same as those of the above-described embodiments, and the following mainly describes the differences between this embodiment (fourth embodiment) and the above-described embodiments, and omits explanations of common parts as much as possible to avoid duplication.
[0088] In the above-described embodiment, the audio signal (channel-based speaker sound) emitted from the speaker of the television 30 was described, but in this embodiment, object-based audio information will be described. The object-based audio information is composed of audio information (audio objects) of each sound source and audio metadata. The audio metadata stores, in chronological order, audio source auxiliary data and position information in three-dimensional space for each sound source (object).
[0089] 10 is a schematic diagram illustrating an overview of the object-based audio information reproduction system of this embodiment. The object audio transmission device 33 transmits audio metadata stored in chronological order to a renderer (signal processing device) 34. The renderer (signal processing device) 34 transmits the object-based audio information transmitted from the object audio transmission device 33 as an optimal audio signal to each of a group of speakers 35 arranged in a three-dimensional space. As a result, the spatial position of each sound source is fixed, and the sound can be transmitted to the smartphone 1 as speaker sound 32.
[0090] In this embodiment, in order to distinguish between the human voice 11 uttered by the user 10 and the speaker sound, the sound source accompanying data in the acoustic metadata is analyzed, and an identification signal is added only to object audio signals that can be recognized as voices.
[0091] 11 is a flowchart showing the main processing procedure in the renderer 34 of this embodiment. In this embodiment, the renderer 34 and the speaker group 35 serve as audio output devices.
[0092] When the process starts (S481), first, an audio metadata acquisition process (S482) is performed to acquire audio metadata from the object audio transmission device 33. Next, audio source attribute information is extracted from the acquired audio metadata (S483). Next, the audio source attribute information extracted in the process of S483 is analyzed to determine whether or not an audio source (dialogue) representing human speech is present (S484). In other words, an audio object (audio source) that is a dialogue is determined to be a speech portion.
[0093] If it is determined in the determination process of S484 that there is no voice source, the process proceeds to the process of S486, which will be described later. If it is determined in the determination process of S484 that there is a voice source, the process proceeds to the process of S485. The process of S485 is a process of adding an identification signal to the audio signal of the voice source, and the identification signal is added only to the audio signal of the voice source.
[0094] Next, the acquired acoustic metadata is analyzed, and an optimal audio signal for each sound source (object) is output to each speaker of the speaker group 35 arranged in three-dimensional space so that the spatial position of each sound source (object) is localized (S486). Next, it is determined whether the acoustic metadata is complete (S487). If it is determined in the determination process of S487 that the acoustic metadata is not complete, the process returns to the acoustic metadata acquisition process of S482 and continues the main process. If it is determined in the determination process of S487 that the acoustic metadata is complete, the main processing procedure in the renderer 34 is terminated (S488).
[0095] The identification signal in this embodiment is realized by the same processing as the identification signal in the above-described embodiment 1. That is, frequency processing, digital watermarking, ultrasonic waves, etc. may be used. The processing of voice recognition of the speaker sound 32 in the smartphone 1 is the same as in the above-described embodiment, and therefore a description thereof will be omitted.
[0096] In this embodiment, an identification signal is added only to the object audio signal of the audio source that can be recognized by voice recognition. As a result, when the smartphone 1 detects the identification signal, the smartphone 1 does not perform processing based on the audio signal, thereby disabling the speaker sound 32 output from the speaker group 35.
[0097] The following describes a fourth embodiment of the present invention. The basic hardware and software configurations of the fourth embodiment are the same as those of the above-described embodiments, and the following mainly describes the differences between this embodiment (fourth embodiment) and the above-described embodiments, and omits explanations of common parts as much as possible to avoid duplication. In the first to third embodiments described above, an identification signal is added to speaker sound 31 output from a television 30, which is an audio output device, and is used to distinguish it from the human voice 11 emitted by a user 10.
[0098] In this embodiment, the speaker sound 31 output from the television 30, which is an audio output device, and the human voice 11 uttered by the user 10 are distinguished from each other without adding the identification signal 1001 to the audio.
[0099] The system configuration of this embodiment is assumed to be the same as that of the first embodiment. Fig. 12 is a flowchart showing a main processing procedure in the television 30 of this embodiment. The processing procedure in Fig. 12 will be explained with reference to the system configuration diagram in Fig. 2. The processing procedure in the television 30 of this embodiment is realized by a program of the control unit 302 stored in the storage unit 304, utilizing a connection to the Internet network 17 via the LAN communication device 361 of the communication unit 306. The processing procedure in Fig. 12 is the same as the processing procedure in Fig. 3 of the first embodiment, and only the processing that differs from the processing procedure in Fig. 3 will be explained. In this embodiment, the processing of adding an identification signal to an audio signal in Fig. 3 (S434) is replaced by a text data extraction and output processing (S441).
[0100] The process of S441 is a process of extracting text data from the subtitle information and sending the extracted text data as an identification signal to the smartphone 1 via the Internet network 17. That is, in this embodiment, the LAN communication device 361 functions as an information transmission unit that transmits information to the outside. By the process of sending the text data to the smartphone 1 (S441), the smartphone 1 can obtain the text data.
[0101] Fig. 13 is a flowchart showing the main processing procedure in the smartphone 1 of this embodiment. The processing procedure in Fig. 13 will be explained with reference to the system configuration diagram in Fig. 5. The processing procedure in the smartphone 1 of this embodiment is realized by the program of the control unit 2 stored in the memory unit 4 using the LAN communication device 61 of the communication unit 6. That is, in this embodiment, the LAN communication device 61 functions as an information receiving unit that receives information from the outside.
[0102] The processing procedure in Fig. 13 is the same as the processing procedure in Fig. 6 in the first embodiment described above, and only the processing different from the processing procedure in Fig. 6 will be described. In this embodiment, the processing for determining the user utterance in Fig. 6 (S404) is a series of processing that continues from the text data reception determination processing (S411) to the determination processing in S413.
[0103] If it is determined in the determination process (S412) of whether text data (first text information) has been received in S411 that no text data has been received, the process proceeds to S405 and subsequent processes. In other words, if no text data has been received, it can be determined that the voice 11 is the user 10's voice and not the speaker sound 31 from the television 30, and the subsequent processes proceed to S405 and subsequent processes as in Fig. 6. If it is determined in the determination process of S411 that text data has been received, the text data (second text information) saved in the process of S403 is compared with the received text data (S412).
[0104] Next, it is determined whether the recognition result stored in the process of S403 matches the received text data (S413).
[0105] If it is determined in the determination process of S413 that the text data saved in the process of S403 matches the received text data, it is determined that the acquired voice signal is speaker sound 31 of television 30, and the process proceeds to S407. If it is determined in the determination process of S413 that the recognition result saved in the process of S411 does not match the received text data, it is determined that the acquired voice signal is the human voice 11 spoken by user 10, and the process proceeds to S405. Therefore, if the text data saved in the process of S403 matches the received text data, voice recognition processing is not performed.
[0106] In this embodiment, if subtitle information is present, the television 30 transmits text data, and by comparing the content of this text data with the recognition results of the acquired audio signal, it is possible to distinguish between the speaker sound 31 emitted by the television 30 and the actual voice 11 spoken by the user 10.
[0107] In this embodiment, the television 30 serving as an audio output device acquires subtitle information from audio information, but it goes without saying that the television 30 serving as an audio output device can also perform voice recognition on the audio signal itself and generate text data. This embodiment also has the same effects as the previous embodiments, and the speaker sound 31 output from the television 30 serving as an audio output device can be disabled. It goes without saying that this embodiment can also be applied to the object audio (embodiment 3) described above.
[0108] The following describes a fifth embodiment of the present invention. The basic hardware and software configurations of the fifth embodiment are the same as those of the above-described embodiments, and the following mainly describes the differences between this embodiment (fifth embodiment) and the above-described embodiments, and omits explanations of common parts as much as possible to avoid duplication. While the fourth embodiment did not consider activation words (voice commands for starting voice recognition), this embodiment makes effective use of activation words.
[0109] Fig. 14 is a flowchart showing a main processing procedure in the television 30 of this embodiment. The processing procedure in Fig. 14 will be explained with reference to the system configuration diagram in Fig. 2. The processing procedure in the television 30 of this embodiment is realized by a program of the control unit 302 stored in the storage unit 304, utilizing a connection to the Internet network 17 via the LAN communication device 361 of the communication unit 306. That is, in this embodiment, the LAN communication device 361 functions as an information transmission unit that transmits information to the outside.
[0110] The processing procedure in Fig. 14 includes some steps that are the same as those in Fig. 12 in the fourth embodiment, but also has some fundamental differences. Therefore, the processing procedure in Fig. 14 will be described below.
[0111] When the process starts (S431), the initial setting is first performed (S442). In the initial setting (S442), the following two settings are performed.
[0112] (1) The activation flag indicating that the activation word has been detected is cleared and the activation flag is disabled.
[0113] (2) Clear the timer that indicates the validity period of the activation word (set the timer value to 0).
[0114] Next, an audio information acquisition process (S432) is performed to acquire audio information. Specifically, acquisition of digital television broadcasting received via the tuner 305 is started. Alternatively, acquisition of audio / video files stored in the storage unit 304 is started.
[0115] Next, it is determined whether or not the audio information acquired in the audio information acquisition process (S432) contains subtitle information (S433). In this embodiment, the text information (closed captions) in digital television broadcasting is used as the subtitle information. Furthermore, a subtitle information file exists in the audio-video file, and three types of information (subtitle number, start time, end time, and subtitle content) are stored. If it is determined in the determination process of S433 that there is no subtitle information (no audio signal), the process proceeds to the determination process of S446, which will be described later. If it is determined in the determination process of S433 that there is subtitle information (an audio signal), the process proceeds to the next determination process (S443).
[0116] In the determination process of S443, the subtitle information is analyzed to determine whether the subtitle information is an activation word. If it is determined in the determination process of S443 that the subtitle information is not an activation word, the process proceeds to the determination process of S446, which will be described later. If it is determined in the determination process of S443 that the subtitle information is an activation word, the process proceeds to the next process (S444). In the process of S444, a timer indicating the effective time of the activation word is cleared and started (the timer value starts counting from 0). Then, the process proceeds to the process of S445.
[0117] In the process of S445, since the hot word is detected, the hot word detection flag is set to 1 and transmitted to the smartphone 1 (S445). Then, the process proceeds to the determination process of S446. This hot word flag corresponds to hot word information.
[0118] In the determination process of S446, it is determined whether the timer has exceeded a specified value (the maximum time of the audio signal section in which the activation word is valid). If it is determined in the determination process of S446 that the time has not expired, the process proceeds to the process of S435. If it is determined in the determination process of S446 that the time has expired, the transition has been made to an audio signal section in which the activation word is invalid, so the activation flag, which is an activation word detection flag, is set to 0 and transmitted to the smartphone 1 (S447). Then, the process proceeds to the process of S435.
[0119] The process of S435 is a process of outputting the audio signal (audio signal within the audio information) acquired in the audio information acquisition process of S432 as speaker sound from the speaker 382. Next, it is determined whether the audio information has ended (S436). If it is determined in the determination process of S436 that the audio information has not ended, the process returns to the audio information acquisition process of S432 and continues. If it is determined in the determination process of S436 that the audio information has ended, the main processing procedure in the television 30 is terminated (S437).
[0120] This completes the description of the processing procedure in Fig. 14. Note that when the user 10 utters an activation word, the smartphone 1 manages the activation word effective voice signal section for the user 10 using a processing procedure equivalent to that in Fig. 14.
[0121] FIG. 15 is a flowchart showing the main processing procedure in the smartphone 1 of this embodiment. The processing procedure in FIG. 15 will be explained with reference to the system configuration diagram in FIG. 5. The processing procedure in the smartphone 1 of this embodiment is realized by the program of the control unit 2 stored in the memory unit 4 using the LAN communication device 61 of the communication unit 6. In other words, in this embodiment, the LAN communication device 61 functions as an information receiving unit that receives information from the outside. The processing procedure in FIG. 15 is equivalent to the processing procedure in FIG. 6 in the above-mentioned embodiment 1, and only processing that differs from the processing procedure in FIG. 6 will be explained.
[0122] In this embodiment, the process of S414 is added immediately after the start, and the determination process (S404) for determining the user utterance is replaced by the determination processes of S415 and S416.
[0123] The processing of S414 is processing for setting the startup flag to 0 as an initial setting and disabling the startup flag. The startup flag information transmitted from the television 30 to the smartphone 1 is sequentially received and sequentially stored as a startup flag in the memory unit 304 of the smartphone 1.
[0124] The determination process of S415 is a determination process for determining whether the text data saved in the process of S403 is a activation word. If it is determined in the determination process of S415 that the text data saved in the process of S403 is not a activation word, the voice signal acquired in S402 is invalidated, and the process proceeds to S407.
[0125] If the determination process of S415 determines that the text data saved in the process of S403 is a wake-up word, the process proceeds to the next determination process (S416). The determination process of S416 determines whether the wake-up flag sequentially saved in the memory unit 304 of the smartphone 1 is 1, indicating enabled, or 0, indicating disabled. If the determination process of S416 determines that the wake-up flag is 1, indicating enabled, the audio signal acquired in the process of S402 is determined to be speaker sound 31 from the television 30. Therefore, the acquired wake-up word is invalidated, and the process proceeds to S407. If the determination process of S416 determines that the wake-up flag is not 1 (the wake-up flag is 0), the audio signal acquired in the process of S402 is determined to be the user's 10's voice 11, and the process proceeds to S405. This concludes the description of the main processing procedure in the smartphone 1.
[0126] In this embodiment, when the speaker sound 31 of the television 30 is the activation word, an activation flag of 1 is transmitted as an identification signal, and the smartphone 1 receives the activation flag of 1, thereby determining and identifying that the activation word is the speaker sound 31. This embodiment has the same effect as the previous embodiment, and can disable the speaker sound 31 output from the television 30, which is an audio output device. Needless to say, this embodiment can also be applied to the object sound (embodiment 3) described above.
[0127] The following describes a sixth embodiment of the present invention. The basic hardware and software configurations of the sixth embodiment are the same as those of the previously described embodiments, and the following mainly describes the differences between this embodiment (sixth embodiment) and the previously described embodiments, and omits explanations of common parts as much as possible to avoid duplication. In the previously described embodiments, the speaker sound 31 was determined using text data, but in this embodiment, the determination is made using an audio signal.
[0128] Fig. 16 is a flowchart showing a main processing procedure in the television 30 of this embodiment. The processing procedure in Fig. 16 will be explained with reference to the system configuration diagram in Fig. 2. The processing procedure in the television 30 of this embodiment is realized by a program of the control unit 302 stored in the storage unit 304, utilizing an external speaker connection via the short-range wireless communication device 363 of the communication unit 306. In other words, in this embodiment, the short-range wireless communication device 363 functions as an information transmission unit that transmits information to the outside. The processing procedure in Fig. 16 will be explained below.
[0129] When the process starts (S451), initial configuration is first performed (S452). In the initial configuration (S452), the device to which the short-range wireless communication device 363 is connected is set. In this embodiment, the smartphone 1 is selected as the device to which the short-range wireless communication device 363 is connected and set as an external speaker. Next, an audio information acquisition process (S453) is performed to acquire audio information. Specifically, playback of a digital television broadcast received via the tuner 305 or playback of an audio / video file stored in the storage unit 304 is started. Next, the audio processor 383 extracts an audio signal from the audio information acquired in the audio information acquisition process of S453 (S454). Next, the audio signal extracted in the process of S454 is output from the speaker 382 of the television 30 (S455). Similarly, the audio signal extracted in the process of S454 is output via the short-range wireless communication device 363 (S456), and the audio signal is transmitted to the smartphone 1.
[0130] Next, it is determined whether the audio information has ended (S457). If it is determined in the determination process of S457 that the audio information has not ended, the process returns to the audio information acquisition process of S432 (S453) and continues. If it is determined in the determination process of S457 that the audio information has ended, the main processing procedure in the television 30 ends (S437). This concludes the description of the processing procedure in FIG. 16.
[0131] In the flowchart of Figure 16, for example, all audio signals extracted from digital television broadcasts are transmitted to smartphone 1, but as described in Example 1, for example, only the audio signals of the subtitle portion may be extracted and transmitted to smartphone 1.
[0132] FIG. 17 is a flowchart showing the main processing procedure in the smartphone 1 of this embodiment. The processing procedure in FIG. 17 will be explained with reference to the system configuration diagram in FIG. 5. The processing procedure in the smartphone 1 of this embodiment is realized by the program of the control unit 2 stored in the storage unit 4 using the short-range wireless communication device 63 of the communication unit 6. In other words, in this embodiment, the short-range wireless communication device 63 functions as an information receiving unit that receives information from the outside. The processing procedure in FIG. 17 will be explained below.
[0133] When the process starts (S421), first, initial setting is performed (S422). In the initial setting (S422), a connection destination device of the short-range wireless communication device 63 is set. In this embodiment, the television 30 is selected and set as the connection destination device of the short-range wireless communication device 63. Next, an audio signal is acquired from the microphone 81 of the audio unit 8 (S423). Similarly, an audio signal from short-range wireless communication to an external speaker is acquired via the short-range wireless communication device 63 (S424).
[0134] Next, the audio signal from the microphone 81 acquired in S423 is compared with the audio signal via near-field wireless communication acquired in S424 (S425). Next, if it is determined in the determination process of S425 that the audio signal from the microphone 81 acquired in S423 matches the audio signal for the external speaker acquired in S424, it is determined that the audio signal is from the television 30, and the audio signal is invalidated. Then, the process proceeds to the determination process of S427. If it is determined in the determination process of S425 that the audio signal from the microphone 81 acquired in S423 does not match the audio signal for the external speaker acquired in S424, it is determined that the audio signal is from the user 10, and the process continues. The subsequent determination process of S405 and the process of S406 have been described in the above-mentioned first embodiment, and therefore will not be described here.
[0135] That is, the audio signal (third audio signal) from the short-range wireless communication is compared with the audio signal (picked audio signal) from the microphone 81, and if it is determined that they match, it is determined that the picked-up audio signal is the second audio signal output by the audio output device. If it is determined that the picked-up audio signal is the second audio signal, the processes of S405 and S406 are not executed, and the voice recognition process is not performed.
[0136] Next, it is determined whether the audio signal has ended (S427). If it is determined in the determination process of S427 that the audio signal has not ended, the process returns to the process of acquiring the audio signal from the microphone 81 in S423 and continues the main process. If it is determined in the determination process of S427 that the audio signal has ended, the main processing procedure in the smartphone 1 is terminated (S428). This concludes the description of the processing procedure in FIG. 17.
[0137] In this embodiment, the audio signal output from the speaker 382 of the television 30 is compared with the audio signal of an external speaker output via short-range wireless communication, and if the two audio signals match, it is determined that the audio signal is the audio signal (speaker sound 31) of the television 30. In this embodiment, the effect is the same as in the previous embodiment, and the speaker sound 31 output from the television 30, which is an audio output device, can be disabled.
[0138] The following describes an eighth embodiment of the present invention. The basic hardware and software configurations of the seventh embodiment are the same as those of the above-described embodiments, and the following mainly describes the differences between this embodiment (seventh embodiment) and the above-described embodiments, and omits explanations of common parts as much as possible to avoid duplication.
[0139] In the above-described embodiment, the television 30 as an audio output device and the smartphone 1 as a voice recognition device function as a system in which the audio signal generated by the audio output device (sound output from the speaker) is invalidated and the audio signal uttered by the user (real voice) is valid. In this embodiment, the smartphone 1 alone invalidates the audio signal generated by the audio output device (sound output from the speaker) and validates the audio signal uttered by the user (real voice).
[0140] In this embodiment, a 3D microphone is used to detect the position of a sound source. The 3D microphone in this embodiment has four unidirectional microphones arranged in different directional directions (front upper left, front lower right, rear lower left, rear upper right), and detects the direction of a sound source based on the signal level (sound pressure) and phase difference of each microphone.
[0141] FIG. 18 is an external view showing an example of a smartphone 1 used in this embodiment. The upper part of FIG. 18 is a diagram showing the front side of the smartphone 1, and the lower part of FIG. 18 is a diagram showing the rear side of the smartphone 1. In the upper part of FIG. 18, the smartphone 1 has a smartphone front surface 183 having a display screen 181 configured with a touch panel and a front camera (also called an in-camera) 182 for taking selfies. In the lower part of FIG. 18, the smartphone 1 has a smartphone back surface 185 having a rear camera (also called an out-camera or simply a camera) 184. The smartphone front surface 183 and the smartphone back surface 185 are connected with a gap therebetween, thereby forming a smartphone side surface that becomes part of the housing.
[0142] An earphone microphone terminal (not shown) is provided on the side of the smartphone 1 facing the camera lens on the front (the left side in FIG. 18 ). An external connection terminal (not shown) is provided on the side of the smartphone 1 facing the opposite side of the camera lens on the front (the right side in FIG. 18 ). A power key and volume keys (not shown) are provided on the upper side of the smartphone 1 facing the camera lens on the front. A card tray (not shown) is provided on the lower side of the smartphone 1 facing the camera lens on the front.
[0143] The smartphone 1 of this embodiment also has a short-range wireless communication function, which is used for wireless communication to acquire position (direction and distance) information and exchange audio data with the television 30. The transmitting and receiving antennas supporting this short-range wireless communication are arranged as follows: a left antenna 186 at the leftmost end of the top edge, a right antenna 187 at the rightmost end, and a central antenna 188 at the center of the opposite side (bottom side) when viewed from the front of the smartphone 1. These three antennas arranged in different positions are used to detect the position of the short-range wireless communication device. That is, in this embodiment, the left antenna 186, the right antenna 187, and the central antenna 188 function as information receiving units that receive information from the outside.
[0144] In particular, the antennas that are components of the short-range wireless communication device 63 are important parts related to calculating the position of the television 30 having a short-range wireless communication function in the present invention. The arrangement of the three short-range wireless communication antennas, consisting of the left antenna 186, right antenna 187, and center antenna 188 shown in Figure 18, is an arrangement that takes into consideration the detection of the position of the television 30.
[0145] In this embodiment, the smartphone 1 is set as the short-range wireless communication device of the television 30, and the television 30 is set as the short-range wireless communication device of the smartphone 1. Fig. 19 is a schematic diagram showing the calculation principle for calculating the relative positions of the smartphone 1 and the television 30.
[0146] The smartphone 1 of this embodiment has three antennas: a left antenna 186, a central antenna 188, and a right antenna 187. The central antenna 188 is located at the center of the opposite side of the left antenna 186 and the right antenna 187. In other words, the distance from the position of the central antenna 188 to the position of the left antenna 186 is equal to the distance from the position of the central antenna 188 to the position of the right antenna 187. In this embodiment, direction detection is performed using AoA (Angle of Arrival), a method of receiving radio waves with multiple antennas and detecting the direction of arrival of the radio waves based on the phase difference of the received radio waves. The position 754 where the direction 751 from the left antenna 186, the direction 752 from the right antenna 187, and the direction 753 from the central antenna 188 intersect is the desired relative position to the television 30. This concludes the explanation of the principle of relative position calculation.
[0147] If the direction of the sound source determined by the 3D microphone matches the relative position of the television 30 determined by AoA, it can be determined that the sound source is the speaker sound 31 emitted by the television 30. By not performing voice recognition on the speaker sound 31 emitted by the television 30, the speaker sound 31 emitted by the television 30 can be disabled.
[0148] Furthermore, a method using directional ultrasound without using a 3D microphone is also possible, and it goes without saying that the effects of this embodiment can be obtained even when directional ultrasound is employed. The smartphone 1, which is a voice recognition device, can detect the direction of the television 30, which is an audio output device, by detecting directional ultrasound output from the television 30. Even when directional ultrasound is used, the direction of the sound source can be determined, and if the relative position with the television 30 matches, it can be determined that the sound source is the speaker sound 31 emitted by the television 30. By not performing voice recognition on the speaker sound 31 emitted by the television 30, the speaker sound 31 emitted by the television 30 can be disabled.
[0149] Although examples of embodiments of the present invention have been described above using Examples 1 to 7, it goes without saying that the present invention can be applied not only to Bluetooth (registered trademark) communication but also to UWB (Ultra Wide Band) communication as short-range wireless communication. Furthermore, in recent years, there have been examples in which Bluetooth (registered trademark) signals are used as remote control signals. It goes without saying that the present invention can be applied by employing Bluetooth (registered trademark) signals instead of infrared rays. It also goes without saying that the present invention can be applied to cases in which the user's voice is output as speaker sound from an audio output device.
[0150] Furthermore, when the voice recognition device is a smartphone 1, the following embodiment may be used. For example, when the smartphone 1 and the user 10 are close to each other, the picked-up voice signal picked up by the microphone 81 is determined to be a first voice signal (human voice 11). This determination may be made, for example, when the distance L between the smartphone 1 and the user 10 is equal to or less than a predetermined threshold, as shown in FIG. 20 . This threshold may be set to a distance such that the user 10 is holding the smartphone 1 in his / her hand and speaking to the smartphone 1. The distance L may be measured, for example, by recognizing a face from an image captured by the front camera 182 and estimating the distance to the face. Furthermore, if the smartphone 1 is equipped with a distance measurement sensor, the distance may be measured using the distance measurement sensor. Alternatively, the distance may be estimated based on the level of the voice signal picked up by the microphone 81.
[0151] 20, it may also be determined whether the front side of the smartphone 1 is facing the direction of the user 10. This determination is possible by recognizing the face from an image captured by the front camera 182. This is because if the front side of the smartphone 1 is facing the direction of the user 10, it can be assumed that the user 10 is talking to the smartphone 1.
[0152] Furthermore, the configuration for realizing the technology of the present invention is not limited to the above-described embodiment, and various modifications are possible. For example, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. All of these fall within the scope of the present invention. Furthermore, the numerical values, messages, etc. appearing in the text and figures are merely examples, and the effects of the present invention will not be impaired even if different ones are used. Some or all of the functions of the present invention described above may be implemented in hardware, for example, by designing an integrated circuit. They may also be implemented in software by a microprocessor unit or the like interpreting and executing a program that realizes each function. Hardware and software may also be used together. The software may be stored in the memory unit 4 of the smartphone 1 before product shipment. It may also be obtained from various server devices on the Internet after product shipment. The software may also be provided on a memory card, optical disc, or the like. The control lines and information lines shown in the figures are those considered necessary for explanation and do not necessarily represent all control lines and information lines on the product. In reality, it is acceptable to consider almost all components to be interconnected.
[0153] 1...smartphone, 2...control unit, 3...system bus, 4...memory unit, 5...sensor unit, 6...communication unit, 7...video unit, 8...audio unit, 9...operation input unit, 10...user, 11...human voice, 30...television, 31...speaker sound, 61...LAN communication device, 63...short-range wireless communication device, 64...infrared communication device, 83...audio processor, 304...memory unit, 361...LAN communication device, 363...short-range wireless communication device, 364...infrared communication device, 383...audio processor.
Claims
1. An audio output device comprising: an audio signal acquisition unit that acquires an audio signal; an audio signal determination unit that determines a speech portion that can be recognized as having a meaning as language from the audio signal acquired by the audio signal acquisition unit; an audio signal processing unit that generates an identification signal for the speech portion of the audio signal acquired by the audio signal acquisition unit to identify the audio signal as speech spoken by a user; and an output unit that outputs the identification signal.
2. An audio output device according to claim 1, wherein the audio signal processing unit adds the identification signal to a portion of the audio signal that is determined to be the speech portion, and the output unit is configured with a speaker that outputs the audio signal to which the identification signal has been added as audio.
3. An audio output device according to claim 2, wherein the audio signal acquisition unit acquires subtitle information as additional information related to the audio signal, and the audio signal determination unit determines that the section in which the subtitle information is valid is the speech portion.
4. An audio output device according to claim 2, wherein the audio signal acquisition unit acquires object-based audio information including acoustic metadata as additional information related to the audio signal, and the audio signal determination unit determines that an audio object whose acoustic metadata is dialogue is the speech portion.
5. An audio output device according to claim 2, wherein the audio signal processing unit forms a suppression section as the identification signal in a specific frequency band of the audio signal acquired by the audio signal acquisition unit.
6. An audio output device according to claim 2, wherein the audio signal processing unit adds a digital watermark signal as the identification signal to the audio signal acquired by the audio signal acquisition unit.
7. An audio output device according to claim 2, wherein the audio signal processing unit adds an ultrasonic signal as the identification signal to the audio signal acquired by the audio signal acquisition unit.
8. An audio output device according to claim 1, wherein the output unit is configured as an information transmission unit that transmits information to the outside, the audio signal acquisition unit acquires subtitle information as additional information related to the audio signal, the audio signal processing unit extracts text information as the identification signal from the subtitle information acquired as the additional information, and the information transmission unit transmits the text information to a voice recognition device.
9. An audio output device as claimed in claim 1, wherein the output unit is configured as an information transmission unit that transmits information to the outside, the audio signal acquisition unit acquires subtitle information as additional information related to the audio signal, the audio signal processing unit extracts text information from the subtitle information acquired as the additional information, and when the text information is a startup word for a voice recognition device, the information transmission unit transmits startup word information indicating that the startup word has been detected to the voice recognition device.
10. A voice output device according to claim 1, wherein the output unit is configured as an information transmission unit that transmits information to the outside, the voice signal processing unit extracts a portion of the voice signal that is determined to be the speech portion as the identification signal, and the information transmission unit transmits the identification signal extracted by the voice signal processing unit to a voice recognition device.
11. A voice recognition device that performs voice recognition on a voice signal, comprising: a microphone that collects external voice; a voice signal discrimination unit that discriminates whether a picked-up voice signal based on voice collected by the microphone is a first voice signal spoken by a user or a second voice signal output by a voice output device; and a voice recognition unit that performs voice recognition processing on the picked-up voice signal and outputs the result of the voice recognition processing, wherein if the voice signal discrimination unit discriminates that the picked-up voice signal is the first voice signal, the voice recognition unit performs voice recognition processing on the picked-up voice signal, and if the voice signal discrimination unit discriminates that the picked-up voice signal is the second voice signal, the voice recognition unit does not perform voice recognition on the picked-up voice signal.
12. A voice recognition device according to claim 11, wherein the voice signal discrimination unit discriminates the picked-up voice signal as the second voice signal when it detects an identification signal added to the voice signal.
13. A speech recognition device according to claim 11, further comprising an information receiving unit that receives first text information from the speech output device, wherein the speech recognition unit outputs second text information obtained by converting the picked-up speech signal into text information, and the speech signal discrimination unit compares the first text information with the second text information, and if it determines that they match, discriminates the picked-up speech signal as the second speech signal.
14. A voice recognition device as claimed in claim 11, further comprising an information receiving unit that receives activation word information indicating that an activation word for activating the voice recognition device has been detected in the voice output device, and when the information receiving unit receives the activation word information, the voice signal determining unit determines that the picked-up voice signal is the second voice signal output by the voice output device.
15. A speech recognition device as claimed in claim 11, further comprising an information receiving unit that receives a third speech signal from the speech output device, wherein the speech signal determining unit compares the third speech signal received by the information receiving unit with the picked-up speech signal, and if it determines that they match, determines that the picked-up speech signal is the second speech signal output by the speech output device.
16. A voice recognition device as claimed in claim 11, further comprising an information receiving unit that receives information from the voice output device, wherein the information receiving unit calculates the relative position of the voice output device based on the received information, and the voice signal discrimination unit detects the direction of the sound source of the sound output from the voice output device based on the picked-up voice signal, compares the direction of the sound source with the direction indicated by the relative position, and if it is determined that they match, discriminates that the picked-up voice signal is the second voice signal output by the voice output device.
17. A speech recognition system comprising: a speech output device according to claim 2; and a speech recognition device according to claim 12.
18. A speech recognition system comprising: a speech output device according to claim 8; and a speech recognition device according to claim 13.
19. A speech recognition system comprising: a speech output device according to claim 9; and a speech recognition device according to claim 14.
20. A speech recognition system comprising: a speech output device according to claim 10; and a speech recognition device according to claim 15.
Citation Information
Patent Citations
Device and system for speech recognition, and interactive device
JP2000227799A
Video audio reproducing apparatus
JP2008160232A
Control program, storage medium, portable communication equipment, program-related information provision device, and program-related information display method
JP2017060059A
Voice recognition device and voice recognition method
JP2019184809A
Smart speaker, processing method, and processing program
JP2022113569A