Electronic device, information processing system, electronic device control method, and program
The electronic device uses imperceptible audio watermarks to secure process execution, addressing voice authentication vulnerabilities by ensuring valid watermarks and user alignment, thus enhancing security.
Patent Information
- Application Number
- JP2022063348
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-06
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2042-04-06
AI Technical Summary
Existing electronic devices that rely on voice authentication for user input face security vulnerabilities as third parties can easily discern the password, compromising the device's security.
An electronic device equipped with a sound acquisition system, determination system, and control system to set a privileged mode using an imperceptible digital watermark in audio, ensuring secure execution of specific processes only when valid audio watermarks are detected and aligned with the user's direction.
Enhances security by allowing secure execution of processes through imperceptible audio watermarks, preventing unauthorized access and maintaining device integrity.
Smart Images

Figure 0007754767000001 
Figure 0007754767000002 
Figure 0007754767000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an electronic device, an information processing system, a control method for an electronic device, and a program. [Background technology]
[0002] Conventionally, a user can input a specific password string into an electronic device (such as a personal computer or smartphone) and, if granted special authority (privilege), have the electronic device perform a specific process (such as deleting data).However, unlike personal computers and smartphones, smart speakers and robots may not have a keyboard or touch panel for inputting strings.
[0003] Therefore, Patent Document 1 discloses a user authentication device that authenticates a user by having the user speak a password. This enables user authentication even in electronic devices that do not have a keyboard or touch panel for inputting characters. As a result, the user can have the electronic device perform specific processing. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2020-64689 [Patent Document 2] Japanese Patent Application Laid-Open No. 2003-5790 [Patent Document 3] Patent Publication No. 2021-5871 Summary of the Invention [Problem to be solved by the invention]
[0005] However, if authentication is performed based on the user's speech, a third party who hears the speech can easily figure out the password required for authentication, which reduces the security of the electronic device.
[0006] Therefore, an object of the present invention is to put an electronic device into a state in which it can execute a specific process in response to a voice, and to prevent a decrease in security of the electronic device. [Means for solving the problem]
[0007] In order to achieve the above object, the present invention employs the following configuration.
[0008] That is, an electronic device according to one aspect of the present invention is an electronic device characterized by having: a sound acquisition means for acquiring sound generated around the electronic device; an execution means for executing a process according to a user's speech when the sound acquisition means acquires the sound; a determination means for determining whether the watermark is valid when the sound acquisition means acquires the sound; and a control means for setting the electronic device to a privileged mode in which the execution means can execute a specific process if the watermark is valid. The specific process is, for example, a process that is desirable to be executed only by a specific user. Specifically, the specific process is, for example, a process related to the security of the electronic device, such as changing (initializing) a password or changing security settings. The electronic device may be, for example, a smart speaker or a robot.
[0009] According to this configuration, the electronic device sets the privileged mode using audio containing a digital watermark that cannot be perceived by the human ear. Therefore, even if a third party listens to the audio, they cannot understand the sound used to set the privileged mode. In addition, when audio containing a digital watermark is played, Even if the voice is being used, it is not easy for a third party to recognize that it is a voice for setting the electronic device to the privileged mode. In other words, a third party cannot understand how the process for setting the electronic device to the privileged mode is being performed. This allows the electronic device to be set to the privileged mode with high security.
[0010] In the above electronic device, the execution means may execute the specific process if the user's utterance instructing the specific process has been completed during the period in which the electronic device is set to the privileged mode. This allows the specific process to be executed in response to the user's utterance completed during the period in which the privileged mode is set. Therefore, the user can cause the electronic device to execute the specific process with just two steps: playing audio containing an audio watermark and uttering an utterance instructing the specific process. Therefore, the user can easily cause the electronic device to execute the specific process.
[0011] In the electronic device, the execution means may prohibit the execution of the specific process if the angle between the direction from which the voice containing the digital watermark was emitted toward the electronic device and the direction from which the user who made the utterance is located toward the electronic device is greater than a predetermined angle. This reduces the possibility that the electronic device will execute the specific process when, for example, a third party other than the user who generated the voice containing the digital watermark utters an utterance instructing the specific process. This improves the security of the electronic device.
[0012] In the electronic device, the execution means may prohibit the execution of the specific process if the user who made the utterance is located in a direction different from the direction from which the voice containing the digital watermark was emitted. This further reduces the possibility that the electronic device will execute the specific process if, for example, a third party other than the user who made the voice containing the digital watermark utters an utterance instructing the specific process. This further improves the security of the electronic device.
[0013] In the electronic device, the control means may set the electronic device in the privileged mode for a predetermined time after the audio acquisition means has acquired the valid digital watermark, and release the privileged mode after the predetermined time has elapsed. This limits the period during which the privileged mode is set, thereby reducing the possibility that a third party will execute a specific process by speaking an utterance instructing the process, thereby improving the security of the electronic device.
[0014] In the above electronic device, the control means may: 1) if the user is speaking during a period in which the electronic device is set to the privileged mode and the speech has not ended, prompt the user to speak again after the period ends, and 2) set the electronic device to the privileged mode for a period of time after the prompt to speak again that is longer than the length of the previous speech. This allows the user to again speak a command to execute a specific process and have the electronic device execute the specific process, even if the privileged mode ends during the period in which the user is speaking. This improves the user's convenience when using the electronic device.
[0015] In the above electronic device, the electronic watermark may include authentication information and information indicating a validity period, and the determination means may determine that the electronic watermark is valid if the authentication information is specified information and the current time is included in the validity period.
[0016] In the electronic device, the execution means may execute the predetermined process when the user utters an utterance instructing a predetermined process other than the specific process, regardless of whether the electronic device is set to the privileged mode. According to this, the user can always cause the electronic device to execute a normal process other than the specific process.
[0017] In the electronic device, the control means may cancel the privileged mode when the user finishes speaking during the period in which the electronic device is set to the privileged mode. This further shortens the period in which the electronic device is set to the privileged mode, thereby reducing the possibility that a third party may force the electronic device to execute a specific process, thereby improving the security of the electronic device.
[0018] In the electronic device, the digital watermark may be expressed by audio having a frequency not included in the frequency range of 20 Hz to 20,000 Hz.
[0019] An information processing system may be provided that includes the electronic device and an audio output device that outputs audio containing the digital watermark, whereby the information processing system can set the electronic device to a privileged mode based on the audio containing the digital watermark output by the audio output device.
[0020] The present invention can be understood as a control method, authentication device, authentication method, information processing device, privilege granting device, information processing method, or privilege granting method for an electronic device that includes at least some of the functions and processes described above. The present invention can also be understood as a program that causes a computer to execute each means of the electronic device (each step of the control method), or a storage medium that non-temporarily stores the program. [Effects of the Invention]
[0021] According to the present invention, it is possible to put an electronic device into a state in which a specific process can be executed in response to a voice, and to prevent a decrease in security of the electronic device. [Brief explanation of the drawings]
[0022] [Figure 1] FIG. 1 is a diagram illustrating an information processing system according to the first embodiment. [Figure 2] FIG. 2 is a diagram showing the internal configuration of each component of the information processing system according to the first embodiment. [Figure 3]FIG. 3 is a flowchart of the execution control process according to the first embodiment. [Figure 4] FIG. 4 is a time chart of the execution control process according to the first embodiment. [Figure 5] FIG. 5 is a flowchart of the music playback process according to the first embodiment. [Figure 6] FIG. 6 is a flowchart of a process during a privileged period according to the second embodiment. [Figure 7] FIG. 7 is a diagram illustrating detection of the direction of sound generation according to the third embodiment. [Figure 8] FIG. 8 is a flowchart of the execution control process according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0023] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited to the described embodiments. Furthermore, not all of the components described in the embodiments are necessarily essential to the present invention.
[0024] <Embodiment 1> An information processing system 1 according to a first embodiment will be described with reference to FIG. 1. The information processing system 1 includes a smart speaker 10, a server 20, and a smartphone 30. Based on a voice emitted from the smartphone 30, the information processing system 1 sets the smart speaker 10 to a mode (privileged mode) in which a specific process can be executed. The specific process may be any process that has been set in advance. The specific process may be, for example, changing (initializing) a password, changing security settings, or deleting data.
[0025] The smart speaker 10 is an electronic device that acquires sounds around itself and executes processing according to the acquired sounds. If the audio contains a digital watermark and the digital watermark is determined to be valid, the smart speaker 10 enters privileged mode. The smart speaker 10 can also communicate with the server 20 via the network 40.
[0026] Here, an audio watermark is information embedded in audio and controlled so that it cannot be perceived by humans. The smart speaker 10 can separate audio containing an audio watermark into the audio watermark and other parts according to frequency and phase. For example, if the audio watermark is expressed using audio with a frequency that is not included in the frequency range of audio that humans can perceive (20 Hz to 20,000 Hz), the smart speaker 10 can separate the audio watermark from other parts according to frequency. In the first embodiment, the audio watermark includes authentication information (key information) and information indicating the validity period of the authentication information. The authentication information corresponds to a password generally used for user authentication.
[0027] The server 20 stores information (digital watermark information) for the smart speaker 10 to determine the validity of the audio digital watermark. The server 20 transmits the digital watermark information to the smart speaker 10 in response to a request from the smart speaker 10. The server 20 also transmits the digital watermark information to the smartphone 30 so that the smartphone 30 can play music containing the audio digital watermark. The digital watermark information includes authentication information and information indicating the validity period of the authentication information.
[0028] The smartphone 30 is an electronic device (audio output device) that plays music. When a user wants to set the smart speaker 10 to privileged mode, the user causes the smartphone 30 to play music containing an audio watermark. The smartphone 30 acquires watermark information from the server 20, embeds the audio watermark in any music, and plays the music. A known echo diffusion method (a method in which an echo with a delay time imperceptible to humans is applied to the original music and the delay time is used as additional data) can be used to embed an audio watermark in music. The audio watermark can also be embedded using other methods besides the echo diffusion method, such as a known periodic phase modulation method or a known spread spectrum method.
[0029] (Configuration of Smart Speaker 10) The internal configuration of the smart speaker 10 will be described with reference to Fig. 2. The smart speaker 10 has a voice acquisition unit 101, a voice separation unit 102, a voice recognition unit 103, a determination unit 104, a mode control unit 105, an execution unit 106, an information update unit 107, a voice output unit 108, and a storage unit 109.
[0030] The voice acquisition unit 101 acquires voices emitted around the smart speaker 10. The voice acquisition unit 101 has, for example, one or more microphones (such as an array microphone).
[0031] The audio separation unit 102 separates the audio acquired by the audio acquisition unit 101 (hereinafter referred to as "acquired audio"). Specifically, first, the audio separation unit 102 separates the acquired audio into audio sources (users and devices). For this audio separation, for example, a technique for separating the audio based on the independence of the acquired audio signals can be used, as described in Patent Document 2. For audio separation, a technique for detecting the audio source direction based on the difference in arrival time of the audio reaching the multiple microphones of the audio acquisition unit 101 and separating the audio for each audio source direction can be used. Furthermore, if the multiple microphones of the audio acquisition unit 101 acquire audio from different directions (beamforming), the audio separation unit 102 may treat the audio acquired by each of the multiple microphones as audio obtained by separating the acquired audio. For example, if the audio acquisition unit 101 acquires the voices of two users and the voice of one smartphone 30, the audio separation unit 102 may treat the audio acquired by each of the multiple microphones as audio obtained by separating the acquired audio. The separation unit 102 separates the acquired sound into three sounds.
[0032] Furthermore, the audio separation unit 102 determines whether any of the separated audio pieces contains an audio watermark. For example, if an audio watermark is embedded in music using an echo diffusion method, the audio separation unit 102 performs an audio watermark extraction process using a known echo diffusion method on each of the separated audio pieces (or only on audio pieces determined to correspond to music). Then, if the audio watermark extraction process succeeds in extracting an audio watermark, the audio separation unit 102 determines that any of the separated audio pieces contains an audio watermark.
[0033] Specifically, extraction of an audio watermark using the echo diffusion method can be achieved by the method described in Patent Document 3. For example, the audio separation unit 102 applies a window function to music containing an audio watermark, performs an FFT (Fast Fourier Transform), takes a logarithm, and then performs an inverse FFT. This allows the audio separation unit 102 to calculate a so-called cepstrum. The audio separation unit 102 can then extract the audio watermark by calculating the cross-correlation between the cepstrum and the echo component for each window length.
[0034] The voice recognition unit 103 recognizes the content of the user's voice from the voice separated by the voice separation unit 102. That is, the voice recognition unit 103 recognizes the content of the user's utterance from the acquired voice.
[0035] Furthermore, the voice recognition unit 103 determines whether the content of the user's utterance is an instruction to perform a specific process. In the simplest example, if the specific process is "initialize the login password," and the utterance contains the words "password" and "initialize," the voice recognition unit 103 can determine that the content of the user's utterance is an instruction to perform a specific process. Note that the technology for recognizing the content of the user's utterance can be the same as the technology used in general smart speakers, and therefore a detailed description thereof will be omitted in this specification.
[0036] When the audio separation unit 102 acquires an audio watermark, the determination unit 104 determines whether the audio watermark is valid (the validity of the audio watermark). Specifically, the determination unit 104 determines whether the authentication information included in the audio watermark corresponds to the authentication information indicated by the watermark information previously stored in the storage unit 109 (for example, whether the two pieces of authentication information are identical or are shifted by a predetermined number of bits). If the two pieces of authentication information correspond to each other, the determination unit 104 determines whether the current time is included in the validity period included in the audio watermark. If the determination unit 104 determines that the two pieces of authentication information correspond to each other and that the current time is included in the validity period included in the audio watermark, it determines that the audio watermark is valid. If the determination unit 104 determines that the two pieces of authentication information do not correspond to each other or that the current time is not included in the validity period included in the audio watermark, it determines that the audio watermark is invalid.
[0037] The mode control unit 105 controls whether to set the smart speaker 10 to the privileged mode depending on whether the audio digital watermark is valid. If the audio digital watermark is valid, the mode control unit 105 sets the smart speaker 10 to the privileged mode. If the audio digital watermark is not valid, the mode control unit 105 does not set the smart speaker 10 to the privileged mode.
[0038] The execution unit 106 executes processing according to the content of the user's utterance recognized by the voice recognition unit 103. The execution unit 106 includes a process execution unit 161 and a privileged process execution unit 162.
[0039] The process execution unit 161 executes processes other than the specific processes (hereinafter referred to as "normal processes"). For example, if the user's utterance instructs "to increase the volume of the smart speaker 10," the processing execution unit 161 increases the volume of the sound emitted from the sound output unit 108. The processing execution unit 161 can execute normal processing regardless of whether the smart speaker 10 is set to the privileged mode or not.
[0040] The privileged process execution unit 162 executes a specific process. Here, the privileged process execution unit 162 executes the specific process only when the smart speaker 10 is set to the privileged mode. In other words, the privileged process execution unit 162 does not execute the specific process unless the smart speaker 10 is set to the privileged mode. In other words, it can be said that "if the smart speaker 10 is not set to the privileged mode, the mode setting unit 105 prohibits the execution unit 106 from executing the specific process."
[0041] The information updating unit 107 updates the digital watermark information stored in the storage unit 109. The information updating unit 107 periodically accesses the server 20 and acquires the latest digital watermark information from the server 20. Then, the information updating unit 107 replaces the digital watermark information stored in the storage unit 109 with the digital watermark information acquired from the server 20. Note that the information updating unit 107 may acquire the latest digital watermark information from the server 20 when it is determined that the audio acquired by the audio acquiring unit 101 includes an audio digital watermark.
[0042] The audio output unit 108 outputs audio (emits audio). The audio output unit 108 includes a speaker. For example, the audio output unit 108 outputs audio to report to the user the result of the processing executed by the execution unit 106.
[0043] The storage unit 109 stores the digital watermark information updated (output) by the information update unit 107. The storage unit 109 may also store a program for controlling each functional unit of the smart speaker 10.
[0044] (Server 20 configuration) The internal configuration of the server 20 will be described with reference to Fig. 2. The server 20 includes an information update unit 201, an information transmission unit 202, and a storage unit 203.
[0045] The information update unit 201 periodically updates the digital watermark information stored in the storage unit 203. Specifically, the information update unit 201 updates the authentication information included in the digital watermark information to new arbitrary information, and updates the validity period information included in the digital watermark information to a new period. The validity period should be as short as possible, for example, 30 minutes from the time when the information update unit 201 updates it.
[0046] In response to a request from the smart speaker 10, the information transmitting unit 202 transmits the digital watermark information stored in the memory unit 203 to the smart speaker 10. In addition, in response to a request from the smartphone 30, the information transmitting unit 202 transmits the digital watermark information stored in the memory unit 203 to the smartphone 30.
[0047] The storage unit 203 stores the digital watermark information.
[0048] (Configuration of Smartphone 30) The internal configuration of the smartphone 30 will be described with reference to Fig. 2. The smartphone 30 includes an information acquisition unit 301, a music generation unit 302, an audio output unit 303, and a storage unit 304.
[0049] The information acquisition unit 301 acquires digital watermark information by making a request to the server 20. The information acquisition unit 301 stores the acquired digital watermark information in the storage unit 304.
[0050] The music generation unit 302 generates music in which an audio digital watermark based on the digital watermark information is embedded (hereinafter referred to as "watermarked music"). The music generation unit 302 stores the watermarked music information (music information) in the storage unit 304. Here, the music in which the audio digital watermark is embedded may be any music. For example, it may be music specified by the user.
[0051] Furthermore, the audio digital watermark is information that includes, for example, information included in the digital watermark information (authentication information and information indicating the validity period). The audio digital watermark may also include authentication information (hereinafter referred to as "encryption information") obtained by encrypting the authentication information included in the digital watermark information according to a certain rule, and information indicating the validity period. If the audio digital watermark includes encryption information, the smart speaker 10 needs to have information for decrypting the encryption information. For this reason, for example, the server 20 may transmit key information for encrypting the authentication information to the smartphone 30, and key information for decrypting the encryption information to the smart speaker 10. In this case, upon acquiring the audio digital watermark, the smart speaker 10 decrypts the encryption information included in the audio digital watermark to extract the authentication information. The determination unit 104 then determines the validity of the audio digital watermark based on whether the extracted authentication information matches the authentication information included in the digital watermark information.
[0052] The audio output unit 303 acquires information about the watermarked music (music information) from the storage unit 304. Then, the audio output unit 303 plays back the watermarked music in response to a user operation.
[0053] The storage unit 304 stores music information and digital watermark information.
[0054] The smart speaker 10, the server 20, and the smartphone 30 can be configured as a computer equipped with a CPU (processor; control device) or memory and storage. In this case, the configuration shown in FIG. 2 is realized by loading a program stored in the storage into the memory and having the CPU execute the program. Alternatively, all or part of the configuration of the smart speaker 10, the server 20, and the smartphone 30 may be configured using an ASIC, FPGA, or the like. Alternatively, all or part of the configuration may be realized by cloud computing or distributed computing.
[0055] (About execution control processing) The execution control process for controlling whether or not a specific process is executed will be described with reference to the flowchart in Figure 3. The process of this flowchart starts when the smart speaker 10 is turned on. The process of this flowchart is realized, for example, by the CPU of the smart speaker 10 executing a program stored in storage.
[0056] In step S1001, the audio separation unit 102 determines whether the audio acquisition unit 101 has acquired watermarked music (whether watermarked music is playing around the smart speaker 10). If it is determined that watermarked music has been acquired, the process proceeds to step S1002. If it is determined that watermarked music has not been acquired, the process proceeds to step S1004.
[0057] In step S1002, the determination unit 104 determines whether the audio digital watermark included in the watermarked music is valid. Specifically, as described above, the determination unit 104 determines whether the audio digital watermark is valid based on the audio digital watermark and the digital watermark information stored in the storage unit 109. If it is determined that the audio digital watermark is valid, the process proceeds to step S1003. If it is determined that the audio digital watermark is not valid (invalid), the process proceeds to step S1004.
[0058] The processing of S1002 may be performed using the server 20. Specifically, when the audio separation unit 102 acquires (extracts) an audio digital watermark, it transmits the audio digital watermark to the server 20. The server 20 determines the validity of the audio digital watermark based on the digital watermark information stored in the storage unit 203. The server 20 then transmits information indicating the validity of the audio digital watermark to the determination unit 104. The determination unit 104 determines whether the audio digital watermark is valid or not based on the information indicating the validity of the audio digital watermark. This eliminates the need to store digital watermark information in the storage unit 109 (i.e., the smart speaker 10). This eliminates the possibility of digital watermark information being stolen from the smart speaker 10, further improving the security of the smart speaker 10. Furthermore, since the process of determining validity in the smart speaker 10 can be simplified, the design of the smart speaker 10 can be simplified.
[0059] In step S1003, the mode control unit 105 sets the smart speaker 10 to the privileged mode for a predetermined time (for example, 20 seconds) after the audio digital watermark has been acquired, and then the mode control unit 105 releases the privileged mode after the predetermined time has elapsed.
[0060] In step S1004, the speech recognition unit 103 determines whether or not the user has spoken. If it is determined that the user has spoken, the process proceeds to step S1012. If it is determined that the user has not spoken, the process proceeds to step S1001.
[0061] In step S1005, the voice recognition unit 103 determines whether or not the user has spoken during the period in which the smart speaker 10 is set to the privileged mode (hereinafter referred to as the "privileged period"). If it is determined that the user has spoken during the privileged period, the process proceeds to step S1006. If it is determined that the user has not spoken during the privileged period, the process proceeds to step S1001. Here, "speech has been made" during the privileged period refers to a state in which the user has uttered at least one word during the privileged period. Therefore, if speech has been continuing since before the privileged period and the speech continues during the privileged period, it is determined that "speech has been made" during the privileged period.
[0062] In step S1005, it may be determined whether or not "speech has started" instead of "speech has been made." That is, if it is determined that the user has started speaking during the privileged period, the process proceeds to step S1006. If it is determined that the user has not started speaking during the privileged period, the process proceeds to step S1001. "Speech has started" refers to a change from a state in which the user has not made any sound for a certain period of time to a state in which the user has started making sound.
[0063] In step S1006, speech recognition unit 103 determines whether the user's speech has ended during the privileged period. If it is determined that the user's speech has ended during the privileged period, the process proceeds to step S1010. If it is determined that the user's speech has not ended during the privileged period, the process proceeds to step S1007. Here, "speech has ended" refers to a change from a state in which the user is making speech to a state in which the user is not making speech for a certain period of time. Note that speech recognition unit 103 may recognize the content of the acquired speech and determine that "speech has ended" when an utterance corresponding to a meaningful sentence has ended.
[0064] In step S1007, the voice recognition unit 103 waits until the user finishes speaking. At this time, the voice recognition unit 103 measures the time of the user's speaking as the previous speaking time.
[0065] In step S1008, the voice output unit 108 requests (prompts) the user to speak again. For example, the voice output unit 108 outputs a voice such as "Please speak again." Furthermore, if the smart speaker 10 has a display, the display may display the words "Please speak again."
[0066] In step S1009, the mode control unit 105 sets the smart speaker 10 in privileged mode for the last speech time plus additional time α after the audio output unit 108 requests another speech in step S1008. The mode control unit 105 releases the privileged mode after the last speech time plus additional time α has elapsed. Here, the additional time α may be any time, such as the same time as the predetermined time in step S1003. For example, the additional time α is 10 seconds. When the processing of step S1009 is completed, the process proceeds to step S1005.
[0067] In step S1010, the speech recognition unit 103 determines whether the utterance made by the user during the privileged period is an utterance instructing a specific process. If it is determined that the utterance made by the user is an utterance instructing a specific process, the process proceeds to step S1011. If it is determined that the utterance made by the user is not an utterance instructing a specific process, the process proceeds to step S1012.
[0068] In step S1011, the privileged process execution unit 162 executes a specific process in response to the user's utterance.
[0069] In step S1012, the process execution unit 161 executes normal processing (processing other than the specific processing) in response to the user's utterance. Note that when the privileged mode is not set in the smart speaker 10 and an utterance is made to instruct a specific processing, the process execution unit 161 outputs a voice such as, for example, "The instructed processing cannot be executed unless the privileged mode is set" from the voice output unit 108.
[0070] 4 is a diagram showing an example of a time chart of an execution control process for executing a specific process. In FIG. 4, the privileged command voice is a voice utterance instructing the specific process. Also, in FIG. 4, it is assumed that the audio digital watermark is valid.
[0071] First, at time t1, the smartphone 30 starts reproducing the audio digital watermark.
[0072] At time t2, the playback of the audio digital watermark ends. Then, at time t2, the audio separation unit 102 determines that watermarked music has been acquired (YES in step S1001), and the determination unit 104 determines that the audio digital watermark is valid (YES in step S1002).
[0073] Also, at time t2, the mode control unit 105 sets the smart speaker 10 to the privileged mode (turns the privileged mode on) for a predetermined time from time t2 (step S1003).
[0074] At time t3, mode control unit 105 cancels the privileged mode (turns the privileged mode off) because a predetermined time has elapsed since time t2. Then, at time t3, voice recognition unit 103 determines that the user has spoken during the privileged period (the period from time t2 to t3) (YES in step S1005) and that the user's speech has ended (YES in step S1006). Furthermore, at time t3, voice recognition unit 103 determines that the user's speech is an utterance instructing a specific process (a privileged command voice utterance) (YES in step S1010). Therefore, privileged process execution unit 162 executes the specific process (step S1011).
[0075] At time t4, the smartphone 30 starts playing back the audio digital watermark again.
[0076] At time t5, playback of the audio watermark ends. Then, at time t5, the audio separation unit 102 determines that watermarked music has been acquired (YES in step S1001), and the determination unit 104 determines that the audio watermark is valid (YES in step S1002). In addition, the mode control unit 105 sets the smart speaker 10 to the privileged mode (turns the privileged mode on) for a predetermined time from time t5 (step S1003).
[0077] At time t6, mode control unit 105 cancels the privileged mode (turns the privileged mode off) because a predetermined time has elapsed since time t5. Then, at time t6, speech recognition unit 103 determines that the user has been speaking during the privileged period (the period from time t5 to t6) (YES in step S1005) and that the user has not finished speaking (NO in step S1006). Furthermore, speech recognition unit 103 waits until the user finishes speaking (until time t7) (step S1007).
[0078] At time t7, the voice output unit 108 requests the user to speak again (step S1008). Also at time t7, the mode control unit 105 sets the smart speaker 10 to the privileged mode for the previous speaking time plus additional time α (step S1009).
[0079] At time t8, mode control unit 105 cancels the privileged mode (turns off the privileged mode) because the previous utterance time plus additional time α has elapsed since time t7. Then, at time t8, voice recognition unit 103 determines that the user has spoken during the privileged period (the period from time t7 to t8) (YES in step S1005) and that the user's utterance has ended (YES in step S1006). Furthermore, at time t8, voice recognition unit 103 determines that the user's utterance is an utterance instructing a specific process (utterance of a privileged command voice) (YES in step S1010). Therefore, privileged process execution unit 162 executes the specific process (step S1011).
[0080] (Music playback processing) The music playback process in which the smartphone 30 plays back watermarked music will be described with reference to the flowchart of Fig. 5. The process of the flowchart of Fig. 5 starts when the user performs an operation to instruct playback of watermarked music.
[0081] In step S2001, the information acquisition unit 301 requests the server 20 for digital watermark information.
[0082] In step S2002, the information acquisition unit 301 determines whether or not digital watermark information has been acquired from the server 20. If it is determined that digital watermark information has been acquired, the process proceeds to step S2003. If it is determined that digital watermark information has not been acquired, the process of step S2002 is repeated.
[0083] In step S2003, the music generation unit 302 generates watermarked music based on the digital watermark information.
[0084] In step S2004, the audio output unit 303 plays back the watermarked music.
[0085] In the first embodiment, the smart speaker 10 sets the privileged mode using music containing an audio digital watermark that is imperceptible to the human ear. Therefore, even if a third party listens to the music, Therefore, it is not possible for a third party to know the sound for setting the privileged mode. Furthermore, even if music containing an audio digital watermark is being played, it is not easy for a third party to recognize that this is music for setting the privileged mode. In other words, a third party cannot know how the process for setting the privileged mode is being performed. Therefore, the smart speaker 10 can be set to the privileged mode with high security.
[0086] Furthermore, if no speech is made within a predetermined time after the audio digital watermark is acquired, the smart speaker 10 will not execute the specific processing even if a speech instructing the specific processing is made. Therefore, even if a third party who perceives that the specific processing is being executed after the user makes a speech instructing the specific processing makes a similar speech instructing the specific processing, the specific processing will not be executed. Therefore, the possibility that the specific processing will be executed by a third party other than the legitimate user can be further reduced.
[0087] Instead of watermarked music, any audio with an embedded audio watermark can be used, i.e., instead of music, an audio watermark can be embedded in a chime sound, a conversation, or the like.
[0088] Furthermore, although the smartphone 30 generates the watermarked music, the server 20 (the generation unit of the server 20) may generate the watermarked music based on the digital watermark information and transmit it to the smartphone 30.
[0089] Although the smart speaker 10 has been given as an example of an electronic device that is set to the privileged mode, any electronic device may be used instead of the smart speaker 10. For example, a robot or a home appliance (such as an air conditioner or a light) may be used instead of the smart speaker 10. Here, the specific process performed by the robot may be, for example, a process of delivering a drink to a specific location.
[0090] <Embodiment 2> In the first embodiment, the smart speaker 10 determines whether or not the user has spoken after the privileged period has ended, and executes processing in accordance with the user's utterance. On the other hand, in the second embodiment, the smart speaker 10 determines whether or not the user has spoken during the privileged period, and then executes processing in accordance with the user's utterance. Specifically, in the second embodiment, when the processing of step S1003 or step S1009 is started, the processing of the flowchart shown in FIG. 6 is performed. Note that the information processing system 1 according to the second embodiment has the same configuration as the information processing system 1 according to the first embodiment.
[0091] In step S1101, the mode control unit 105 sets the smart speaker 10 to the privileged mode.
[0092] In step S1102, speech recognition unit 103 determines whether an utterance was made during the privileged period and whether the utterance ended during the privileged period. If it is determined that an utterance was made during the privileged period and whether the utterance ended during the privileged period, the process proceeds to step S1105. On the other hand, if it is determined that an utterance was not made during the privileged period or whether the utterance did not end during the privileged period, the process proceeds to step S1103.
[0093] In step S1103, the mode control unit 105 determines whether a specific time (a predetermined time, or the previous speech time + additional time) has elapsed since the smart speaker 10 was set to the privileged mode. If it is determined that the specific time has elapsed since the smart speaker 10 was set to the privileged mode, the process proceeds to step S1104. If it is determined that the specific time has not elapsed since the smart speaker 10 was set to the privileged mode, the process proceeds to step S1105. Proceed to step S1102.
[0094] In step S1104, mode control unit 105 cancels the privileged mode, and the process proceeds to step S1005.
[0095] In step S1105, mode control unit 105 cancels the privileged mode, and the process proceeds to step S1010.
[0096] According to the second embodiment, the smart speaker 10 can execute a specific process immediately after an utterance instructing the specific process ends during the privileged period. This improves user convenience. On the other hand, by having the smart speaker 10 wait until a specific period has elapsed since being set to the privileged mode and then execute the specific process, as in the first embodiment, the overall process can be simplified, and the configuration of the smart speaker 10 can be simplified.
[0097] <Embodiment 3> In the first embodiment, the smart speaker 10 executes a specific process when an utterance instructing the specific process is made during the privileged period. In this case, for example, if a third party other than the user who played the watermarked music from the smartphone 30 makes an utterance instructing the specific process during the privileged period, the smart speaker 10 executes the specific process. This may result in a decrease in the security of the smart speaker 10.
[0098] Therefore, in the third embodiment, the smart speaker 10 performs a specific process only when a user who has played watermarked music from the smartphone 30 utters an utterance instructing the specific process during the privileged period. Note that the information processing system 1 according to the third embodiment has the same configuration as the information processing system 1 according to the first embodiment.
[0099] In the third embodiment, the voice acquisition unit 101 has a plurality of microphones 121 as shown in FIG. 7. The voice separation unit 102 has a direction detection unit that detects the direction of voice generation. The direction detection unit detects the direction from which each voice input to each of the plurality of microphones 121 is generated, depending on the volume of the voice. For example, if a microphone 121 located south of the center of the smart speaker 10 acquires the voice of the user 601 louder than any of the other microphones 121, the direction detection unit can detect that the direction in which the user 601 who generated the voice is located is south. The direction detection unit may determine the direction from which the voice is generated from one of four directions: "north," "south," "east," and "west," or may detect the direction more precisely.
[0100] Fig. 8 shows a flowchart of the execution control process according to the third embodiment. In the process of the flowchart in Fig. 8, steps S3001 and S3002 are added to the process of the flowchart in Fig. 3. Therefore, only steps S3001 and S3002 will be described below.
[0101] The process of step S3001 is executed when it is determined in step S1002 that the audio digital watermark is valid. In step S3001, the direction detection unit detects the direction from which the watermarked music is coming. For example, as shown in FIG. 7, if the watermarked music is coming from a smartphone 30 located south of the smart speaker 10, the direction detection unit detects south as the direction of origin.
[0102] The process of step S3002 is executed when it is determined in step S1010 that the utterance made by the user during the privileged period is an utterance instructing a specific process. In step S3002, the direction detection unit compares the direction detected in step S3001 with the utterance made during the privileged period. It is determined whether the directions in which the two users are present are the same. If it is determined that the two directions are the same, the process proceeds to step S1011. If it is determined that the two directions are not the same (different), the process proceeds to step S1012.
[0103] According to the third embodiment, the specific process will not be executed unless the direction of the watermarked music (audio digital watermark) and the direction of the utterance instructing the specific process (the direction of the user instructing the specific process) are the same. Therefore, only the voice of the user who generated the watermarked music can be used to instruct the specific process. This reduces the possibility of a third party executing the specific process, thereby improving the security of the smart speaker 10.
[0104] The specific process may be executed not only when the direction in which the watermarked music is generated and the direction in which the utterance instructing the specific process is generated (the direction in which the user instructing the specific process is located) are the same, but also when the angle between the two directions is within a predetermined angle. In other words, if the angle between the two directions is not within the predetermined angle (if the angle between the two directions is greater than the predetermined angle), the smart speaker 10 will not execute the specific process (prohibit execution of the specific process) even if it is set to privileged mode. The smaller the predetermined angle, the higher the security of the smart speaker 10, but the larger the angle, the wider the range in which the user can move. The predetermined angle is, for example, 90 degrees.
[0105] Even in the case where the possibility of a third party performing a specific process is reduced as in the third embodiment, for example, if a user who owns the smartphone 30 lends the smartphone 30 to another user, and the other user generates watermarked music from the smartphone 30, the other user can then cause the smart speaker 10 to perform the specific process. In other words, it is also possible to easily transfer authority to perform a specific process between multiple users.
[0106] The configurations and processes described in the above-described embodiments of the present invention can be used in any combination. [Explanation of symbols]
[0107] 10: Smart speaker, 20: Server, 30: Smartphone 101: voice acquisition unit, 102: voice separation unit, 103: voice recognition unit, 104: Determination unit, 105: Mode control unit, 106: Execution unit, 107: Information update unit, 108: Audio output unit, 109: Storage unit
Claims
1. An electronic device, a sound acquisition means for acquiring sounds generated around the electronic device; an execution means for executing a process according to the user's utterance when the voice acquisition means acquires the user's utterance; a determination means for determining whether the digital watermark is valid when the audio acquisition means acquires the audio including the digital watermark; a control means for setting the electronic device to a privileged mode in which the execution means can execute a specific process if the digital watermark is valid; An electronic device comprising:
2. the execution means executes the specific process if the user's utterance instructing the specific process has finished during the period in which the electronic device is set to the privileged mode.
2. The electronic device according to claim 1, wherein the electronic device is a semiconductor device.
3. the execution means prohibits the execution of the specific process if an angle between a direction from which the voice including the digital watermark is emitted toward the electronic device and a direction from which the user who made the utterance is located toward the electronic device is greater than a predetermined angle.
3. The electronic device according to claim 2.
4. the execution means prohibits the execution of the specific process if the user who made the utterance is located in a direction different from the direction from which the voice including the digital watermark was uttered relative to the electronic device.
3. The electronic device according to claim 2.
5. the control means sets the electronic device in the privileged mode for a predetermined time after the audio acquisition means has acquired the valid digital watermark, and releases the privileged mode after the predetermined time has elapsed.
5. The electronic device according to claim 1, wherein the first and second electrodes are electrically connected to the first and second electrodes.
6. The control means: 1) if the user is speaking during a period in which the electronic device is set to the privileged mode and the speaking has not ended, prompts the user to speak again after the period has ended; and 2) sets the electronic device to the privileged mode for a period of time after prompting the user to speak again that is longer than the length of the previous speaking.
5. The electronic device according to claim 1, wherein the first and second electrodes are electrically connected to the first and second electrodes.
7. the digital watermark includes authentication information and information indicating a validity period; the determination means determines that the digital watermark is valid when the authentication information is predetermined information and the current time is included in the validity period; 5. The electronic device according to claim 1, wherein the first and second electrodes are electrically connected to the first and second electrodes.
8. the execution means executes the predetermined process when the user makes an utterance instructing a predetermined process other than the specific process, regardless of whether the electronic device is set to the privileged mode or not.
5. The electronic device according to claim 1, wherein the first and second electrodes are electrically connected to the first and second electrodes.
9. the control means cancels the privileged mode when the user finishes speaking during a period in which the electronic device is set to the privileged mode.
5. The electronic device according to claim 1, wherein the first and second electrodes are electrically connected to the first and second electrodes.
10. The digital watermark is expressed by audio having a frequency not included in the frequency range of 20 Hz to 20,000 Hz.
5. The electronic device according to claim 1, wherein the first and second electrodes are electrically connected to the first and second electrodes.
11. An electronic device according to any one of claims 1 to 4; an audio output device that outputs audio including the digital watermark; An information processing system comprising:
12. A method for controlling an electronic device, comprising: a sound acquisition step of acquiring sound generated around the electronic device; an execution step of executing a process according to the user's utterance when the user's utterance is acquired in the voice acquisition step; a determination step of determining whether or not the digital watermark is valid when the audio including the digital watermark is acquired in the audio acquisition step; a control step of setting the electronic device to a privileged mode in which a specific process can be executed in the execution step if the digital watermark is valid; 1. A method for controlling an electronic device, comprising:
13. A program for causing a computer to execute each step of the control method according to claim 12.
Citation Information
Patent Citations
Method and device for voice separation of compound voice data, method and device for specifying speaker, computer program, and recording medium
JP2003005790A
Device and related method for enabling authentication using a digital music authentication token.
JP2011517349A
Authentication system, authentication method, and program
JP2019185605A
Smart speaker, secure element, program, information processing method, and distribution method
JP2020010144A
User authentication device for authenticating user, program executed by user authentication device, program executed by input device for authenticating user, and computer system having user authentication device and input device
JP2020064689A