Acoustic characteristics of speech-enabled computer systems
By using acoustic features in voice-enabled computer systems to prevent playback attacks, the problem of voice-based authentication systems being vulnerable to playback attacks is solved, achieving higher security and reliability.
Patent Information
- Application Number
- CN202011502746.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-12-18
AI Technical Summary
Voice-based authentication systems are vulnerable to replay attacks, and unauthorized users can fake identity authentication by recording the voice of an authorized user and playing it later.
By playing and mixing specific noise patterns as acoustic features while the user speaks, a speech-enabled computer system captures mixed audio of these acoustic features and user speech and analyzes at local or backend servers to verify the presence or absence of acoustic features, thereby preventing playback attacks.
It effectively prevents playback attacks, improves the security and reliability of voice-based authentication, and ensures that the system can effectively detect and reject unauthorized playback attempts.
Smart Images

Figure CN113012715B_ABST
Abstract
Description
Background Art
[0001] The present disclosure relates generally to voice-based authentication in voice-enabled computer systems, and more particularly, to acoustic signatures for voice-enabled computer systems.
[0002] Voice-enabled computer systems are becoming increasingly popular. Examples include smartphones with "intelligent" assistants, smart speakers, and other devices that can respond to voice input. Voice-enabled computer systems typically use a microphone to detect user speech and use a speaker to play an audible response using a synthesized voice. Some voice-enabled computer systems can process detected speech locally to determine a specific request (or command) from the user, and then communicate with other devices or computer systems (usually referred to as servers) as needed to process the request and determine the response (which can be an action taken or a speech response or both). Other voice-enabled computer systems can forward recorded audio to a "backend" server that processes the speech and determines the request. In addition, other voice-enabled computer systems use a combination of local (client-based) and backend (server-based) processing to respond to user input.
[0003] Voice-enabled computer systems can support a range of user interactions involving varying degrees of security risk. Some interactions, such as requests for weather forecasts or requests to play music, pose little security risk regardless of who makes the request. Other interactions, such as requests to unlock a door or access a bank account, can pose significant security risks if made by an unauthorized person. Therefore, some voice-enabled devices can use voice-based authentication techniques to confirm that the request is made by an authorized person before responding to the request. Voice-based authentication can include comparing the audio characteristics of a recorded request with known audio characteristics of the authorized person's voice. For example, a user can speak a passphrase, which can be compared to a recorded version to assess frequency characteristics, speech rate, word pronunciation, and / or other characteristics.
[0004] Voice-based authentication may be vulnerable to "replay" attacks, in which an unauthorized person obtains a recording of an authorized person's voice and later plays the recording back in an attempt to impersonate a voice-enabled computer system. If the recording is of high enough quality, the voice-based authentication process may be unable to distinguish the recording from the authorized person's real-time speech. Therefore, it may be desirable to prevent or detect replay attacks. Summary of the invention
[0005] Embodiments disclosed herein relate to acoustic features that can be used in conjunction with a voice-enabled computer system. In some embodiments, an acoustic feature can be a specific noise pattern (or other sound) that is played when a user speaks and mixed with the user's speech in an acoustic channel. The microphone of the voice-enabled computer system can capture a mixture of the acoustic feature and the user's voice as recorded audio. The voice-enabled computer system can analyze the recorded audio (locally or at a backend server) to verify whether the expected acoustic feature is present and / or whether a previous acoustic feature is not present.
[0006] Embodiments of the present invention relate to a method performed by a speech-enabled computer system. A method may include obtaining a current random number and operating a speaker of the speech-enabled computer system to generate a current acoustic feature based on the current random number. The method may further include: operating a microphone of the speech-enabled computer system while generating the current acoustic feature to record audio including the current acoustic feature and speech input; and verifying the recorded audio based at least in part on the current acoustic feature. In the event that the recorded audio passes verification, the method may also include: processing the recorded audio to extract the speech input; processing the speech input to determine an action to be performed; and performing the action.
[0007] Some embodiments of the present invention are directed to a server computer, comprising a processor, a memory, and a computer-readable medium coupled to the processor. The computer-readable medium may have code stored therein, the code being executable by the processor to implement a method, the method comprising: generating a random number; providing the random number to a voice-enabled client having a speaker and a microphone and capable of generating an acoustic feature based on the random number; receiving an audio recording from the voice-enabled client; verifying the audio recording based at least in part on detecting the acoustic feature in the audio recording; and processing a user utterance in the audio recording only if the audio recording passes verification.
[0008] Some embodiments of the present invention relate to a speech-enabled computer system, comprising a processor, a memory, a speaker operable by the processor, a microphone operable by the processor, and a computer-readable medium coupled to the processor. The computer-readable medium may have code stored therein, the code being executable by the processor to implement a method, the method comprising: obtaining a random number; operating the speaker to generate an acoustic feature based on the random number; and while generating the acoustic feature, operating the microphone to record an audio recording comprising the acoustic feature and speech input.
[0009] The following detailed description and attachedFigure 1 This will provide a better understanding of embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A simplified block diagram of a speech-enabled computer system is shown in accordance with some embodiments.
[0011] Figure 2 A flow chart illustrating a process for generating and validating acoustic signatures that may be used in some embodiments.
[0012] Figure 3 An example of amplitude modulation is shown which may be used in some embodiments.
[0013] Figure 4 An example of frequency modulation is shown which may be used in some embodiments.
[0014] Figure 5 An example of digital modulation is shown which may be used in some embodiments.
[0015] Figure 6 A flow chart illustrating a process for decoding a random number that may be used in some embodiments.
[0016] Figure 7 A simplified example of decoding a random number according to some embodiments is shown.
[0017] Figure 8 Shown is a simplified representation of an acoustic signature that may be produced according to some embodiments.
[0018] Fig. 9 is a flow chart of a verification process according to some embodiments.
[0019] Fig.10 A communication diagram illustrating client / server interactions according to some embodiments.
[0020] the term
[0021] The following terminology may be used herein.
[0022] A "server computer" may include a powerful computer or cluster of computers. For example, a server computer may be a mainframe, a cluster of small computers, or a group of servers that function as a unit. In one example, a server computer may be a database server coupled to a network server. A server computer may include one or more computing devices and may use any of a variety of computing structures, arrangements, and compilations to service requests from one or more client computers.
[0023] A "client" or "client computer" may include a computer system or other electronic device that communicates with a server computer to make requests to the server computer and receive responses. For example, a client may be a laptop or desktop computer, a mobile phone, a tablet computer, a smart speaker, a smart home management device, or any other user-operable electronic device.
[0024] "Memory" may include suitable one or more devices that can store electronic data. Suitable memory may include non-transitory computer-readable media that stores instructions that can be executed by a processor to implement the desired method. Examples of memory may include one or more memory chips, disk drives, and the like. Such memory may operate using any suitable electrical, optical, and / or magnetic operating modes.
[0025] A "processor" may include any suitable one or more data computing devices. A processor may include one or more microprocessors working together to perform the desired functions. A processor may include a CPU including at least one high-speed data processor sufficient to execute program components for executing user and / or system generated requests. The CPU may be a microprocessor such as AMD's Athlon, Duron, and / or Opteron; IBM and / or Motorola's PowerPC; IBM and Sony's Cell processor; Intel's Celeron, Itanium, Pentium, Xeon, and / or XScale; and / or similar processors.
[0026] A "communication device" may include any electronic device that can provide communication capabilities, including through a mobile phone (wireless) network, a wireless data network (e.g., 3G, 4G or similar network), Wi-Fi, Wi-Max, or any other communication medium that can provide access to a network such as the Internet or a private network. Examples of communication devices include mobile phones (e.g., cellular phones), PDAs, tablet computers, netbooks, laptop computers, personal music players, handheld dedicated readers, wearable devices (e.g., watches), vehicles (e.g., cars), etc. A communication device may include any suitable hardware and software for performing such functions, and may also include multiple devices or components (e.g., when a device remotely accesses a network by being fastened to another device (i.e., using the other device as a repeater), the two devices together may be considered a single communication device). A communication device may store and capture a user's voice recording, and may store the recording locally and / or forward the recording to another device (e.g., a server computer) for processing. The mobile device may store the recording on a secure memory element.
[0027] A "user" may include an individual who operates a voice-enabled computer system by speaking to the voice-enabled computer system. In some embodiments, a user may be associated with one or more personal accounts and / or devices.
[0028] A "random number" may include any number or bit string, etc., generated and valid for a single transaction. Random numbers may be generated using a random or pseudo-random process, many examples of which are known in fields such as cryptography. DETAILED DESCRIPTION
[0029] The following description of exemplary embodiments of the present invention is presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form described, and those skilled in the art will appreciate that many modifications and variations are possible. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention and various embodiments with various modifications to suit the particular use contemplated.
[0030] Certain embodiments described herein relate to acoustic features in combination with a voice-enabled computer system. The acoustic feature can be, for example, a dynamically generated noise pattern (or other sound) that is played by a speaker of the voice-enabled computer system when the user speaks and mixed with the user's speech in an acoustic channel. The microphone of the voice-enabled computer system can capture a mixture of the acoustic feature and the user's voice as recorded audio. The voice-enabled computer system can analyze the recorded audio (locally or at a back-end server) to verify whether the expected acoustic feature exists and / or whether there is no previous acoustic feature, and if the verification analysis fails, the request can be rejected. In some embodiments, the use of acoustic features in a voice-based identity authentication process can improve reliability, for example, by providing protection to prevent replay attacks.
[0031] Figure 1 A simplified block diagram of a voice-enabled computer system 100 is shown in accordance with some embodiments. The system 100 includes a voice-enabled client 102 and a server 104 communicatively coupled via a network 106 (eg, the Internet).
[0032] The voice-enabled client 102 may be any electronic device capable of receiving and responding to spoken commands. Examples include smartphones, smart speakers, smart home controllers, or other devices that implement voice response features. In some cases, the voice-enabled client 102 may be a personal device associated with a particular user (e.g., a smartphone); in other cases, the voice-enabled client 102 may be a shared device used daily by multiple users (e.g., a smart home controller). The voice-enabled client 102 may include various components, such as a speaker 110, a microphone 112, a network interface 114, and a processor 116. The speaker 110 may include any device or component capable of converting an input signal (e.g., a digital or analog electronic signal received via a wired or wireless interface) into sound (i.e., a pressure wave in a compressible medium such as air). The microphone 112 may include any device or component capable of converting sound into an output signal (e.g., a digital or analog electronic signal that can be recorded and / or analyzed). The network interface 114 may include hardware and software components that support communication via the network 104. For example, the network interface 114 may support wireless communication protocols conforming to standards such as Wi-Fi, LTE, 5G, etc. and / or wired communication protocols such as Ethernet protocols. Many embodiments of speakers, microphones, and network interfaces are known in the art, and detailed descriptions thereof are omitted.
[0033] The processor 116 may include one or more programmable logic circuits, such as a microprocessor or microcontroller, an ASIC, an FPGA, etc. By executing appropriate program code, the processor 116 may be configured to perform various operations, including the operations described herein. In some embodiments, the processor 116 may have an associated memory or storage subsystem ( Figure 1 110 ), which stores program code to be executed and data that may be generated or consumed in the process of executing the program code. In some embodiments, the program code may include code that implements a feature generator module 120 and a speech processing module 122. The feature generator module 120 may generate acoustic features to be played by the speaker 110; examples are described below. The speech processing module 122 may receive an output signal from the microphone 112 and perform various processing operations on the signal. Examples of such operations include noise reduction, speech parsing (e.g., identifying words based on audio signals), and the like. In some embodiments, the speech processing module 122 may also generate a speech response signal to be played by the speaker 110. The operation of the speech processing module 122 may be implemented, for example, using conventional techniques.
[0034] Server 104 may be any server computer, server farm, cloud-based computing system, etc. Server 104 may include network interface 132, processor 134, and user records 136. Similar to network interface 114, network interface 132 may include hardware and software components that support communication via network 104. Processor 134 may include one or more programmable logic circuits, such as microprocessors or microcontrollers, ASICs, FPGAs, etc. By executing appropriate program code, processor 134 may be configured to perform various operations, including the operations described herein. In some embodiments, processor 134 may have an associated memory or storage subsystem ( Figure 1 140 ), which stores program code to be executed and data that may be generated or consumed in the process of executing the program code. In some embodiments, the program code may include code that implements a speech interpreter 140, a command execution module 142, a random number generator 146, and a feature verifier 148. The speech interpreter 140 can interpret an audio recording provided from a voice-enabled client 102 to identify a user request and determine one or more operations to be performed in response to the request. The command execution module 142 can perform an operation in response to the speech interpreter 140 and send the result of the operation to the voice-enabled client 102. The operations of the speech interpreter 140 and the command execution module 142 can be implemented, for example, using conventional techniques. A random number generator 146 can be used to generate a random number (e.g., a number or a bit string) that will be used by a feature generator 120 of the voice-enabled client 102 when generating acoustic features; an example is described below. The feature verifier 148 can analyze an audio recording received from a voice-enabled client 102 to detect whether an acoustic feature is present; an example is described below. The user record 136 may include information about the user of the voice-enabled client 102. For example, the user record 136 for a particular user may include a voiceprint 150, which may include any record of unique characteristics of the user's voice that may be used for voice-based authentication; and a random number list 152, which may store one or more random numbers that were previously used to generate an acoustic signature and are no longer considered valid. Other user-specific information related to specific services provided by the server 104 (e.g., bank account information or other account information) may also be stored in the user record 136.
[0035] In operation, the voice-enabled client 102 (supported by the server 104) may provide a voice response interface to support user interaction with the voice-enabled client 102 and / or other devices to which the voice-enabled client 102 may be connected (e.g., via the network 106). Thus, for example, a user may speak a request, such as, "What's today's weather forecast?" or "What's my bank balance?" The microphone 112 may pick up the sound and provide a corresponding electronic signal to the voice processing module 122. The voice processing module 122 may perform signal processing, for example, to determine that the request should be processed by the server 106. Thus, the recorded electronic signal may be sent to the server 106.
[0036] The server 106 can process the recorded electronic signal using the speech interpreter 142 to extract the request, and can call the command execution module 144 to perform a response operation (e.g., retrieve the weather forecast or bank balance). The command execution module 144 can return the response to the voice-enabled client 102, for example, as an audio file to be played using the speaker 110, or as data to be converted into audio by the speech processing module 122 and played using the speaker 110.
[0037] In some (or all) instances, the server 104 (or the voice-enabled client 102) may use voice-based authentication techniques to verify the identity of the user. For example, the server 104 may store the voiceprint 150. When a voice-based identity check is required, the voice-enabled client 102 may prompt the user to provide input for verification. For example, the user may be prompted to speak a passphrase, which is sent to the server 104. The server 104 may compare the passphrase to the stored voiceprint 150 and confirm the identity of the user based on the result of the comparison.
[0038] According to some embodiments described herein, voice-based authentication or other operations of a voice-enabled computer system 100 may include the use of acoustic features. As used herein, an "acoustic feature" may include a sound pattern that is intentionally played into an environment in which a user is speaking so that a microphone that picks up the user's voice may also pick up the acoustic feature. Acoustic features may be dynamic, which means that different acoustic features are generated for different recording events. In some embodiments, acoustic features may be designed to make it difficult or impossible for a party that does not understand a particular sound pattern to filter out the acoustic feature from a recording that includes a mixture of user speech and acoustic features, while making the user's speech unavailable for voice-based authentication, making it difficult or impossible to use such recordings in a replay attack. For example, an acoustic feature may include a frequency within the frequency range of human speech, and the acoustic feature may change over time (frequency and / or volume) in an unpredictable manner. In some embodiments, a device having information defining the acoustic features played during a particular recording event, such as a server 104 (or client 102), may use acoustic features to confirm the presence of a user. Additionally or alternatively, a device such as server 104 (or client 102) may also detect a replay attack by detecting "old" acoustic features in a recording of a user's speech. Specific examples of generating and verifying acoustic features are described below.
[0039] It should be appreciated that the voice-enabled computer system 100 is illustrative and that many variations and modifications are possible. For example, the voice-enabled client 102 may have Figure 1Other components not shown in the figure, such as a display, a touch screen, a keypad or other visual or tactile (or other) user interface components. The voice-supported client 102 can also be implemented using multiple discrete devices. For example, a speaker 110 and / or a microphone 112 can be provided in an earplug, earphone or headset connected to a larger device such as a smart phone (using a wired or wireless connection). It should be noted that the speaker located in the earplug, earphone or headset may not be the best choice for generating acoustic features, because the sound from such a speaker may not be picked up by the microphone. In this case, the speaker in the larger device to which the earplug, earphone or headset is connected can be used instead. The division of operations between the voice-supported client 102 and the server 104 can also be modified. For example, some voice-supported clients can simply provide recorded utterances to the server for interpretation, while other voice-supported clients can process recorded utterances locally to determine user requests, and then relay the requests to the appropriate server as needed. In some cases, the voice-supported client may be able to interpret and respond to certain requests locally without involving a server. Thus, in various embodiments, voice-based authentication can be performed locally on a voice-enabled client, or remotely at a server (or both). It should be understood that the acoustic features described herein can be used by any device (including a server and / or client) that performs voice-based authentication in any situation where voice-based authentication is being performed. As described below, the acoustic features can be specified by the same server or client that verifies the acoustic features.
[0040] Figure 2 A flow chart showing a process 200 for generating and validating acoustic features that may be used in some embodiments. The process 200 may be performed at Figure 1 The process 200 may be implemented in the voice-enabled computer system 100 or in any other voice-enabled computer system. The process 200 may begin at any point where the voice-enabled system determines that an acoustic feature is useful, such as when a user is prompted to speak a passphrase or provide other voice input that may be used to authenticate the user.
[0041] At box 202, a voice-enabled device in a voice-enabled computer system (e.g., a voice-enabled client 102) can obtain a current random number. In some embodiments, the random number can be obtained from a random number generator 146 of a server 104. According to an embodiment, the client 102 can request a random number from the server 104, or the server 104 can determine, for example, based on a specific request from the client 102 that an acoustic feature should be used, and can provide the random number to the client 102. In other embodiments, the random number can be generated locally by the voice-enabled client; this can be useful, for example, where the voice-enabled client performs voice-based authentication locally. The random number can be, for example, a randomly generated number, or other tokens or data that can be used to determine a sound sequence to be generated. Specific examples are described below.
[0042] At box 204, the voice-enabled device may operate a speaker (e.g., speaker 110) to generate an acoustic feature (e.g., a sound sequence) based on the current random number. At the same time, at box 206, the voice-enabled device may operate a microphone (e.g., microphone 112) to record audio. The acoustic feature and the speech input from the user may overlap in time and may be recorded as a single audio recording including a mixture of the acoustic feature and the speech input from the user. Recording audio at box 206 may include storing a digital or analog representation of the audio signal in any storage medium including a short-term storage device such as a cache or buffer for a period of time long enough to support verification and other audio processing operations associated with the voice-enabled computer system.
[0043] At box 208, the recorded audio can be verified based at least in part on the presence of acoustic features. In some embodiments, verification can include analyzing the recorded audio to determine whether the acoustic features generated at box 204 are present in the recorded audio. Verification can also include other operations, such as analyzing the recorded audio to determine whether there are different acoustic features from the previous iteration of process 200. Detecting the acoustic features from the previous iteration of process 200 can indicate a replay attack using the audio recorded by the intruder during the previous iteration, and can be the basis for invalidating the recorded audio. A specific example of the verification process is described below. Depending on the implementation, verification can be performed by a server responding to the request, or locally performed on a voice-enabled client that records audio at box 206. In some embodiments, the verification at box 208 can also include other operations, such as extracting the user's voice from the recorded audio, and comparing the extracted voice with the user's stored voiceprint.
[0044] At box 210, process 200 can determine whether the recorded audio is valid or invalid based on the verification result at box 208. If the recorded audio is valid, process 200 can continue to operate on the speech input. For example, at box 212, process 200 can process the recorded audio to extract the speech input; at box 214, process 200 can process the speech input to determine the action to be performed; at box 216, process 200 can perform the action. If at box 210, the recorded audio is invalid, at box 218, process 200 can ignore the recorded audio. In some embodiments, ignoring the recorded audio can include notifying the user that the input is rejected. If necessary, the user can be prompted to try again, or the user's further activities can be restricted (for example, if a replay attack is suspected).
[0045] It should be understood that process 200 is illustrative and may be subject to change or modification. For example, operations described with reference to a single box may be performed at different times, operations described with reference to different boxes may be combined into a single operation, the order of operations may be changed, and some operations may be omitted entirely. As long as the acoustic features are played in time and overlap with the recording of the user's speech, the acoustic features may be used during the verification operation as described herein. In some embodiments, process 200 may be used whenever a user provides voice input to a voice-enabled computer system. For example, some voice-enabled computer systems operate by listening to a specific activation phrase (e.g., "Hey Assistant") and begin processing other voice-based inputs in response to detecting the activation phrase. Therefore, a voice-enabled computer system with an activation phrase may obtain a random number and begin playing the acoustic features in response to detecting the activation phrase. In other embodiments, the use of process 200 may be more selective. For example, a user may interact with a voice-enabled computer system without an acoustic feature until the user requests an action involving sensitive information, at which point process 200 may be called. An example of a selective call of process 200 is described below.
[0046] In some embodiments, the acoustic features may be produced as modulated tones within the frequency range of human speech. Figure 3-5 An example of base (or carrier) frequency modulation that may be used in some embodiments is shown.
[0047] Figure 3 An example of amplitude modulation is shown. A given frequency f is modulated with amplitude according to input wave 304. c The carrier wave 302 is modulated to produce a modulated wave 306. When played by a speaker, the modulated wave 306 can produce a sound of constant pitch (frequency) whose loudness (amplitude) varies over time.
[0048] Figure 4An example of frequency modulation is shown. A given frequency f is modulated according to an input wave 404. c The carrier wave 402 is modulated to produce a modulated wave 406. In this example, the deviation from the carrier frequency is a function of the amplitude of the input wave 404. When played by a speaker, the modulated wave 406 can produce a sound of constant loudness (amplitude) but variable pitch (frequency).
[0049] Figure 5 An example of digital modulation is shown. A given frequency f is digitally modulated according to an input binary signal 504. c The carrier wave 502 is modulated to produce a modulated wave 506. In this example, when the binary signal is in the "1" state, the frequency increases above the carrier frequency, and when the binary signal is in the "0" state, the frequency decreases below the carrier frequency. When played by a speaker, the modulated wave 506 can produce a sound of constant loudness (amplitude) but variable pitch (frequency).
[0050] Figure 3 , 4 The modulation schemes of and 5 are illustrative, and other modulation schemes may be used. In some embodiments, the selection of a modulation scheme may be one element of the acoustic signature, and multiple modulation schemes may be combined. It is contemplated that the acoustic signature may be audible to the user and may sound like noise.
[0051] In some embodiments, random numbers are used to define the acoustic signature for a particular case of process 200. The random number may be a bit string (or number) of arbitrary length generated using a process such that knowledge of previous random numbers cannot be used to predict future random numbers. A random process, a pseudo-random process, or any other process with a sufficient degree of unpredictability may be used. The generation of the random number may be performed by a server (or client) that will verify the acoustic signature. (For example, the generation of the random number may be performed by the random number generator 146 of the server 104 in the system 100.)
[0052] To generate the acoustic signature, the random number may be "decoded" to define the corresponding sound wave. The decoding of the random number may be performed by the device that generated the acoustic signature code (e.g., by the feature generator 120 of the voice-enabled client 102 in the system 100). The decoding scheme may vary, as long as the same decoding scheme is used for generating the acoustic signature and for subsequently verifying the acoustic signature.
[0053] In some embodiments, the decoding scheme may be a frequency hopping scheme with variable modulation. Figure 6 Flowchart showing a process 600 for decoding a random number that may be used in some embodiments. Process 600 may be performed, for example, in Figure 1The feature generator 120 of the voice-enabled client 102 or implemented in any other voice-enabled device.
[0054] At box 602, the process 600 may receive a random number (r), for example, from the server 104 or another random number generator. At box 604, the process 600 may separate the random number into two components (r1 and r2). For example, if the random number is an N-bit string of an integer N, the first component r1 may be defined as the first N / 2 bits, and the second component r2 may be defined as the last N / 2 bits. Other separation techniques may be used, and the lengths of the two components need not be equal. At box 606, the process 600 may use the first component r1 to define a frequency hopping sequence. The frequency hopping sequence may be a sequence of carrier frequencies used during consecutive time intervals. At box 608, the process 600 may use the second component r2 to define a modulation input to be applied to the carrier frequency during each time interval. In some embodiments, different modulation inputs for different time intervals may be defined. At box 610, the process 600 may generate a corresponding drive signal for a loudspeaker. For example, the drive signal for the first time interval may be generated by modulating the carrier frequency of the first time interval (determined according to the first component r1 ) using the modulation input of the first time interval (determined according to the second component r2 ), and so on.
[0055] Figure 7 A simplified example of random number decoding according to an embodiment of process 600 is shown. In this example, it is assumed that the frequency hopping sequence includes four time intervals. At each time interval, one of a set of carrier frequencies and one of a set of modulation modes are allocated based on the received random number r (700). The random number r can be interpreted as a set of eight numbers (shown as decimal numbers for convenience). The random number r can be divided into components r1 and r2 by allocating the first four numbers to the first component r1 (702) and the last four numbers to the second component r2 (704). Other schemes can be used, such as allocating spare numbers to spare components.
[0056] Table 706 shows an explanation of components r1 and r2 according to some embodiments. For the purpose of this example, it is assumed that a set of at least five different carrier frequencies has been defined, and a set of at least five different modulation modes has also been defined. Each carrier frequency is mapped to an index (f1, f2, etc.), and each modulation mode is mapped to an index (m1, m2, etc.). For each time interval, the acoustic feature is determined by applying the modulation mode identified by the corresponding element of r2 to the carrier frequency identified by the number at the corresponding element of r1. As shown in Table 706, during the first time interval, the carrier frequency is f1 (because the first element of r1 is 1), and the modulation mode is m5 (because the first element of r2 is 5). Therefore, during the first time interval, the acoustic feature corresponds to the carrier frequency f1 modulated according to the modulation mode m5. During the second time interval, the carrier frequency is f2 (because the second element of r1 is 2), and the modulation mode is m4 (because the second element of r2 is 4). Therefore, during the second time interval, the acoustic feature corresponds to the carrier frequency f2 modulated according to the modulation mode m4. Similarly, during the third time interval, the acoustic signature corresponds to carrier frequency f5 modulated according to the modulation mode ml, and during the fourth time interval, the acoustic signature corresponds to carrier frequency f3 modulated according to the modulation mode ml.
[0057] Figure 8 A simplified graphical representation of an acoustic signature that may be generated according to table 702 in some embodiments is shown. Time is shown on the horizontal axis, and different carrier frequencies are represented on the vertical axis. For each time interval (t1 to t4), a carrier frequency is selected according to the corresponding element of r1. The time interval may be a short duration, such as 1-10 ms, 10-50 ms, etc. The carrier frequencies f1 to f5 are different frequencies, and the signal may jump from one frequency to the next. In addition, during each time interval, a modulation pattern ( m5, m4, m1, m1). Each modulation mode may include amplitude modulation, frequency modulation, digital modulation or any other modulation mode.
[0058] It should be understood that Figure 6-8A process and modulation scheme that can be used to generate an acoustic feature based on a random number are described. Any combination of carrier frequencies and / or modulation modes, as well as any number of carrier frequencies and / or modulation modes, can be used. In some embodiments, it may be desirable to select a carrier frequency within a "voice range", which can be a frequency range associated with general human speech or associated with the speech of a particular user. Using frequencies within the sound range can prevent the acoustic feature from being filtered out by noise reduction in a microphone or associated signal processing component. In addition, as described below, the presence of an "old" acoustic feature can be an indication of a replay attack, and using frequencies within the sound range can make it more difficult for an intruder to filter out the "old" acoustic feature from a recording of a user's voice. The volume of the acoustic feature can be changed as desired, as long as it is loud enough to be picked up by the microphone used in a particular situation. In the case where the acoustic feature uses a frequency within the human hearing range, the feature can be audible, and a low volume may be desired.
[0059] According to some embodiments, acoustic features may be used to verify the integrity of an acoustic channel, for example in conjunction with verifying a user's identity. Fig. 9 is a flow diagram of a verification process 900 that may be used, for example, at block 208 of process 200, according to some embodiments. Process 900 may be performed, for example, by feature verifier 148 of server 102 of system 100 or by any other device that performs voice-based identity authentication.
[0060] Process 900 may begin at box 902 by receiving an audio signal, which may be a digital or analog signal, depending on the implementation. Assume that the audio signal includes a combination (or mixture) of user speech and acoustic features; for example, the audio signal may correspond to the audio recorded at box 206 of process 200. At box 904, process 900 may determine the "current" acoustic features (i.e., the acoustic features that may have been playing when the audio signal was recorded). For example, feature verifier 148 (or any other device performing process 900) may receive a random number for generating the current acoustic features. In some embodiments, feature verifier 148 may receive the random number from random number generator 146; other techniques for providing random numbers may also be used. Based on the random number and an applicable decoding scheme for generating acoustic features based on the random number (e.g., as described in reference Figure 6-8 As described above, process 900 can determine the expected frequency and / or amplitude pattern of the current acoustic feature.
[0061] At box 906, process 900 can determine whether there is a current acoustic feature in the received audio signal. Conventional signal analysis techniques or other techniques can be used to determine whether the received audio signal includes a component corresponding to the current acoustic feature. If there is no current acoustic feature, then at box 910, process 900 can treat the received audio signal as an invalid input. In various embodiments, processing the audio signal as invalid can include any one or all of the following: ignoring the audio signal, prompting the user to try again (this can include providing a different random number for generating a different acoustic feature for the next attempt), generating a notification of an invalid access attempt to the user, or other actions as needed.
[0062] If there is a current acoustic feature at block 906, then at block 912, process 900 may determine one or more "previous" acoustic features to detect. Each previous acoustic feature may be an acoustic feature that would be generated by the voice-enabled computer system in response to a random number generated in conjunction with a previous authentication process. In some embodiments, server 104 (or other device performing process 900) may store random numbers that have been previously used for a particular user or voice-enabled client, e.g., Figure 1 The random number list 152 shown in FIG. 150 is a list of random numbers. Each stored random number can be used to determine a possible previous acoustic signature. In some embodiments, there may be an upper limit on the number or lifetime of stored random numbers; however, no specific limit needs to be imposed, and at block 912, any number of previous acoustic signatures may be determined.
[0063] At box 914, process 900 can determine whether there are any previous acoustic features in the received audio signal. For example, using each previous acoustic feature, the same signal analysis technique applied at box 906 can be applied at box 914. In this example, the presence of the previous acoustic feature is considered to indicate that the user's voice was originally recorded when the previous acoustic feature was the current acoustic feature and is now playing. Therefore, if the result of box 914 is to determine that there are previous acoustic features, then at box 910, process 900 can treat the received audio signal as invalid input, which helps prevent replay attacks. In some embodiments, process 900 can notify the user (e.g., via different channels, such as email or text messaging) that suspicious voice input has been received and that the user's security may be compromised; the notification can also provide suggestions for remedial measures or other information as needed.
[0064] If no previous acoustic features exist at block 914 (and current acoustic features exist as determined at block 906), then at block 918, process 900 may accept the input and perform further processing on the audio signal. Any type of processing operation is supported. For example, if the expected user input includes a passphrase, further processing at block 918 may include, for example, detecting the passphrase, comparing the received speech pattern to a stored voiceprint (e.g., Figure 1 The process 900 may be used to compare the user's voiceprint 150 with the user's voiceprint 150, and determine whether to verify the user's identity based on the result. Other processing operations may include: other voice-based authentication operations; parsing speech input to determine the command to be executed or extracting information to be used when executing the command (e.g., access code, transaction amount, etc.); executing the command using speech input; and so on. In some embodiments, before performing voice-based authentication or other processing operations, process 900 may filter out the current acoustic features from the audio signal, for example, using conventional techniques for removing known signal components from mixed audio signals, to obtain a noise-free speech signal. As described above, the acoustic features may have frequency overlap with the user's voice, which may lead to errors in matching speech patterns (e.g., the spectrum of the user's voice) and / or speech interpretation. Filtering out the acoustic features may provide a cleaner speech signal and reduce these types of errors, depending on the specific processing to be performed; of course, it is not necessary to filter out the acoustic features.
[0065] It should be noted that process 900 can reliably filter out the current acoustic feature (if necessary) because the device performing process 900 has information defining the current acoustic feature, which makes it simple and straightforward to identify and filter out the current acoustic feature from the audio signal. However, a device or entity that lacks the information defining the acoustic feature will not be able to filter the acoustic feature out of the mixed audio signal without distorting the voice component of the signal. Therefore, if an intruder records the user while an acoustic feature of the type described herein is played, the intruder will not be able to filter out the acoustic feature while still retaining a recording that can satisfy voice-based authentication. (It should be noted that recording without the acoustic feature and later replaying it may not be detected as a replay attack by process 900.)
[0066] In various embodiments, acoustic features may be generated in any case where user speech is captured by a voice-enabled device. However, acoustic features are likely to be audible, and it may be desirable to use acoustic features selectively. Therefore, in some embodiments, acoustic features may be selectively used in conjunction with verifying the identity of a user. In one instance, a user of a voice-enabled device may request sensitive information (e.g., a bank balance or health record) or authorize a transaction (e.g., pay a bill or send health data to a third party). Before satisfying the request, the voice-enabled device or server that satisfies the request may require the user to perform identity authentication, e.g., by saying a passphrase or providing some other vocal input for voice-based identity authentication. Acoustic features may be played when the user says the passphrase. If the voice-enabled device or server that uses the passphrase to verify the user's identity also has an acoustic feature, verification of the acoustic feature may be used to prevent an intruder from using a recording of the user who said the passphrase.
[0067] Fig.10 A communication diagram illustrating client / server interaction incorporating selective use of acoustic features according to some embodiments. Client 1002 may be a voice-enabled client, such as Figure 1 The voice-enabled client 102. The server 1004 can be any server that receives and responds to voice requests, such as Figure 1 The communication between the client 1002 and the server 1004 can be carried out using a secure channel, such as a virtual private network (VPN), a secure HTTP session (e.g., using HTTPS or SSL protocol), or any other channel that is sufficient to protect any sensitive information that can be exchanged between the client 1002 and the server 1004 from eavesdropping by a third party. It is assumed that the server 1004 has access to sensitive information that should only be provided to authorized users. It is also assumed that the server 1004 uses voice-based authentication to determine whether a specific authorized user has made a request.
[0068] At block 1010, client 1002 sends a request involving sensitive information to server 1004. In some embodiments, client 1002 may determine that the request involves sensitive information and alert server 1004 to this fact; in other embodiments, server 1004 may make the determination. For example, the request may be a request to check a bank account balance.
[0069] At block 1012, in response to the request involving sensitive information, the server 1004 may generate a random number for acoustic feature generation (e.g., as described above). At block 1014, the server 1004 may send the random number to the client 1002. In some embodiments, the server 1004 may send the random number along with a request for the client 1002 to prompt the user to enter a passphrase or other input for a voice-based authentication process.
[0070] At box 1016, the client 1002 may receive the random number. At box 1018, the client 1002 may generate an acoustic feature based on the random number, for example, according to the above process 600. At box 1020, the client 1002 may play the acoustic feature into the environment (e.g., using speaker 110) while recording the sound from the environment (e.g., using microphone 112). At box 1022, the client 102 may send the recording to the server 1004. The recording is expected to include a mixture of the acoustic feature and the user's voice, which occurs in the audio channel. It should be understood that the client 1002 may perform some processing operations on the recording before sending it to the server 1004, such as verifying whether the speech component exists.
[0071] At block 1024, the server 1004 may perform a verification operation on the received recording. The verification operation may include verifying the acoustic features (e.g., using Fig. 9 900) and performing voice-based authentication after verifying the acoustic features. If the verification operation is successful, at box 1026, the server 1004 can process the user's request. In some embodiments, the processed request is received at box 1010; the recording for authentication does not need to include any specific request. At box 1028, the server 1004 can send a response to the request to the client 1002. If the verification at box 1024 is unsuccessful, the response can be an error response or other response indicating that the request is rejected.
[0072] At block 1030 , the client 1002 may receive a response from the server, and at block 1032 , the client 1002 may play an audible response for the user.
[0073] It should be understood that process 1000 can be performed repeatedly. In each case where a request involving sensitive information is received, server 1004 can generate a new random number (e.g., using a random process) so that different transactions use different acoustic features. Previous random numbers can be stored, such as described above, and the verification operation at block 1024 can include detecting the presence of an acoustic feature based on the previous random number (e.g., as described above with reference to process 900), which can indicate a replay attack.
[0074] It should also be understood that Fig.10 The division of activities between the client and the server shown in is illustrative and modifiable. For example, speech processing activities may be performed in part by the client and in part by the server. The system component that performs verification of the acoustic feature may be the same component (or under common control with) that generates the random number that defines the acoustic feature, so that the system component that performs verification of the acoustic feature does not have to rely on an independent device to accurately report what acoustic feature was used.
[0075] Although the foregoing description refers to specific embodiments, those skilled in the art will appreciate that the description is not exhaustive of all embodiments. Many variations and modifications are possible. For example, the acoustic features may include any number of frequency components and modulation schemes. The frequency components may be selected based on general characteristics of human speech or based on characteristics of a particular user's speech. For example, where the acoustic features are used in conjunction with voice-based authentication, the spectrum of the authorized user's voice may be known, and the acoustic features may include frequency components within the spectrum of the authorized user's voice, which may make it more difficult for an intruder to filter out the acoustic features from a recording.
[0076] The process of determining the acoustic feature based on the random number can also be changed. As in the above example, the random number can be used to determine the frequency and / or modulation pattern of the acoustic feature. The total duration of the acoustic feature can be long or short as needed. In some embodiments, the acoustic feature is long enough that it can cover the entire expected speech input. In other embodiments, the acoustic feature can be shorter than the expected speech input and can be played in a loop or in one or more bursts while the user speaks.
[0077] As described above, the use of acoustic features can be selectively triggered based on specific user requests. The set of requests that trigger the use of acoustic features may depend on the specific capabilities of a given voice-enabled client and / or the requirements of the server that processes voice-based requests. Some voice-enabled computer systems may operate on a transaction model, where a user initiates a request, and in executing the request, multiple back-and-forth exchanges with a voice-enabled device may occur. The acoustic feature can be used for any or all of the back-and-forth exchanges in a transaction, depending on the desired level of security. In some embodiments, the acoustic feature can be played when the user speaks the passphrase. Alternatively, the passphrase can be recorded without playing the acoustic feature, and the acoustic feature can be applied during one or more other exchanges during the transaction.
[0078] Any computer system mentioned herein can use any suitable number of subsystems. In some embodiments, the computer system includes a single computer device, wherein the subsystem can be a component of the computer device. In other embodiments, the computer system can include multiple computer devices, each of which is a subsystem with internal components.
[0079] A computer system may include multiple components or subsystems connected together, for example, by external interfaces or by internal interfaces. In some embodiments, the computer system, subsystems, or devices may communicate via a network. In such cases, one computer may be considered a client and another computer may be considered a server, where each computer may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.
[0080] It should be understood that any embodiment of the present invention can be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or a field programmable gate array) and / or using computer software, wherein a general-purpose programmable processor is modular or integrated. As used herein, a processor includes a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and the teachings provided herein, a person of ordinary skill in the art will know and understand other ways and / or methods of implementing embodiments of the present invention using hardware and combinations of hardware and software.
[0081] Any software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language such as Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission, suitable media including random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disk (CD) or a digital versatile disk (DVD), flash memory, etc. The computer-readable medium may be any combination of such storage or transmission devices.
[0082] Such programs can also be encoded and transmitted using carrier signals adapted to be transmitted via wired, optical and / or wireless networks that meet multiple protocols including the Internet. Therefore, the computer-readable medium according to an embodiment of the present invention can be created using data signals encoded with such programs. The computer-readable medium encoded with program code can be encapsulated with compatible devices or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium can reside on or in a single computer product (e.g., hard drive, CD or entire computer system), and can be present on or in different computer products in a system or network. The computer system may include a monitor, printer or other suitable display for providing any result mentioned herein to the user.
[0083] Any method described herein can be performed completely or in part with a computer system comprising one or more processors that can be configured to perform these steps. Therefore, an embodiment may relate to a computer system that is configured to perform the steps of any method described herein, may have different components that perform corresponding steps or corresponding step groups. Although presented with numbered steps, the method steps herein may also be performed simultaneously or in different orders. In addition, the part of these steps can be used together with the part of other steps from other methods. Equally, all or part of a step can be optional. In addition, any step of any method can be performed with module, circuit or other means for performing these steps.
[0084] Without departing from the spirit and scope of the embodiments of the invention, the specific details of the specific embodiments may be combined in any suitable manner. However, other embodiments of the invention may be directed to specific embodiments related to each individual aspect, or specific combinations of these individual aspects.
[0085] Unless expressly indicated to the contrary, the recitation of "a / an" or "the" is intended to mean "one or more." Unless expressly indicated to the contrary, the use of "or" is intended to mean an "inclusive or" rather than an "exclusive or."
[0086] All patents, patent applications, publications, and descriptions mentioned herein are incorporated by reference in their entirety for all purposes. No admission is made that they are prior art.
[0087] The above description is illustrative and not restrictive. After reading this disclosure, many variations of the present invention will become apparent to those skilled in the art. Therefore, the scope of patent protection should not be determined with reference to the above description, but should be determined with reference to the attached claims and their full scope or equivalent.
Claims
1. A method performed by a speech-enabled computer system, the method comprising: Get the current random number; Operating a speaker of the voice-enabled computer system to generate a current acoustic signature based on the current random number, comprising: defining a frequency hopping sequence for the current acoustic signature based on a first portion of the current random number, wherein the frequency hopping sequence includes frequencies within a range that overlaps with a vocalization range of an authorized user; defining a modulation input based on a second portion of the current random number; and generating a sound at the speaker based on the frequency hopping sequence and the modulation input; operating a microphone of the speech-enabled computer system while generating the current acoustic signature to record audio including the current acoustic signature and speech input; and authenticating the recorded audio based at least in part on the current acoustic characteristics; and In the event that the recorded audio is verified: processing the recorded audio to extract the speech input; processing the speech input to determine an action to be performed; and Perform the action.
2. The method of claim 1, wherein verifying the recorded audio comprises: determining whether the recorded audio includes the current acoustic feature; as well as determining whether the recorded audio includes different acoustic features based on the previous random number, Wherein, in a case where the recorded audio includes the current acoustic feature but does not include the different acoustic feature, the recorded audio passes the verification.
3. The method according to claim 1, further comprising: The user is prompted to speak so that the user speaks while the current acoustic feature is being generated. 4 . The method of claim 1 , wherein the speech input comprises a passphrase, and processing the speech input comprises performing voice-based authentication on the speech input.
5. The method of claim 1, wherein processing the recorded audio comprises filtering out the current acoustic feature from the recorded audio using the current random number.
6. A server computer, comprising: processor; Memory; as well as A computer readable medium coupled to the processor, the computer readable medium having code stored therein, the code executable by the processor to implement a method, the method comprising: Generate random numbers; providing the random number to a voice-enabled client having a speaker and a microphone and capable of generating an acoustic feature based on the random number; receiving an audio recording from the voice-enabled client; authenticating the audio recording based at least in part on detecting the acoustic signature in the audio recording, wherein the acoustic signature is based on an expected frequency hopping sequence that is based on a first portion of the random number and includes frequencies within a range that overlaps with a range of utterances of an authorized user, and an expected modulation input that is based on a second portion of the random number; and The user utterance in the audio recording is processed only if the audio recording passes the verification.
7. The server computer of claim 6, wherein the method further comprises adding the random number to a list of previous random numbers after verifying the audio recording.
8. The server computer of claim 7, wherein verifying the audio recording further comprises: determining whether old acoustic features are present in the audio recording based on the previous list of random numbers, If the old acoustic features exist, the audio recording fails the verification.
9. The server computer of claim 6, wherein processing the user utterance in the audio recording comprises filtering out the acoustic features.
10. The server computer of claim 9, wherein processing the user utterance in the audio recording comprises performing speech-based identity verification after filtering out the acoustic features.
11. The server computer of claim 6, wherein providing the random number to the voice-enabled client is performed in response to a request from the voice-enabled client involving sensitive information.
12. A speech-enabled computer system comprising: processor; Memory; a speaker operable by the processor; a microphone operable by the processor; as well as A computer readable medium coupled to the processor, the computer readable medium having code stored therein, the code executable by the processor to implement a method, the method comprising: Get a random number; Operating the speaker to generate an acoustic signature based on the random number, comprising: defining a frequency hopping sequence for the current acoustic signature based on a first portion of the current random number, wherein the frequency hopping sequence includes frequencies within a range overlapping a vocalization range of an authorized user; defining a modulation input based on a second portion of the current random number; and generating a sound at the speaker based on the frequency hopping sequence and the modulation input; and While generating the acoustic feature, the microphone is operated to record an audio recording including the acoustic feature and speech input.
13. The voice-enabled computer system of claim 12, wherein the method further comprises: authenticating the audio recording based at least in part on detecting the acoustic feature; and The speech input is processed only if the audio recording passes verification.
14. The voice-enabled computer system of claim 12, wherein the random number is obtained from a server computer, and the method further comprises: sending the audio recording to the server computer for processing; receiving a response from the server computer; as well as The response is provided to the user.
15. The voice-enabled computer system of claim 12, wherein obtaining the random number and operating the speaker to generate the acoustic signature are performed in response to receiving a user request for voice-based identity authentication.
16. The voice-enabled computer system of claim 12, wherein the method further comprises: detecting a spoken input corresponding to an activation phrase of the speech-enabled computer system, Wherein obtaining the current random number and operating the speaker are performed in response to detecting the speech input corresponding to the activation phrase.
Citation Information
Patent Citations
Voice biometrics systems and methods
US20190311722A1