Method and System for Speech Detection

By combining air-conducting speech and bone-conducting speech signals in wearable devices, the problem of distinguishing between user-spontaneous speech and non-user-spontaneous speech is solved, and more accurate speech recognition and security is achieved.

CN113767431BActive Publication Date: 2025-06-13CIRRUS LOGIC INT SEMICON LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080031842.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-30
Filing Date
2020-05-26
Publication Date
2025-06-13
Estimated Expiration
2040-05-26

AI Technical Summary

Technical Problem

The prior art is difficult to determine whether the person wearing wearable accessories is the speaker, especially in speech recognition systems, and it is necessary to distinguish between user-spontaneous voice and non-user-spontaneous voice.

Method used

The bone conduction speech signal is detected by detecting the air-conducting speech signal using the first microphone of the device and the bone conduction sensor, and filtering these signals to obtain a component of speech clarity. The components of the two signals are then compared, and if the difference exceeds the threshold, it is determined that the voice is not generated by the device user.

Benefits of technology

It effectively distinguishes spontaneous voice from non-spontaneous voice for people wearing wearable devices, improves the accuracy and security of the speech recognition system, and prevents illegal voice interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113767431B_ABST
    Figure CN113767431B_ABST
Patent Text Reader

Abstract

A method for a user of a device to detect their own voice is provided. A first microphone of the device is used to detect a first signal, and the first signal represents air-conducted voice. A bone conduction sensor of the device is used to detect a second signal, and the second signal represents bone-conducted voice. The first signal is filtered to obtain a component of the first signal at voice clarity, and the second signal is filtered to obtain a component of the second signal at the voice clarity. The component of the first signal at the voice clarity is compared with the component of the second signal at the voice clarity, and if the difference between the component of the first signal at the voice clarity and the component of the second signal at the voice clarity exceeds a threshold, it is determined that the voice is not generated by the user of the device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to voice detection and, in particular, to detecting when a speaker is a person using a device, such as a person wearing a wearable accessory such as headphones. Background of the Invention

[0003] Wearable accessories such as headphones, smart glasses, and smart watches are common. There are situations where, when voice is detected, it is desirable to know whether the person speaking is the person wearing a particular accessory. For example, when an accessory is used in combination with a device such as a smart phone having a voice recognition function, it is useful to know whether the detected voice is spoken by the person wearing the accessory. In such a case, the voice spoken by the person wearing the accessory can be provided to the voice recognition function so that any spoken commands can be executed, while in some cases, voices not spoken by the person wearing the accessory can be ignored. Summary of the Invention

[0004] According to embodiments described herein, a method and system are provided for reducing or avoiding one or more of the above disadvantages.

[0005] According to a first aspect of the present invention, a method for detecting a user's own voice for a device is provided, the method comprising:

[0006] Detecting a first signal using a first microphone of the device, the first signal representing air-conducted voice;

[0007] Detecting a second signal using a bone conduction sensor of the device, the second signal representing bone-conducted voice;

[0008] Filtering the first signal to obtain a component of the first signal at voice clarity;

[0009] Filtering the second signal to obtain a component of the second signal at the voice clarity;

[0010] Comparing the component of the first signal at the voice clarity with the component of the second signal at the voice clarity; and

[0011] If a difference between the component of the first signal at the voice clarity and the component of the second signal at the voice clarity exceeds a threshold, determining that the voice is not generated by the user of the device.

[0012] According to a second aspect of the present invention, a system for detecting a user's own voice for a device is provided, the system comprising:

[0013] An input for receiving a first signal representing air-conducted speech from a first microphone of the device and a second signal representing bone-conducted speech from a bone-conduction sensor of the device;

[0014] At least one filter for filtering the first signal to obtain a component of the first signal at speech intelligibility and filtering the second signal to obtain a component of the second signal at the speech intelligibility;

[0015] A comparator for comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility; and

[0016] A processor for determining that the speech is not generated by the user of the device if a difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold.

[0017] According to a third aspect of the present invention, there is provided a method for detecting a spoofing attack on a speaker recognition system, the method comprising:

[0018] Detecting a first signal using a first microphone of a device, the first signal representing air-conducted speech;

[0019] Detecting a second signal using a bone-conduction sensor of the device, the second signal representing bone-conducted speech;

[0020] Filtering the first signal to obtain a component of the first signal at speech intelligibility;

[0021] Filtering the second signal to obtain a component of the second signal at the speech intelligibility;

[0022] Comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility;

[0023] Determining that the speech is not generated by the user of the device if a difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold; and

[0024] Performing speaker recognition on the first signal representing speech if the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility does not exceed the threshold.

[0025] According to a fourth aspect of the present invention, there is provided a speaker recognition system, comprising:

[0026] an input for receiving a first signal representing air-conducted speech from a first microphone of the device and a second signal representing bone-conducted speech from a bone conduction sensor of the device;

[0027] at least one filter for filtering the first signal to obtain a component of the first signal at speech intelligibility and filtering the second signal to obtain a component of the second signal at the speech intelligibility;

[0028] a comparator for comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility;

[0029] a processor for determining that the speech is not generated by a user of the device if a difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold; and

[0030] a speaker recognition block for performing speaker recognition on the first signal representing the speech if the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility does not exceed the threshold. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] For a better understanding of the present invention and to more clearly show how the present invention may be practiced, reference is now made, by way of example only, to the accompanying drawings, in which:

[0032] Figure 1 is a schematic diagram of an electronic device and related accessories;

[0033] Figure 2 is a further schematic diagram of an electronic device and accessories;

[0034] Figure 3 is a flowchart showing a method;

[0035] Figure 4 shows a wearable device;

[0036] Figure 5 is a block diagram showing a part of the system as described herein;

[0037] Figure 6 is showing Figure 5 a block diagram of a part of the system;

[0038] Figure 7 shows the stages in the method of Figure 3 ;

[0039] Figure 8 further shows the stages in the method of Figure 3 ;

[0040] Figure 9 is a block diagram showing one form of the system as described herein;

[0041] Figure 10 shows the stages in the use of the system of Figure 9 ;

[0042] Figure 11 shows the stages in the use of the system of Figure 9 ;

[0043] Figure 12 is a block diagram showing an alternative form of the system as described herein; and

[0044] Figure 13 is a block diagram showing an alternative form of the system as described herein. DETAILED DESCRIPTION

[0045] The following description sets forth exemplary embodiments in accordance with the present disclosure. Other exemplary embodiments and implementations will be apparent to those of ordinary skill in the art. Additionally, those of ordinary skill in the art will recognize that various equivalent techniques may be applied in lieu of or in combination with the embodiments discussed below, and all such equivalents will be considered to be covered by the present disclosure.

[0046] For clarity, it will be noted here that this specification refers to speaker recognition and speech recognition, which are intended to have different meanings. Speaker recognition refers to a technique that provides information about the identity of the speaker. For example, speaker recognition can determine the identity of the speaker from a set of previously registered individuals, or can provide information indicating whether the speaker is a particular individual for identification or authentication purposes. Speech recognition refers to a technique for determining what is said and / or the meaning rather than identifying the speaker.

[0047] Figure 1An apparatus according to an aspect of the present invention is shown. The apparatus can be of any suitable type, such as a mobile computing device (e.g., laptop or tablet computer), a gaming console, a remote control device, a home automation controller including a home temperature or lighting control system or a household appliance, a toy, a machine such as a robot, an audio player, a video player, etc. However, in this illustrative example, the apparatus is a mobile phone, and specifically a smart phone 10 having a microphone 12 for detecting sound. The smart phone 10 can be used as a control interface for controlling any other additional device or system through suitable software.

[0048] Figure 1 An accessory is also shown, in this case a wireless headset 30, which in this example is in the form of in-ear headphones or earbuds. The headset 30 can be one of a pair of headphones or part of a headset, or can be used alone. A wireless headset 30 is shown, but a headset wired to the device can also be used. The accessory can be any suitable device that can be worn by a person and used in combination with the device. For example, the accessory can be a smart watch or a pair of smart glasses.

[0049] Figure 2 is a schematic diagram showing the form of the smart phone 10 and the wireless headset 30.

[0050] Specifically, Figure 2 various interconnecting components of the smart phone 10 and the wireless headset 30 are shown. It should be understood that the smart phone 10 and the wireless headset 30 will actually include many other components, but the following description is sufficient for understanding the present invention. Additionally, it should be understood that components similar to those Figure 2 shown can be included in any suitable device and any suitable wearable accessory.

[0051] Therefore, Figure 2 it is shown that the smart phone 10 includes the aforementioned microphone 12.

[0052] Figure 2 A memory 14 is also shown, which can actually be provided as a single component or multiple components. The memory 14 is provided for storing data and program instructions.

[0053] Figure 2 An accessory 16 is also shown, which can actually again be provided as a single component or multiple components. For example, one component of the processor 16 can be the application processor of the smart phone 10.

[0054] Therefore, the memory 14 can act as a tangible computer-readable medium storing code for causing the processor 16 to execute the method described below.

[0055] Figure 2 Also shown is a transceiver 18, which is provided to allow the smart phone 10 to communicate with an external network. For example, the transceiver 18 may include circuitry for establishing an Internet connection via a WiFi local area network or via a cellular network.

[0056] In addition, the transceiver 18 allows the smart phone 10 to communicate with the wireless headset 30, for example, using Bluetooth or another short-range wireless communication protocol.

[0057] Figure 2 Also shown is an audio processing circuit 20 for performing operations on the audio signals detected by the microphone 12 as needed. For example, the audio processing circuit 20 may filter the audio signals or perform other signal processing operations.

[0058] Figure 2 The wireless headset 30 is shown to include a transceiver 32, which allows the wireless headset 30 to communicate with the smart phone 10, for example, using Bluetooth or another short-range wireless communication protocol.

[0059] Figure 2 The wireless headset 30 is also shown to include a first sensor 34 and a second sensor 36, which will be described in more detail below. Signals generated by the first sensor 34 and the second sensor 36 in response to an external stimulus are transmitted to the smart phone 10 via the transceiver 32.

[0060] Thus, in the illustrated embodiment, signals are generated by sensors on the accessory, and these signals are transmitted to the host device, where the signals are processed. In other embodiments, the signals are processed on the accessory itself.

[0061] In this embodiment, the smart phone 10 is provided with a speaker recognition function and a control function. Thus, the smart phone 10 is capable of performing various functions in response to an oral command from a registered user. The speaker recognition function is capable of distinguishing an oral command from a registered user from the same command spoken by a different person. Thus, certain embodiments of the present invention relate to the operation of a smart phone or another portable electronic device (such as a tablet computer or a laptop computer, a gaming console, a home control system, a home entertainment system, an in-vehicle entertainment system, a household appliance, etc.) having a certain voice operability, where the speaker recognition function is performed in the device intended to execute the oral command. Certain other embodiments relate to a system in which the speaker recognition function is performed on a smart phone or other device, and if the speaker recognition function can confirm that the speaker is a registered user, then the smart phone or other device then transmits the command to a separate device.

[0062] In some embodiments, when performing a speaker recognition function on the smart phone 10 or other devices located near the user, a transceiver 18 is used to transmit an oral command to a remote speech recognition system that determines the meaning of the oral command. For example, the speech recognition system may be located on one or more remote servers in a cloud computing environment. A signal based on the meaning of the oral command then returns to the smart phone 10 or other local device.

[0063] When an oral command is received, it is generally necessary to perform a speaker verification process to confirm that the speaker is an enrolled user of the system. It is known to perform speaker verification by first performing a process of enrolling the user to obtain a user voice model. Then, when it is desired to determine whether a particular test input is the voice of the user, a first score is obtained by comparing the test input with the user voice model. Additionally, a score normalization process may be performed. For example, the test input may also be compared with multiple voice models obtained from multiple other speakers. These comparisons give multiple group scores, and statistics describing the multiple group scores may be obtained. The statistics may then be used to normalize the first score to obtain a normalized score, and the normalized score may be used for speaker verification.

[0064] In the embodiments described herein, it may be determined whether the detected speech is spoken by a person wearing the accessory 30. This is referred to as "own voice detection". If it is determined that the detected speech is spoken by a person wearing the accessory 30, a voice signal may be sent for speaker recognition and / or speech recognition. If it is determined that the detected speech is not spoken by a person wearing the accessory 30, it may be determined that the voice signal should not be sent for speaker recognition and / or speech recognition.

[0065] Figure 3 is a flowchart showing an example of a method according to the present disclosure, specifically a method for own voice detection of a user wearing a wearable device (i.e., a method for detecting whether a person wearing a wearable device is the speaking person).

[0066] Essentially the same method may be used to detect whether a person holding a handheld device such as a mobile phone is the speaking person.

[0067] Figure 4 The form of the wearable device is shown in more detail in one embodiment. Specifically, Figure 4 it shows the earphone 30 worn in the ear canal 70 of the wearer.

[0068] In this example, Figure 2 the first sensor 34 shown therein takes the form of an ear external microphone 72, i.e., a microphone that detects acoustic signals in the air around the wearer's ear. In this example, Figure 2The second sensor 36 shown in FIG. takes the form of a bone conduction sensor 74. This can be an in-ear microphone that is capable of detecting an acoustic signal in the wearer's ear canal, the acoustic signal being generated by the wearer's voice and transmitted through the bones of the wearer's head, and the in-ear microphone may also be capable of detecting vibrations of the wearer's ear canal itself. Alternatively, the bone conduction sensor 74 can be an accelerometer that is positioned such that it contacts the wearer's ear canal and can detect contact vibrations caused by the wearer's voice and transmitted through the bones and / or soft tissues of the wearer's head.

[0069] Similarly, if the wearable device is a pair of smart glasses, the first sensor can take the form of an external directional microphone that picks up the wearer's voice (and other sounds) through air conduction, while the second sensor can take the form of an accelerometer that is positioned to contact the wearer's head and can detect contact vibrations caused by the wearer's voice and transmitted through the bones and / or soft tissues of the wearer's head.

[0070] Similarly, if the wearable device is a smartwatch, the first sensor can take the form of an external directional microphone that picks up the wearer's voice (and other sounds) through air conduction, while the second sensor can take the form of an accelerometer that is positioned to contact the wearer's wrist and can detect contact vibrations caused by the wearer's voice and transmitted through the bones and / or soft tissues of the wearer.

[0071] If the method is used to detect whether a person holding a handheld device such as a mobile phone is the speaker, the first sensor can be a microphone 12 that picks up the wearer's voice (and other sounds) through air conduction, while the second sensor can take the form of an accelerometer that is positioned inside the handheld device (and thus not visible in Figure 1 ), and the accelerometer can detect contact vibrations caused by the user's voice and transmitted through the bones and / or soft tissues of the wearer. When the handheld device is pressed against the user's head, the accelerometer can detect contact vibrations caused by the wearer's voice and transmitted through the bones and / or soft tissues of the user's head, while when the handheld device is not pressed against the user's head, the accelerometer can detect contact vibrations caused by the wearer's voice and transmitted through the bones and / or soft tissues of the user's arm and hand.

[0072] In Figure 3 step 50 of, the method then includes using the first microphone 72 of the wearable device 30 to detect a first signal representative of air-conducted speech. That is, when the wearer speaks, the sound leaves their mouth and travels through the air and can be detected by the microphone 72.

[0073] InFigure 3 In step 52 of, the method further includes using the bone conduction sensor 74 of the wearable device 30 to detect a second signal indicative of bone conduction speech. That is, when the wearer speaks, vibrations are conducted through the bones of their head (and / or at least to some extent can be conducted through the surrounding soft tissues) and can be detected by the bone conduction sensor 74.

[0074] In principle, the self-speech detection process can be achieved by comparing the signals generated by the microphone 72 and the bone conduction sensor 74.

[0075] However, since a typical bone conduction sensor 74 is based on an accelerometer and accelerometers typically operate at a low sampling rate (in the range of 100 Hz - 1 kHz), the signal generated by the bone conduction sensor 74 has a limited frequency range. Additionally, typical bone conduction sensors are prone to picking up contact noise (such as noise generated by the wearer's head turning or by contact with other objects). Furthermore, it has been found that audible speech is typically transmitted more effectively via bone conduction than inaudible speech.

[0076] Therefore, Figure 5 A filter circuit is shown for filtering the signals generated by the microphone 72 and the bone conduction sensor 74 such that the signals are more useful for the purpose of self-speech detection.

[0077] Specifically, Figure 5 A signal from the first sensor 72 received at the first input 90 is shown (i.e., the first signal mentioned at step 50 of Figure 3 ), and a signal from the second sensor 72 received at the second input 92 is shown (i.e., the second signal mentioned at step 52 of Figure 3 ).

[0078] As described above, it can be expected that audible speech is transmitted more effectively via bone conduction than inaudible speech. Given this, it is expected that the second signal will contain significant signal content only during periods when the wearer's speech contains audible speech.

[0079] Therefore, the first signal received at the first input 90 is passed to the audible speech detection block 94. When it is determined that the first signal represents audible speech, this generates a flag.

[0080] Voiced speech can be identified by, for example, the following: using a deep neural network (DNN) trained according to a gold reference, such as with Praat software; performing autocorrelation with a unit delay on the speech signal (since voiced speech has higher autocorrelation for non-zero lags); performing linear predictive coding (LPC) analysis (since the initial reflection coefficients are a good indicator of voiced speech); looking at the zero-crossing rate of the speech signal (since unvoiced speech has a higher zero-crossing rate); looking at the short-term energy of the signal (which tends to be higher for voiced speech); tracking the first formant frequency F0 (since unvoiced speech does not contain the first formant frequency); examining the error in linear predictive coding (LPC) analysis (since the LPC prediction error is lower for voiced speech); using automatic speech recognition to identify the words being spoken, and thus classifying the speech into voiced and unvoiced speech; or fusing any or all of the above.

[0081] In Figure 3 In step 54 of the method of , the first signal is filtered to obtain the component of the first signal at speech intelligibility. Thus, the first signal received at the first input 90 is passed to the first intelligibility filter 96.

[0082] In Figure 3 In step 56 of the method of , the second signal is filtered to obtain the component of the second signal at speech intelligibility. Thus, the second signal received at the second input 92 is passed to the second intelligibility filter 98.

[0083] Figure 6 is a schematic diagram showing in more detail the form of the first intelligibility filter 96 and the second intelligibility filter 98.

[0084] In each case, the corresponding input signal is passed to a low-pass filter 110, the cut-off frequency of which can be in the range of, for example, 1 kHz. The low-pass filtered signal is passed to an envelope detector 112 for detecting the envelope of the filtered signal. The resulting envelope signal can optionally be passed to a decimator 114 and then to a band-pass filter 116, which is tuned to allow signals with typical intelligibility to pass. For example, the band-pass filter 116 can have a passband between 5 Hz and 15 Hz or between 5 Hz and 10 Hz.

[0085] Thus, the intelligibility filters 96, 98 detect the power modulated at frequencies corresponding to the typical intelligibility of speech, where intelligibility is the rate at which the speaker speaks, which can be measured, for example, as the rate at which the speaker produces different speech or makes phone calls.

[0086] In Figure 3In step 58 of the method, the component of the first signal in speech intelligibility is compared with the component of the second signal in speech intelligibility. Thus, as Figure 5 shown, the outputs of the first intelligibility filter 96 and the second intelligibility filter 98 are passed to the comparison and determination block 100.

[0087] In Figure 3 step 60 of the method, if the difference between the component of the first signal in speech intelligibility and the component of the second signal in speech intelligibility exceeds a threshold, it is determined that the speech is not generated by the user wearing the wearable device. Conversely, if the difference between the component of the first signal in speech intelligibility and the component of the second signal in speech intelligibility does not exceed the threshold, it can be determined that the speech is generated by the user wearing the wearable device.

[0088] More specifically, in an embodiment including the audible speech detection block 94, the comparison may be performed only when a flag indicating that the first signal represents audible speech is generated. The entire filtered first signal may be passed to the comparison and determination block 100, where the comparison is performed only when a flag indicating that the first signal represents audible speech is generated, or the inaudible speech may be rejected, and only those segments of the filtered first signal that represent audible speech may be passed to the comparison and determination block 100.

[0089] Thus, processing the signals detected by the first sensor and the second sensor is such that when the wearer of the wearable device is the speaker, it is expected that the processed version of the first signal will be similar to the processed version of the second signal. In contrast, if the wearer of the wearable device is not the speaker, the first sensor can still detect air-conducted speech, but the second sensor will not be able to detect any bone-conducted speech, and thus it is expected that the processed version of the first signal will be very different from the processed version of the second signal.

[0090] Figure 7 And Figure 8 are examples in this regard, showing the amplitudes of the processed versions of the first signal and the second signal in these two cases. In each figure, the signal is divided into (for example) frames with a duration of 20 ms, and the amplitude during each frame period is plotted over time.

[0091] Specifically, Figure 7 shows the case where the wearer of the wearable device is the speaker, so the processed version of the first signal 130 is similar to the processed version of the second signal 132.

[0092] In contrast, Figure 8Illustrates a situation where the wearer of the wearable device is not the speaker. Thus, although the processed version of the first signal 140 contains components generated by air-conducted speech, it is quite different from the processed version of the second signal 142 because it does not contain any components generated by bone-conducted speech.

[0093] There are different methods for performing a comparison between the components of the first signal at speech clarity and the components of the second signal at speech clarity to determine whether the difference exceeds a threshold.

[0094] Figure 9 Is a schematic diagram showing the first form of the comparison and determination block 100 from Figure 5 where the processed version of the first signal (i.e., the component of the first signal at speech clarity) is generated by the first clarity filter 96 and passed to the first input 120 of the comparison and determination block 100. The processed version of the second signal (i.e., the component of the second signal at speech clarity) is generated by the second clarity filter 98 and passed to the second input 122 of the comparison and determination block 100.

[0095] Then the signal at the first input 120 is passed to block 124 where an empirical cumulative distribution function (ECDF) is formed. Similarly, the signal at the second input 122 is then passed to block 126 where an empirical cumulative distribution function (ECDF) is formed.

[0096] In each case, the ECDF is calculated frame by frame using the amplitude of the signal during the frame. Then, for each possible signal amplitude, the ECDF indicates the proportion of frames in which the actual signal amplitude is below that level.

[0097] Figure 10 Illustrates two ECDFs calculated in a situation similar to Figure 7 where the two signals are generally similar. It can thus also be seen that the two ECDFs (i.e., the ECDF 142 formed by the first signal and the ECDF 144 formed by the second signal) are generally similar.

[0098] One measure of the similarity between the two ECDFs 142, 144 is to measure the maximum vertical distance between them, which is d1 in this case. Alternatively, several measurements of the vertical distance between the two ECDFs can be made and then summed or averaged to arrive at a suitable measure.

[0099] Figure 11 Illustrates in a situation similar to Figure 8Two ECDFs calculated in the case where the two signals are significantly different. Thus, it can also be seen that the two ECDFs (i.e., ECDF 152 formed by the first signal and ECDF 154 formed by the second signal) are significantly different. Specifically, since the second signal does not contain any component generated by bone-conducted speech, its level is generally much lower than that of the first signal, and thus the form of the ECDF indicates that the amplitude of the second signal is generally lower.

[0100] In this case, the maximum vertical distance between ECDFs 152 and 154 is d2.

[0101] More generally, the step of comparing the components of the first signal and the second signal in terms of speech intelligibility may include forming corresponding first and second distribution functions from these components and calculating the value of the statistical distance between the second distribution function and the first distribution function.

[0102] For example, as described above, the value of the statistical distance between the second distribution function and the first distribution function can be calculated as:

[0103] d KS =max{|F1 - F2|}

[0104] where

[0105] F1 is the first distribution function, and

[0106] F2 is the second distribution function, and thus

[0107] │F1 - F2│ is the vertical distance between the two functions at a given frequency.

[0108] Alternatively, the value of the statistical distance between the second distribution function and the first distribution function can be calculated as:

[0109] d IN =∫|F1 - F2|df

[0110] where

[0111] F1 is the first distribution function, and

[0112] F2 is the second distribution function, and thus

[0113] │F1 - F2│ is the vertical distance between the two functions at a given frequency.

[0114] Alternatively, the value of the statistical distance between the second distribution function and the first distribution function can be calculated as:

[0115]

[0116] Or, more specifically, when p = 2:

[0117]

[0118] wherein

[0119] F1 is a first distribution function, and

[0120] F2 is a second distribution function, thus

[0121] │F1 - F2│ is the vertical distance between the two functions at a given frequency.

[0122] In other examples, the step of comparing components can use a machine learning system that has been trained to distinguish components generated from the wearer's own voice and the voice of a non - wearer.

[0123] Although examples are given here where the distribution function is a cumulative distribution function, other distribution functions (such as probability distribution functions) and appropriate methods for comparing these functions can also be used. The method for comparison can include using a machine learning system as described above.

[0124] Thus, returning to Figure 9 , the ECDF is passed to block 128, which calculates the statistical distance d between the ECDFs and passes it to block 130, where the statistical distance is compared with a threshold θ.

[0125] As discussed in step 60 of the method of reference Figure 3 , if the statistical distance between the ECDFs generated by the components of the first signal at speech intelligibility and the components of the second signal at speech intelligibility exceeds the threshold θ, it is determined that the speech is not generated by the user wearing the wearable device. Conversely, if the statistical distance between the ECDFs generated by the components of the first signal at speech intelligibility and the components of the second signal at speech intelligibility does not exceed the threshold θ, it can be determined that the speech is generated by the user wearing the wearable device.

[0126] Figure 12 is a schematic diagram showing a second form of the comparison and determination block 100 from Figure 5 , where the processed version of the first signal (i.e., the component of the first signal at speech intelligibility) is generated by the first intelligibility filter 96 and passed to the first input 160 of the comparison and determination block 100. The processed version of the second signal (i.e., the component of the second signal at speech intelligibility) is generated by the second intelligibility filter 98 and passed to the second input 162 of the comparison and determination block 100.

[0127] The signal at the first input 160 is subtracted from the signal at the second input 162 in the subtractor 164, and the difference Δ is passed to the comparison block 166. In some embodiments, the value of the difference Δ calculated in each frame is compared with a threshold. If the difference exceeds the threshold in any frame, it can be determined that the speech is not generated by the user wearing the wearable device. In other embodiments, statistical analysis is performed on the values of the difference Δ calculated over multiple frames. For example, the mean or moving mean of Δ calculated over a block of 20 frames can be compared with a threshold. As another example, the median of Δ calculated over a block of frames can be compared with a threshold. In any of these cases, if the parameter calculated from the individual differences exceeds the threshold, it can be determined that the speech (or at least the relevant portion of the speech from which the difference is calculated) is not generated by the user wearing the wearable device.

[0128] Figure 13 is a schematic block diagram of a system using the self-speech detection method described previously.

[0129] Specifically, Figure 13 shows a first sensor 200 and a second sensor 202 that can be disposed on the wearable device. As referenced Figure 4 described, the first sensor 200 can take the form of a microphone that detects acoustic signals transmitted through the air, while the second sensor 202 can take the form of an accelerometer for detecting signals transmitted through bone conduction (including through the wearer's soft tissue).

[0130] Signals from the first sensor 200 and the second sensor 202 are passed to the wear detection block 204, which compares the signals generated by the sensors and determines whether the wearable device is being worn at this time. For example, when the wearable device is an earphone, the wear detection block 204 can take the form of an in-ear detection block. In the case of an earphone, for example as Figure 4 shown, if the earphone is outside the user's ear, the signals detected by the sensors 200 and 202 are substantially the same, but if the earphone is worn, they are significantly different. Therefore, comparing the signals allows determination of whether the wearable device is being worn.

[0131] Other systems can be used to detect whether the wearable device is being worn, and these systems can use either or neither of the sensors 200, 202. For example, an optical sensor, a conductivity sensor, or a proximity sensor can be used.

[0132] Additionally, the wear detection block 204 can be configured to perform "liveness detection", i.e., determine whether the wearable device is being worn by a living person at this time. For example, the signals generated by the second sensor 202 can be analyzed to detect signs of the wearer's pulse to confirm that a wearable device such as headphones or a watch is being worn by a person, rather than placed in or on an inanimate object.

[0133] When the method is used to detect whether a person holding a handheld device such as a mobile phone is the speaking person, the detection block 204 can be configured to determine whether the device is being held by the user (rather than, for example, being used when placed on a table or other surface). For example, the detection block can receive signals from the sensor 202 or from one or more separate sensors, which can be used to determine whether the device is being held by the user and / or pressed against the user's head. For example, an optical sensor, a conductivity sensor, or a proximity sensor can be used again.

[0134] If it is determined that the wearable device is being worn, or the handheld device is being held, the signal from the detection block 204 is used to close the switches 206, 208, and the signals from the sensors 200, 202 are passed to Figure 5 inputs 90, 92 of the shown circuit.

[0135] As described in reference Figure 5 The comparison and determination block 100 generates an output signal indicating whether the detected speech is spoken by the person wearing the wearable device.

[0136] In this exemplary system, the signal from the first sensor 200 is also passed to the voice keyword detection block 220. The voice keyword detection block 220 detects a specific "wake word" that is used by the user of the device to wake the device from a low-power standby mode and place the device in a mode where voice recognition can be performed.

[0137] When the voice keyword detection block 220 detects a specific "wake word", the signal from the first sensor 200 is sent to the speaker recognition block 222.

[0138] If the output signal from the comparison and determination block 100 indicates that the detected speech is not spoken by the person wearing the wearable device, this is considered spoofing, and thus it is not desirable to perform voice recognition on the detected speech.

[0139] However, if the output signal from the comparison and determination block 100 indicates that the detected speech is spoken by the person wearing the wearable device, the speaker recognition block 222 may perform a speaker recognition process on the signal from the first sensor 200. Generally speaking, the speaker recognition process extracts features from the speech signal and compares them with the features of a model generated by registering known users into the speaker recognition system. If the comparison finds that the speech features are similar enough to the model, the speaker is determined to be the registered user with a high enough probability level.

[0140] When the wearable device is the earphone, the speaker recognition process may be omitted and replaced by an ear biometric process for identifying the wearer of the earphone. For example, this can be done by examining the signal generated by an in-ear microphone (which can also act as the second sensor 202) provided on the earphone and comparing the signal features with the acoustic model of the registered user's ear. If it can be determined with sufficient confidence for the relevant application that the earphone is worn by the registered user and the speech is the speech of the person wearing the earphone, this can serve as a form of speaker recognition.

[0141] If it is determined that the speaker is the registered user (either through the output from the speaker recognition block 222 or by using the confirmation that the wearable device is worn by the registered user), the signal is sent to the speech recognition block 224, which also receives the signal from the first sensor 200.

[0142] The speech recognition block 224 then performs speech recognition processing on the received signal. For example, if the speech recognition block detects that the speech contains a command, it may send an output signal to a further application to act on the command.

[0143] Therefore, the availability of the bone conduction signal can be used for the purpose of self-speech detection. In addition, the result of self-speech detection can be used to improve the reliability of the speaker recognition and speech recognition systems.

[0144] Those skilled in the art will recognize that some aspects of the above-described apparatus and methods may be embodied as processor control code, for example, located on a non-volatile carrier medium such as a magnetic disk, a CD- or DVD-ROM, a programmed memory such as a read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. For many applications, embodiments of the present invention will be implemented on a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), or an FPGA (Field Programmable Gate Array). Thus, the code may include conventional program code or microcode, or (for example) code for setting or controlling an ASIC or FPGA. The code may also include code for dynamically configuring a reconfigurable device such as a reprogrammable logic gate array. Similarly, the code may include code for a hardware description language such as Verilog TM or VHDL (Very High Speed Integrated Circuit Hardware Description Language). Those skilled in the art should understand that the code may be distributed among multiple coupled components that communicate with each other. In appropriate cases, code for running on a field (reprogrammable) programmable analog array or similar device to configure analog hardware may also be used to implement the embodiments.

[0145] Note that, as used herein, the term module should be used to refer to a functional unit or block that can be implemented at least in part by dedicated hardware components such as custom-defined circuits and / or at least in part by one or more software processors or suitable code running on a suitable general-purpose processor, etc. A module itself may include other modules or functional units. A module may be provided by multiple components or sub-modules, and the components or sub-modules need not be co-located, but may be arranged on different integrated circuits and / or run on different processors.

[0146] Embodiments may be implemented in a host device, especially a portable and / or battery-powered host device such as a mobile computing device (e.g., a laptop or tablet computer), a gaming console, a remote control device, a home automation controller including a home temperature or lighting control system or a household appliance, a toy, a machine such as a robot, an audio player, a video player, or a mobile phone (e.g., a smart phone).

[0147] It should be noted that the above embodiments illustrate rather than limit the present invention, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. The word "comprising" does not exclude the presence of elements or steps other than those listed in the claims, "a" or "an" does not exclude a plurality, and a single feature or other unit may perform the functions of several units recited in the claims. Any reference numerals or labels in the claims should not be construed as limiting their scope.

Claims

1. A method for detecting the user's own voice of a device, the method comprises: detecting a first signal using a first microphone of the device, the first signal representing air-conducted voice; detecting a second signal using a bone conduction sensor of the device, the second signal representing bone-conducted voice; filtering the first signal to obtain a component of the first signal at speech intelligibility; filtering the second signal to obtain a component of the second signal at the speech intelligibility; comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility; and if the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold, determining that the voice is not generated by the user of the device.

2. The method according to claim 1, which comprises: detecting a signal component in the first signal representing voiced speech; and comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility only during a time period when a signal component representing voiced speech exists in the first signal.

3. The method according to claim 1 or 2, which comprises: filtering the first signal in a first band-pass filter to obtain the component of the first signal at the speech intelligibility; and filtering the second signal in a second band-pass filter to obtain the component of the second signal at the speech intelligibility; wherein the first band-pass filter and the second band-pass filter have respective passbands including a frequency range of 5 Hz to 15 Hz.

4. The method according to claim 3, which further comprises: performing low-pass filtering on the first signal and the second signal before filtering the first signal in the first band-pass filter and filtering the second signal in the second band-pass filter, and detecting the envelope of each filtered signal.

5. The method according to claim 1 or 2, wherein: comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility comprises: forming a cumulative distribution function of the component of the first signal at the speech intelligibility; forming a cumulative distribution function of the component of the second signal at the speech intelligibility; and determining the difference between the cumulative distribution function of the component of the first signal at the speech intelligibility and the cumulative distribution function of the component of the second signal at the speech intelligibility, And wherein if the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold, it is determined that the speech is not generated by the user of the device, including if the difference between the cumulative distribution function of the component of the first signal at the speech intelligibility and the cumulative distribution function of the component of the second signal at the speech intelligibility exceeds a threshold, it is determined that the speech is not generated by the user of the device.

6. The method according to claim 5, wherein comprises: obtaining, in each of a plurality of frames, the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility; and forming the cumulative distribution function frame by frame using the amplitudes of the respective signals during the plurality of frames.

7. The method according to claim 1 or 2, wherein: comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility comprises: subtracting the component of the second signal at the speech intelligibility from the component of the first signal at the speech intelligibility; and wherein if the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold, it is determined that the speech is not generated by the user of the device, including if the result of subtracting the component of the second signal at the speech intelligibility from the component of the first signal at the speech intelligibility exceeds a threshold, it is determined that the speech is not generated by the user of the device.

8. The method according to claim 7, wherein comprises: obtaining, in each of a plurality of frames, the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility; subtracting the component of the second signal at the speech intelligibility from the component of the first signal at the speech intelligibility frame by frame; and forming the result of subtracting the component of the second signal at the speech intelligibility from the component of the first signal at the speech intelligibility from a plurality of values calculated frame by frame.

9. The method according to claim 1 or 2, wherein the device is a wearable device, and the user of the device is the wearer of the device.

10. The method according to claim 9, wherein the wearable device is an earphone, smart glasses or a smart watch.

11. The method according to claim 1 or 2, wherein the device is a handheld device.

12. The method according to claim 11, wherein the handheld device is a mobile phone.

13. A system for detecting a user's own voice of a device, the system comprises: an input for receiving a first signal representing air-conducted speech from a first microphone of the device and a second signal representing bone-conducted speech from a bone conduction sensor of the device; At least one filter, the at least one filter being configured to filter the first signal to obtain a component of the first signal at speech intelligibility and to filter the second signal to obtain a component of the second signal at the speech intelligibility; A comparator, the comparator being configured to compare the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility; And A processor, the processor being configured to determine that the speech is not generated by the user of the device if a difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold.

14. The system according to claim 13, wherein Comprises: Wherein the at least one filter comprises at least one band - pass filter for filtering the first signal and for filtering the second signal; And Wherein the at least one band - pass filter has a pass - band including a frequency range of 5 Hz to 15 Hz.

15. The system according to claim 14, wherein the at least one filter Comprises: At least one low - pass filter, the at least one low - pass filter being configured to perform low - pass filtering on the first signal and the second signal before filtering the first signal and the second signal in the at least one band - pass filter, and An envelope detector, the envelope detector being configured to detect the envelope of each filtered signal.

16. The system according to any one of claims 13 to 15, wherein the device is a wearable device, and the user of the device is the wearer of the device, and wherein the system is implemented in a device separated from the wearable device.

17. The system according to claim 16, wherein the wearable device is an earphone, smart glasses or a smart watch.

18. The system according to any one of claims 13 to 15, wherein the device is a handheld device, and the system is implemented in the device.

19. The system according to claim 18, wherein the handheld device is a mobile phone.

20. A method for detecting a spoofing attack on a speaker recognition system, the method Comprises: Detecting a first signal using a first microphone of a device, the first signal representing air - conducted speech; Detecting a second signal using a bone - conduction sensor of the device, the second signal representing bone - conducted speech; Filtering the first signal to obtain a component of the first signal at speech intelligibility; Filtering the second signal to obtain a component of the second signal at the speech intelligibility; Comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility; Determining that the speech is not generated by the user of the device if a difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold; And If the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility does not exceed the threshold, perform speaker recognition on the first signal representing speech.

21. The method according to claim 20, wherein comprises: determining whether the device is worn or held, and performing at least the step of comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility only when it is determined that the device is worn or held.

22. The method according to claim 20 or 21, further comprises: performing voice keyword detection on the first signal representing speech; and performing speaker recognition on the first signal representing speech only when a predetermined voice keyword is detected in the first signal representing speech.

23. The method according to claim 20 or 21, further comprises: if the result of performing speaker recognition is to determine that the speech is generated by a registered user, performing speech recognition on the first signal representing speech.

24. The method according to claim 20 or 21, wherein the device is a wearable device, and the user of the device is the wearer of the device.

25. The method according to claim 24, wherein the wearable device is an earphone, smart glasses or a smart watch.

26. The method according to claim 20 or 21, wherein the device is a handheld device.

27. The method according to claim 26, wherein the handheld device is a mobile phone.

28. A speaker recognition system, which comprises: an input for receiving a first signal representing air-conducted speech from a first microphone of a device and a second signal representing bone-conducted speech from a bone conduction sensor of the device; at least one filter for filtering the first signal to obtain a component of the first signal at speech intelligibility and filtering the second signal to obtain a component of the second signal at the speech intelligibility; a comparator for comparing the component of the first signal at the speech intelligibility with the component of the second signal at the speech intelligibility; a processor for determining that the speech is not generated by the user of the device if the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility exceeds a threshold; and a speaker recognition block for performing speaker recognition on the first signal representing the speech if the difference between the component of the first signal at the speech intelligibility and the component of the second signal at the speech intelligibility does not exceed the threshold.

29. The system according to claim 28, wherein comprises: a detection circuit for determining whether the device is worn or held, Wherein the system is configured to perform at least the step of comparing the component of the first signal at the speech clarity with the component of the second signal at the speech clarity only when it is determined that the device is worn or held.

30. The system according to claim 28 or 29, further comprising: a voice keyword detection block for receiving the first signal representing speech, wherein the system is configured to perform speaker recognition on the first signal representing speech only when a predetermined voice keyword is detected in the first signal representing speech.

31. The system according to claim 28 or 29, wherein the system is configured to perform speech recognition on the first signal representing speech if the result of performing speaker recognition is to determine that the speech is generated by a registered user.

32. The system according to claim 28 or 29, wherein the device is a wearable device, and the user of the device is the wearer of the device.

33. The system according to claim 32, wherein the wearable device is an earphone, smart glasses or a smart watch.

34. The system according to claim 28 or 29, wherein the device is a handheld device.

35. The system according to claim 34, wherein the handheld device is a mobile phone.

Citation Information

Patent Citations

  • Head mounted multi-sensory audio input system

    CN1591568A

  • Ultrasonic and multimodality assisted hearing

    US20100040249A1