Inertial sensor unit and method for detecting speech activity
By employing two-stage signal processing and variable operating modes with an inertial sensor unit, the problems of misidentification and high energy consumption in microphone systems in noisy environments are solved, achieving low-energy and high-efficiency speech activity detection.
Patent Information
- Application Number
- CN202110743348.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-01
- Filing Date
- 2021-07-01
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2041-07-01
AI Technical Summary
In the existing technology, microphone-based speech recognition systems are prone to misidentification in noisy environments, and speech recognition algorithms of battery-powered devices consume a lot of energy, making it difficult to achieve efficient and energy-saving speech activity detection.
An inertial sensor unit is used to detect speech activity through a two-stage signal processing method. The first stage is a simple threshold comparison, and the second stage is a more complex analysis and processing. Combined with variable operating modes and component activation strategies, energy consumption is reduced.
It achieves efficient recognition of language activities under low energy consumption conditions, reduces misrecognition, and improves the system's power saving and recognition accuracy.
Smart Images

Figure CN113884176B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to an inertial sensor unit and a method for detecting speech activity by means of an inertial sensor unit. The invention particularly relates to an inertial sensor unit for a head-wearable device. BACKGROUND
[0002] In the context of speech activity recognition, acceleration sensors (accelerometers) can be used in order to improve the quality of speech recognition. For example, the signals of acceleration sensors can be used in order to improve the signal-to-noise ratio or in order to perform an automatic gain adjustment.
[0003] From US 2017 / 263267 A1 a system and a method for performing an automatic gain adjustment in the case of using an acceleration sensor in a headset is known. Here, a speech signal is detected by analyzing the acceleration signal of the acceleration sensor. Here, in a first step, a signal pre-processing by means of a high-pass filter and a low-pass filter is performed. In a second step, the acceleration signal is analyzed by a threshold comparison of the absolute amplitude or of an extracted envelope curve. The speech recognition can be performed by a threshold comparison of the correlation of the acceleration signal with respect to two axes within a short time window.
[0004] US 2013 / 196715 A1 relates to an adapted noise suppression for speech activity recognition.
[0005] From US 10397687 B2 a signal processing device for earbud headset speech recognition is known. Here, speech features are derived based on the signals of acceleration sensors. The microphone signal is manipulated depending on the derived speech features, for example by using a Kalman filter, by a signal-to-noise ratio estimation or the like.
[0006] US 2014 / 093091 A1 relates to a system for recognizing speech activity of a user in the case of using an acceleration sensor. In the speech activity recognition, not only the signal of the acceleration sensor is considered, but also the signal of a microphone.
[0007] US 2017 / 365249 A1 relates to a system for performing an automatic speech recognition in the case of using an endpoint marker, which is generated by means of a speech activity detector based on an acceleration sensor. The speech activity recognition is performed depending on the signal of the acceleration sensor and on the signal of a microphone.
[0008] In battery-powered or rechargeable devices, such as in-ear headphones, microphone-based systems that continuously detect audio data for keyword recognition or speech recognition require high power consumption, which is necessary for data detection and speech processing. Here, speech recognition algorithms typically run on an external digital signal processor (DSP) that combines signals from the microphone and an accelerometer sensor.
[0009] Furthermore, microphone-based speech activity recognition is prone to errors in speech recognition, especially due to interference noise present in noisy environments. Summary of the Invention
[0010] The present invention relates to an inertial sensor unit having features as described below and a method for detecting language activity by means of the inertial sensor unit.
[0011] Preferred embodiments are described below.
[0012] Therefore, according to a first aspect, the present invention relates to an inertial sensor unit. The inertial sensor unit includes sensor elements for detecting motion and vibration and converting motion and vibration into electrical sensor signals. Furthermore, the inertial sensor unit includes signal processing devices for analyzing and processing the sensor signals, with a particular objective: detecting vibrations caused by linguistic activity. Additionally, the inertial sensor unit includes an interface for signaling the detected linguistic activity. The signal processing devices include a first processing stage and a second processing stage for the sensor signals, wherein the first processing stage is designed to check for a first criterion for the presence of linguistic activity, and the second processing stage is designed to check for at least one additional second criterion for the presence of linguistic activity. The second processing stage is only executed if the sensor signal has passed through the first processing stage and the first criterion for the presence of linguistic activity is met. The signal processing devices are designed to manipulate the interface to signal the linguistic activity only if the sensor signal has passed through the second processing stage and at least one additional second criterion for the presence of linguistic activity is met.
[0013] Therefore, according to a second aspect, the present invention relates to a method for detecting language activity using an inertial sensor unit, the inertial sensor unit comprising at least one sensor element, a signal processing device, and an interface for signaling the detected language activity. Motion and vibration are detected by the at least one sensor element and converted into at least one electrical sensor signal. The sensor signal is analyzed and processed by the signal processing device. A first criterion for the presence of language activity is checked. Only if the first criterion for the presence of language activity is met is at least one additional second criterion for the presence of language activity checked. Only if the at least one additional second criterion for the presence of language activity is met is the interface manipulated to signal the language activity.
[0014] Advantages of the present invention
[0015] This invention provides a particularly power-efficient inertial sensor unit. The analysis and processing of the sensor signal generated by the sensor element is performed in two stages. In the first step, analysis based on the current measurement data is performed by a first processing stage. This involves, for example, simple analysis by comparing thresholds at the current measurement point. Only when a first criterion for the presence of linguistic activity is met is a more complex analysis method applied by a second processing stage in the second step. Here, for example, values stored in a buffer can be considered.
[0016] Therefore, according to the present invention, at least two analysis and processing methods are used in a time-variable manner. The first analysis and processing method can be based on the current data point, and the second analysis and processing method can be based on multiple data points in the buffer. Data in the buffer is stored and analyzed only after a first criterion of the first analysis and processing method is met.
[0017] According to another embodiment of the inertial sensor unit, at least two sensor elements are provided for detecting motion and vibration in different spatial directions. The inertial sensor unit can particularly include sensor elements that detect acceleration or rotation along or about different axes.
[0018] According to another embodiment of the inertial sensor unit, at least one accelerometer sensor element and / or at least one rotation speed sensor element are provided. The inertial sensor unit can include a biaxial or triaxial accelerometer and / or a rotation speed sensor.
[0019] According to another embodiment of the inertial sensor unit, the signal processing device further includes at least one signal filter, particularly a high-pass filter and / or a band-pass filter, for preprocessing the sensor signal, and at least one analog-to-digital converter for the sensor signal. The signal filter can have variable filter parameters. The analog-to-digital converter can have a variable sampling rate. The signal filter can be configured to suppress or filter acceleration signals generated not through speech activity but through the user's motion.
[0020] According to another embodiment of the inertial sensor unit, different operating modes can be implemented in such a way that the various components of the inertial sensor unit are optionally activatable or deactivatable, and / or the various components of the inertial sensor unit can operate in different operating modes. For example, the sensor components can be activated or deactivated in an axial manner, or the second processing stage can be activated or deactivated. Additionally or alternatively, for example, the analog-to-digital converter can be operable in different operating modes. Thus, in the first operating mode, the inertial sensor unit operates in a particularly power-saving manner. According to one embodiment of the inertial sensor unit, this can be achieved through a low data rate, a low over-sampling rate (OSR), or by measuring only using a single axis. Once the first criterion is met, the inertial sensor unit can automatically transition to the second operating mode. Then, the inertial sensor unit stores the measurement data in a buffer, and once a predefined number of measurement data points have been stored, more complex analysis processing of the buffer contents is performed. If the second criterion is met, a signal is generated through language detection.
[0021] The power-saving implementation scheme is ensured by the following method: the inertial sensor unit automatically switches between two operating modes and thus enables variable power-saving analysis processing that requires a small amount of computation and can be achieved by shutting down individual axes or by configuring the oversampling rate.
[0022] According to another embodiment of the inertial sensor unit, in a first operating mode, a first processing stage of the signal processing device operates in the first operating mode, and a second processing stage is turned off. In a second operating mode, the first processing stage of the signal processing device operates in the second operating mode, and the second processing stage is activated. The signal processing device is designed to automatically switch between the first and second operating modes based on whether a first and / or second criterion for the presence of linguistic activity is met.
[0023] According to another embodiment of the inertial sensor unit, the current consumption in the first operating mode is less than the current consumption in the second operating mode. In the first operating mode, at least one parameter can be configured or optimized in the following manner:
[0024] Compared to the second operating mode, the data rate is selected to be lower, for example, 2kHz.
[0025] Compared to the second operating mode, a higher noise level and a lower oversampling rate are set.
[0026] Only one active axis of the sensor element is used for analysis and processing. Therefore, it is possible for only one channel to be active in the analog-to-digital converter, or for only one analog-to-digital converter to be active.
[0027] In the second operating mode, at least one parameter can be configured or optimized in the following ways:
[0028] Compared to the first operating mode, a higher data rate is selected, such as 4kHz or 8kHz.
[0029] Data is stored in an active First-In-First-Out (FIFO) memory.
[0030] Multiple axes of the sensor element are active. For example, two axes can be active, such as X and Z, or three axes can be active, namely X, Y, and Z.
[0031] Compared to the first operating mode, a lower noise level and a higher oversampling rate are set. The conversion rate can be improved compared to the first operating mode.
[0032] According to another embodiment of the inertial sensor unit, the first processing stage includes at least one comparator that compares the current signal amplitude of the sensor signal with at least one threshold to determine whether a first criterion for the presence of speech activity is met. This enables the differentiation of speech activity from other movements of the user.
[0033] According to another embodiment of the inertial sensor unit, the second processing stage of the signal processing device includes: a buffer for buffering a defined number of successive sampled values of the sensor signal, and a signal analysis device for determining at least one signal characteristic based on the buffered sampled values and for comparing the at least one signal characteristic with at least one additional second criterion for the presence of linguistic activity. This enables the identification of the actual presence of linguistic activity in a power-saving manner.
[0034] According to another embodiment of the inertial sensor unit, the signal analysis device is designed to compare at least one obtained signal characteristic with at least one additional third criterion to identify at least one additional cause for the sensor signal. This allows other causes, such as shaking, tapping, or scraping motions by the user on the device, to be ruled out.
[0035] According to another embodiment of the inertial sensor unit, it is possible to signal speech activity to an external system. This can be achieved, for example, via an interrupt method. For instance, the digital signal processor (DSP) can be woken up. Thus, the inertial sensor unit can wake up the entire system to reduce the required data transfer between the DSP and the host CPU, for example, by having the DSP operate normally under default conditions. It is in sleep mode. Current consumption is reduced by integrating the detection of speech activity into the inertial sensor unit. Furthermore, the entire system only needs to be woken up when speech activity is detected.
[0036] According to another embodiment of the method for detecting language activity, sensor signals are preprocessed using a signal processing device. The preprocessing of the sensor signals includes signal filtering, particularly high-pass filtering and / or band-pass filtering, and includes analog-to-digital conversion in which the analog sensor signals are sampled and digitized, such that the digitized sensor signals exist in the form of a sequence of sampled values.
[0037] According to another embodiment of the method for detecting language activity, a first criterion for the presence of language activity is checked by comparing the current signal amplitude or current sample value of the sensor signal with at least one threshold.
[0038] According to another embodiment of the method for detecting language activity, as a first criterion for the presence of language activity, it is checked whether the current signal amplitude or current sample value of the sensor signal is greater than a first threshold and / or less than a second threshold within a predetermined duration.
[0039] According to another embodiment of the method for detecting language activity, when a first criterion for the presence of language activity is met, a predetermined number N of successive sampled values of the sensor signal are buffered in a buffer of a signal processing device, at least one signal characteristic is obtained based on the buffered sampled values, and the at least one signal characteristic is compared with at least one additional second criterion for the presence of language activity.
[0040] According to another embodiment of the method for detecting language activity, when a first criterion for the presence of language activity is met, at least one obtained signal characteristic is compared with at least one additional third criterion in order to identify at least one additional cause for the sensor signal.
[0041] According to another embodiment of the method for detecting language activity, if only a first criterion for the presence of language activity is checked, the inertial sensor unit operates in a first operating mode, wherein when at least one additional second criterion for the presence of language activity is checked, the inertial sensor unit operates in a second operating mode, and wherein the switching between the first operating mode and the second operating mode is automatic based on whether the first criterion for the presence of language activity and / or at least one additional second criterion is met.
[0042] According to another embodiment of the method for detecting language activity, different operating modes of the inertial sensor unit are implemented by: optionally activating or deactivating the individual components of the inertial sensor unit, and / or having the individual components of the inertial sensor unit operate in different operating modes. Attached Figure Description
[0043] The attached diagram shows:
[0044] Figure 1 A schematic block diagram of an inertial sensor unit according to one embodiment of the invention is shown.
[0045] Figure 2 A schematic diagram showing two operating modes is provided.
[0046] Figure 3 A schematic diagram showing an acceleration signal detected by an inertial sensor unit according to an embodiment of the present invention;
[0047] Figure 4 A flowchart is shown for a method of detecting language activity using an inertial sensor unit according to an embodiment of the present invention;
[0048] Figure 5 A flowchart illustrating a method for detecting language activity using an inertial sensor unit according to another embodiment of the invention; and
[0049] Figure 6 A flowchart is shown for a method of detecting language activity using an inertial sensor unit according to another embodiment of the invention. Detailed Implementation
[0050] Figure 1A schematic block diagram of an inertial sensor unit 1 is shown, which can be used, for example, in portable devices, particularly in-ear headphones, headsets, helmets, or smart glasses.
[0051] The inertial sensor unit 1 includes a device 5 for energy management, a clock generator 6, and control logic 7. Furthermore, the inertial sensor unit 1 includes an interface 4 for signaling detected speech activity. Additionally, the inertial sensor unit 1 includes at least one sensor element 2 for detecting motion and vibration and converting motion and vibration into electrical sensor signals. For example, an acceleration sensor element 2 can be provided for measuring acceleration along mutually perpendicular axes X, Y, and Z. Furthermore, a rotational speed sensor element can be provided for measuring rotation about mutually perpendicular axes X', Y', and Z', wherein the axis for measuring acceleration and the axis for measuring rotation can be the same. Therefore, motion and vibration, preferably in different spatial directions, can be detected.
[0052] Furthermore, the inertial sensor unit 1 includes a signal processing device 3 for analyzing and processing sensor signals, particularly for detecting vibrations caused by speech activity. The signal processing device 3 includes an analog-to-digital converter 34 that digitizes the sensor signal from at least one sensor element 2. The analog-to-digital converter 34 can have a variable sampling rate. The signal output from the analog-to-digital converter 34 is preprocessed by a signal filter 33. The signal filter 33 can have variable filter parameters. The signal filter 33 can include a high-pass filter and / or a band-pass filter.
[0053] The signal processing device 3 includes a first processing stage 31 and a second processing stage 32 for sensor signals. The first processing stage 31 checks for a first criterion regarding the presence of speech activity. The second processing stage 32 checks for a second criterion regarding the presence of speech activity. The second processing stage 32 only proceeds if the sensor signal has passed through the first processing stage 31 and the first criterion for the presence of speech activity is met. For example, the first processing stage 31 can determine, using a comparator, whether the current signal amplitude of the sensor signal exceeds a threshold. If so, the first criterion for the presence of speech activity is met.
[0054] The signal processing device 3 is designed to control the interface 4 to signal the speech activity only when the sensor signal has passed through the second processing stage 32 and at least one additional second criterion for the presence of speech activity is met.
[0055] The second processing stage 32 of the signal processing device 3 includes: a buffer 35 for buffering a defined number of successive sampled values of the sensor signal; and a signal analysis device 36 for determining at least one signal characteristic based on the buffered sampled values and for comparing the at least one signal characteristic with at least one additional second criterion for the presence of speech activity. Furthermore, the signal analysis device 36 can compare the determined signal characteristic with at least one additional third criterion to identify at least one additional cause for the sensor signal. This allows other causes, such as shaking or scraping motions, to be ruled out. The signal analysis device 36 and the first processing stage 31 are part of a speech activity detection unit 37. This speech activity detection unit can store the state of speech activity (detected / not detected) in a register 8 or output the state of speech activity via interrupt logic 9.
[0056] Figure 2 A schematic diagram of two operating modes, M1 and M2, is shown, in which the inertial sensor unit 1 can operate. Here, individual components of the inertial sensor unit 1 can be activated, deactivated, or operated in different operating modes. For example, sensor element 2 can be activated or deactivated in an axial manner, or the second processing stage 32 can be activated or deactivated. Additionally or alternatively, for example, the analog-to-digital converter 34 can operate in different operating modes M1 and M2.
[0057] In the first operating mode M1, the first processing stage 31 of the signal processing device 3 can operate in the first operating mode M1, and the second processing stage 32 can be turned off. In the second operating mode M2, the first processing stage 31 of the signal processing device 3 can operate in the second operating mode M2, and the second processing stage 32 is activated. The signal processing device 3 automatically switches between the first operating mode M1 and the second operating mode M2 based on whether the first and / or second criteria for the presence of language activity are met.
[0058] In the first operating mode M1, measurements can be performed at a low data rate, using a low oversampling rate of the signal, or by using only a single axis (with the remaining AD converter channels disabled). The second operating mode M2 is designed to perform data detection as accurately and quickly as possible. For example, the second operating mode M2 is achieved by using a higher data rate, a higher oversampling rate, and / or by measuring all axes (two or three).
[0059] Figure 3 A schematic diagram showing the acceleration signal detected by inertial sensor unit 1 is presented. Figure 3Above, the acceleration 'a' measured by sensor element 2, in 1sb (least significant bit), is plotted for a given time 't' (in seconds) and against three different axes x, y, and z. Figure 3 In the middle, the corresponding data after passing through high-pass filter 33 is plotted. Box R shows the area where the speech signal appears. Figure 3 Below, the average amplitudes magx and mazz in the x and z directions are plotted, with the first threshold T1 and the second threshold T2 also shown. If the average amplitudes magx and mazz are between the first threshold T1 and the second threshold T2, then the criteria for the existence of linguistic activity are met.
[0060] Figure 4 A flowchart is shown for a method of detecting language activity using an inertial sensor unit, particularly the aforementioned inertial sensor unit 1. Conversely, the inertial sensor unit 1 can also be configured to perform one of the methods described below. The inertial sensor unit 1 includes a sensor element 2, a signal processing device 3, and an interface 4 for signaling the detected language activity.
[0061] In the first method step S11, motion and vibration are detected by sensor element 2 and converted into electrical sensor signals. In the second method step S12, the sensor signals are preprocessed, for example, using a high-pass filter and / or a low-pass filter, to remove signal components corresponding to typical user movements. For example, a band-pass filter that allows a frequency range between 250 Hz and 2 kHz can be used.
[0062] In method step S13, it is checked whether a first criterion is met. This involves the following condition: the condition is used to check whether language signals may be involved. In particular, it can be checked whether the absolute value of the preprocessed sensor signal is between a first threshold and a second threshold. If this is not the case, the method is aborted or repeated.
[0063] Otherwise, in method step S14, the data is cached in buffer 35, which has a capacity of [size missing]. N, where N is an integer.
[0064] In method step S15, features are extracted from the data stored in buffer 35. For example, average amplitude, zero-crossing rate, and similar features can be extracted.
[0065] In another method step S16, the extracted features are analyzed according to a second criterion to determine whether they meet the conditions for language activity. Language activity may be detected. The control interface 4 is used to signal the language activity.
[0066] Figure 5 A flowchart is shown for a method of detecting language activity using an inertial sensor unit, particularly the aforementioned inertial sensor unit 1. The inertial sensor unit 1 includes a sensor element 2, a signal processing device 3, and an interface 4 for notifying the detected language activity by signaling.
[0067] In the first method step S21, motion and vibration are again detected by sensor element 2 and converted into electrical sensor signals. In the second method step S22, as described above, preprocessing of the sensor signals is performed.
[0068] In method step S23, as described above, it is checked whether the absolute value of the preprocessed sensor signal is between the first threshold and the second threshold. If not, the method is terminated or repeated.
[0069] Otherwise, in method step S24, the data is cached in buffer 35.
[0070] In method step S25, as described above, features are extracted from the data stored in buffer 35.
[0071] In method step S26, the data is analyzed and processed to identify other user behaviors. For example, shaking, tapping, gripping, or manipulating the device by a user can be identified. These other user behaviors can be identified if one or more data points in buffer 35 have values higher than a maximum threshold. Methods such as Fourier analysis, spectral analysis, or wavelet analysis can be used to identify these other user behaviors.
[0072] In method step S27, the analysis process checks whether such additional user behavior has been identified. If so, no language activity is involved, and the method is either terminated or repeated.
[0073] Otherwise, in method step S26, as described above, the extracted features are analyzed according to the second criterion to determine whether they meet the conditions for language activity. Language activity may be detected. The control interface 4 is used to signal the language activity.
[0074] During method steps S21 to S23, inertial sensor unit 1 operates in the first operating mode M1. During method steps S24 to S28, inertial sensor unit 2 operates in the second operating mode M2.
[0075] Figure 6 A flowchart is shown for a method of detecting language activity using an inertial sensor unit, particularly the aforementioned inertial sensor unit 1. The inertial sensor unit 1 includes a sensor element 2, a signal processing device 3, and an interface 4 for notifying the detected language activity by signaling.
[0076] In the first method step S31, as described above, motion and vibration are detected by sensor element 2 and converted into electrical sensor signals. In the second method step S32, as described above, preprocessing of the sensor signals is performed.
[0077] In method step S33, the sensor signal, i.e. the detected value, is stored in buffer 35.
[0078] In method step S34, it is checked whether the absolute value of the preprocessed sensor signal is between the first threshold and the second threshold. If not, method step S31 is repeated.
[0079] Otherwise, in method step S35, check if the buffer is filled with N new values. If not, then re-execute method step S31.
[0080] In method step S36, as described above, features are extracted from the data stored in buffer 35.
[0081] In method step S37, as described above, the extracted features are analyzed according to the second criterion to determine whether they meet the conditions for language activity. Language activity may be detected. The control interface 4 is used to signal the language activity.
[0082] Therefore, method steps S31 to S35, M1 are always performed, while method steps S36 to S37, M2 are only performed if the analysis and processing S34 and S35 are successful.
[0083] This method enables variable analysis processing of the data stored in buffer 35, wherein analysis processing of buffer 35 is performed only after criterion 1 (simple threshold comparison) is met, and a predefined number of measurement data are additionally stored after criterion 1 is met.
[0084] According to this embodiment, in order to perform analysis and processing, the measurement data from buffer 35 is organized such that the measurement data before and after the verification standard are included in the analysis and processing window. This enables power-efficient detection of speech activity. The detection is highly accurate because it also considers the signal before the standard, and it is fast due to the short waiting time.
Claims
1. An inertial sensor unit (1), said inertial sensor unit comprising at least: a. A sensor element (2), said sensor element (2) being used to detect motion and vibration and convert said motion and vibration into electrical sensor signals, b. A signal processing device (3), said signal processing device being used to analyze and process the sensor signal, having the objective of: detecting vibrations caused by speech activity, and c. Interface (4), the interface being used to signal the detected language activity, in, The signal processing device (3) includes a first processing stage (31) and a second processing stage (32) for the sensor signal, wherein the first processing stage (31) is designed to check for a first criterion for the presence of language activity, and the second processing stage (32) is designed to check for at least one additional second criterion for the presence of language activity. The second processing level (32) is only processed if the sensor signal has passed through the first processing level (31) and the first criterion for the presence of language activity is met. The signal processing device (3) is designed to manipulate the interface (4) to signal the speech activity only when the sensor signal has passed through the second processing stage (32) and at least one additional second criterion for the presence of speech activity is met. in, The inertial sensor unit (1) can achieve different operating modes by the following means: each component of the inertial sensor unit (1) is activatable or deactivable, and / or each component of the inertial sensor unit can operate in different operating modes. in, In the first operating mode, the first processing stage (31) of the signal processing device (3) operates in the first operating mode, and the second processing stage (32) is turned off. In the second operating mode, the first processing stage (31) of the signal processing device (3) operates in the second operating mode, and the second processing stage (32) is active. The signal processing device (3) is designed to automatically switch between the first operating mode and the second operating mode based on whether the first and / or second criteria for the presence of language activity are met. In the first operating mode, measurements can be performed at a low data rate, using a low oversampling rate of the signal and with only a single axis. In the second operating mode, measurements can be performed at a higher data rate, using a higher oversampling rate of the signal and with the aid of all axes.
2. The inertial sensor unit (1) according to claim 1, characterized in that, It is equipped with at least two sensor elements (2) for detecting motion and vibration in different spatial directions.
3. The inertial sensor unit (1) according to claim 1 or 2, characterized in that, It is equipped with at least one acceleration sensor element (2) and / or at least one speed sensor element (2).
4. The inertial sensor unit (1) according to any one of claims 1 to 3, characterized in that, The signal processing device (3) further includes: At least one signal filter (33) for preprocessing the sensor signal, At least one analog-to-digital converter (34) for the sensor signal.
5. The inertial sensor unit (1) according to claim 4, characterized in that, The signal filter (33) includes a high-pass filter and / or a band-pass filter.
6. The inertial sensor unit (1) according to any one of claims 1 to 5, characterized in that, The current consumption in the first operating mode is less than the current consumption in the second operating mode.
7. The inertial sensor unit (1) according to any one of claims 1 to 6, characterized in that, The first processing stage (31) includes at least one comparator that compares the current signal amplitude of the sensor signal with at least one threshold to determine whether a first criterion for the presence of language activity is met.
8. The inertial sensor unit (1) according to any one of claims 1 to 7, characterized in that, The second processing stage (32) of the signal processing device (3) includes at least: Buffer (35), the buffer being used to buffer a defined number of successive sampled values of the sensor signal, A signal analysis device (36) is used to determine at least one signal characteristic based on the cached sample values and to compare the at least one signal characteristic with at least one additional second criterion for the presence of language activity.
9. The inertial sensor unit (1) according to claim 8, characterized in that, The signal analysis device (36) is designed to compare at least one obtained signal characteristic with at least one additional third criterion in order to identify at least one additional cause for the sensor signal.
10. A method for detecting language activity using an inertial sensor unit (1), the inertial sensor unit comprising at least one sensor element (2), a signal processing device (3), and an interface (4) for signaling the detected language activity. in, Motion and vibration are detected by the at least one sensor element (2) and converted into at least one electrical sensor signal. The sensor signal is analyzed and processed using the signal processing device (3). Specifically, the first criterion for the existence of a language activity is checked. Only when the first criterion for the existence of a language activity is met is at least one additional second criterion for the existence of a language activity checked. And only when the at least one additional second criterion for the existence of a language activity is met is the interface (4) manipulated to signal the language activity. In the first operating mode, measurements are performed at a low data rate, using a low oversampling rate of the signal and with only a single axis. In the second operating mode, measurements are performed at a higher data rate, using a higher oversampling rate of the signal and with the aid of all axes.
11. The method according to claim 10, wherein, The sensor signal is preprocessed using the signal processing device (3), and the preprocessing of the sensor signal includes at least: Signal filtering, Analog-to-digital conversion, in which analog sensor signals are sampled and digitized, such that the digitized sensor signals exist in the form of a sequence of sampled values.
12. The method according to claim 11, wherein, The signal filtering includes high-pass filtering and / or band-pass filtering.
13. The method according to any one of claims 10 to 12, wherein, The first criterion for examining the existence of language activity is as follows: The current signal amplitude or current sample value of the sensor signal is compared with at least one threshold.
14. The method according to any one of claims 10 to 13, wherein, As a first criterion for the existence of language activity, it is checked whether the current signal amplitude or current sample value of the sensor signal for a pre-given duration is greater than a first threshold and / or less than a second threshold.
15. The method according to any one of claims 11 to 14, characterized in that, When the first criterion for the existence of language activity is met... The sensor signal is buffered in the buffer (35) of the signal processing device (3) at a predetermined number N of consecutive sampled values. At least one signal characteristic is obtained based on the cached sample values. The at least one signal characteristic is compared with at least one additional second criterion for the existence of language activity.
16. The method according to claim 15, characterized in that, When the first criterion for the existence of language activity is met, at least one signal characteristic is compared with at least one additional third criterion in order to identify at least one additional cause for the sensor signal.
17. The method according to any one of claims 10 to 16, characterized in that, If only the first criterion for the existence of language activity is checked, the inertial sensor unit (1) operates in the first operating mode; when the at least one additional second criterion for the existence of language activity is checked, the inertial sensor unit (1) operates in the second operating mode and automatically switches between the first operating mode and the second operating mode depending on whether the first criterion for the existence of language activity and / or the at least one additional second criterion is met.
18. The method according to claim 17, characterized in that, The different operating modes of the inertial sensor unit (1) can be achieved by: activating or deactivating the various components of the inertial sensor unit (1), and / or having the various components of the inertial sensor unit operate in different operating modes.
Citation Information
Patent Citations
Earbud speech estimation
US10397687B2
Adjusted noise suppression and voice activity detection
US20130196715A1
System and method of detecting a user's voice activity using an accelerometer
US20140093091A1
System and method for performing automatic gain control using an accelerometer in a headset
US20170263267A1
System and method of performing automatic speech recognition using end-pointing markers generated using accelerometer-based voice activity detector
US20170365249A1