Voice activity detection method, voice wake-up system, and electronic device
By periodically acquiring the energy and zero-crossing feature count of the voice analog signal in the voice wake-up system, and combining it with the temporal change characteristics for dual judgment, the problem of high false wake-up rate in traditional voice wake-up systems is solved, achieving higher detection accuracy and power consumption optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU ANYKA MICROELECTRONICS CO LTD
- Filing Date
- 2026-06-25
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional voice wake-up systems have a high false detection rate in their voice activity detection units, leading to frequent false wake-ups of high-power units and wasted power.
By periodically acquiring the energy count and zero-crossing feature count of the voice analog signal, and combining the energy timing change characteristics and zero-crossing timing change characteristics, a dual judgment is made to output a wake-up signal, thereby reducing false wake-ups.
It improves the accuracy of voice activity detection, reduces the number of false wake-ups, and avoids power waste caused by frequent activation of high-power units.
Smart Images

Figure CN122493850A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice information processing technology, and in particular to a voice activity detection method, a voice wake-up system, and an electronic device. Background Technology
[0002] With the widespread adoption of smart terminals such as smart homes, in-vehicle systems, and wearables, traditional manual interaction methods are no longer sufficient to meet the demand for convenient, contactless operation. Leveraging upgrades in voice processing and embedded computing technology, voice wake-up systems have emerged, enabling seamless device activation via voice commands and becoming a crucial entry point for intelligent human-computer interaction.
[0003] Traditional voice wake-up systems' voice activity detection units can only make a rough judgment on the voice analog signal, resulting in a high false detection rate and frequent false wake-ups of subsequent high-power units, thus wasting the power of the voice wake-up system. Summary of the Invention
[0004] Therefore, it is necessary to provide a voice activity detection method, a voice wake-up system, and an electronic device to address the aforementioned technical problems.
[0005] In a first aspect, this application provides a voice activity detection method applied to a voice activity detection unit in a voice wake-up system. The voice wake-up system further includes an analog microelectromechanical system (MEMS) microphone unit and a wake-up control unit. The voice activity detection unit is connected to both the analog MEMS microphone unit and the wake-up control unit. The method includes:
[0006] By controlling the microphone unit of the analog microelectromechanical system to periodically turn on to output an analog voice signal, at least two consecutive energy count values are obtained based on the signal amplitude of the analog voice signal, and at least two consecutive zero-crossing feature count values are obtained based on the number of zero-crossings of the analog voice signal.
[0007] Based on the energy count value of the current cycle and the energy count value of the previous cycle, the energy time series change characteristics are obtained, and based on the zero-crossing characteristic count value of the current cycle and the zero-crossing characteristic count value of the previous cycle, the zero-crossing time series change characteristics are obtained.
[0008] Based on the energy timing change characteristics, a first voice activity determination result is obtained, and based on the zero-crossing timing change characteristics, a second voice activity determination result is obtained. The first voice activity determination result and the second voice activity determination result are used to characterize whether voice activity exists.
[0009] If both the first voice activity determination result and the second voice activity determination result indicate the presence of voice activity, a wake-up signal is output to the wake-up control unit to wake up the wake-up control unit.
[0010] In one embodiment, obtaining at least two consecutive cycles of energy count values based on the signal amplitude of the speech analog signal includes:
[0011] A comparator is used to perform positive and negative threshold discrimination on the speech analog signal of each cycle, so as to obtain the positive and negative discrimination results for each cycle.
[0012] The positive and negative discrimination results of each cycle are logically synthesized to determine the level state of the energy indication signal for each cycle;
[0013] The duration of the high-level state of the energy indicator signal in each cycle is counted to obtain the energy count value for each cycle.
[0014] In one embodiment, obtaining the energy time-series change characteristics based on the energy count value of the current period and the energy count value of the previous period includes:
[0015] The energy count value of the current period is summed with the energy count value of the previous period to obtain the sum of the energy counts.
[0016] The energy count difference is obtained by subtracting the energy count value of the current cycle from the energy count value of the previous cycle.
[0017] Based on the sum of the energy counts and the difference between the energy counts, the energy temporal variation characteristics are obtained.
[0018] In one embodiment, obtaining the first voice activity determination result based on the energy temporal change characteristics includes:
[0019] If the sum of the energy counts is greater than the first preset threshold, and the difference in the energy counts is greater than the first preset difference threshold, then the first voice activity determination result is that there is voice activity.
[0020] If the sum of the energy counts is less than or equal to a first preset threshold, or the difference in the energy counts is less than or equal to a first preset difference threshold, then the first voice activity determination result is that there is no voice activity.
[0021] In one embodiment, obtaining the zero-crossing time-series change characteristics based on the zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period includes:
[0022] The zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period are summed to obtain the zero-crossing feature count sum value.
[0023] The zero-crossing feature count value of the current period is subtracted from the zero-crossing feature count value of the previous period to obtain the zero-crossing feature count difference value.
[0024] The zero-crossing time-series change characteristics are obtained based on the sum of the zero-crossing feature counts and the difference between the zero-crossing feature counts.
[0025] In one embodiment, obtaining the second speech activity determination result based on the zero-crossing timing change characteristics includes:
[0026] If the sum of the zero-crossing feature counts is within the second preset threshold range, and the difference between the zero-crossing feature counts is greater than the second preset threshold, then the second speech activity determination result is that speech activity exists.
[0027] If the zero-crossing feature count and value are outside the second preset threshold range, or if the difference in the zero-crossing feature count is less than or equal to the second preset difference threshold, then the second voice activity determination result is that there is no voice activity.
[0028] Secondly, this application provides a voice activity detection device applied to a voice activity detection unit in a voice wake-up system. The voice wake-up system further includes an analog microelectromechanical system (MEMS) microphone unit and a wake-up control unit. The voice activity detection unit is connected to the analog MEMS microphone unit and the wake-up control unit, respectively. The device includes:
[0029] The counting value acquisition module is used to control the microphone unit of the analog microelectromechanical system to periodically turn on to output a voice analog signal, obtain energy count values for at least two consecutive cycles based on the signal amplitude of the voice analog signal, and obtain zero-crossing feature count values for at least two consecutive cycles based on the number of zero-crossings of the voice analog signal.
[0030] The time-series change feature acquisition module is used to obtain energy time-series change features based on the energy count value of the current period and the energy count value of the previous period, and to obtain zero-crossing time-series change features based on the zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period.
[0031] The determination result acquisition module is used to obtain a first voice activity determination result based on the energy timing change characteristics, and to obtain a second voice activity determination result based on the zero-crossing timing change characteristics. The first voice activity determination result and the second voice activity determination result are used to characterize whether voice activity exists.
[0032] The wake-up module of the wake-up control unit is used to output a wake-up signal to the wake-up control unit when both the first voice activity determination result and the second voice activity determination result indicate the presence of voice activity, so as to wake up the wake-up control unit.
[0033] Thirdly, this application also provides a voice wake-up system, the system comprising an analog microelectromechanical system microphone unit, a voice activity detection unit, and a wake-up control unit; the voice activity detection unit is connected to the analog microelectromechanical system microphone unit and the wake-up control unit respectively;
[0034] The analog microelectromechanical system microphone unit is used to output analog voice signals;
[0035] The voice activity detection unit is used to perform the steps of the method described in the above embodiments;
[0036] The wake-up control unit is used to execute pre-set functional operations upon receiving a wake-up signal.
[0037] In one embodiment, the voice wake-up system further includes a power management chip connected to the analog microelectromechanical system (MEMS) microphone unit. The power management chip is used to intermittently supply power to the analog MEMS microphone unit when the analog MEMS microphone unit does not receive voice, and to synchronously control the opening and closing of the voice activity detection unit.
[0038] In one embodiment, the pre-set functional operations include analog-to-digital conversion, voice feature detection, and keyword detection. The wake-up control unit includes a low-precision analog-to-digital converter and a high-precision analog-to-digital converter. The low-precision analog-to-digital converter is used to convert the voice analog signal into a digital signal during the voice feature detection stage. The high-precision analog-to-digital converter is used to convert the voice analog signal into a digital signal during the keyword detection stage. The analog-to-digital conversion accuracy of the low-precision analog-to-digital converter is lower than a preset accuracy threshold, and the analog-to-digital conversion accuracy of the high-precision analog-to-digital converter is higher than the preset accuracy threshold.
[0039] Fourthly, this application also provides an electronic device, including a processor and a voice wake-up system as described in the above embodiments. The processor is connected to the wake-up control unit of the voice wake-up system, and the processor is used to operate when a valid wake-up command is received. The valid wake-up command is an instruction output when a preset wake-up word is recognized during keyword detection of the voice analog signal.
[0040] Fifthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, the computer program being executed by a processor using the methods described above.
[0041] Sixthly, this application also provides a computer program product. The computer program product includes a computer program that is executed by a processor using the methods described above.
[0042] This application obtains the first and second voice activity determination results based on the temporal change characteristics of the energy count value and zero-crossing feature count value of the current cycle and the previous cycle, in order to determine whether to output a wake-up signal to wake up the wake-up control unit. This can avoid the misjudgment of voice activity caused by judging whether to output a wake-up signal based solely on the energy count value and zero-crossing feature count value of a single cycle. It can improve the accuracy of voice activity detection, reduce the number of false wake-ups, and avoid the power waste caused by frequently activating subsequent high-power units. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a schematic diagram of the architecture of a traditional voice wake-up system in one embodiment;
[0045] Figure 2 This is a schematic diagram of the architecture of a traditional voice wake-up system with an added human voice feature detection step in one embodiment.
[0046] Figure 3 This is a diagram illustrating the application environment of a voice activity detection method in one embodiment;
[0047] Figure 4 This is a flowchart illustrating a speech activity detection method in one embodiment;
[0048] Figure 5 This is a schematic diagram of the process for obtaining the energy count value for each cycle in one embodiment;
[0049] Figure 6 This is a schematic diagram of the process for obtaining the zero-crossing feature count value for each cycle in one embodiment;
[0050] Figure 7 This is a structural block diagram of a voice activity detection device in one embodiment;
[0051] Figure 8 This is a schematic diagram of the operation mechanism of a voice wake-up system in one embodiment;
[0052] Figure 9 This is a schematic diagram of the wake-up determination process in one embodiment;
[0053] Figure 10 This is a flowchart illustrating the voice wake-up system in one embodiment;
[0054] Figure 11 This is a diagram of the internal structure of an electronic device in one embodiment. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0057] With the rapid development of artificial intelligence technology, voice wake-up and voice interaction functions have become standard features of wearable devices. Since these devices generally use lithium batteries for power, effectively extending battery life has become one of the key technical challenges. Therefore, voice wake-up systems need to minimize power consumption while ensuring full functionality.
[0058] Traditional voice wake-up system architecture, such as Figure 1 As shown, it includes the following core units: an analog micro-electro-mechanical system (MIC MEMS) microphone unit, a voice activity detection (VAD) unit, a high-precision analog-to-digital converter (e.g., a summation-delta analog-to-digital converter, or sigma-delta ADC for short), and a keyword detection unit.
[0059] Traditional voice wake-up systems typically involve two stages. The first stage involves a voice activity detection unit that monitors the analog audio signal output from the microphone unit of an analog microelectromechanical system (MEMS). When the voice activity detection unit detects potential voice activity based on the analog audio signal, it generates a wake-up signal. This wake-up signal activates the subsequent high-precision analog-to-digital converter (ADC) and keyword detection unit. The second stage involves the activated ADC performing high-fidelity sampling of the input analog audio signal and converting it into a digital signal. This digital signal is then passed to the keyword detection unit for accurate keyword recognition. Upon recognizing a preset wake-up word, the keyword detection unit outputs a valid wake-up command to instruct the wearable device to enter working mode.
[0060] However, traditional voice wake-up system architectures suffer from significant power consumption issues: the voice activity detection unit can only make a rough judgment on the analog voice signal, resulting in a high false detection rate and frequent false wake-ups of subsequent high-power units. Meanwhile, the keyword detection unit relies on a high-precision analog-to-digital converter and a high-performance central processing unit (CPU) or neural processing unit (NPU), consuming extremely high power, and a large number of false wake-up events lead to severe power waste.
[0061] To alleviate this problem, the industry has proposed an improvement: adding a human voice feature detection step before the keyword detection unit in traditional voice wake-up systems, such as... Figure 2 As shown, since the computational complexity required for the voice feature detection unit is much lower than that for keyword recognition, a low-power, low-computing-power central processing unit or neural network processor can be used to implement voice feature detection. Only when the voice feature detection unit confirms that the input signal contains valid voice components will the voice wake-up system enter the more energy-intensive keyword detection stage. Nevertheless, traditional technology still has significant shortcomings: the false wake-up rate of the voice activity detection unit is still too high, resulting in wasted power.
[0062] To address the problem of excessively high false wake-up rate in voice activity detection units, resulting in wasted power, embodiments of this application provide a voice activity detection method that can be applied to, for example... Figure 3 The voice wake-up system shown includes a voice activity detection unit. The system further includes an analog microelectromechanical system (MEMS) microphone unit and a wake-up control unit, with the voice activity detection unit connected to both. The wake-up control unit may include an analog-to-digital converter, a voice feature detection unit, and a keyword detection unit. The voice activity detection method provided in this application embodiment is as follows: Figure 4 As shown, the procedure includes steps S401 to S404. Wherein:
[0063] Step S401: By controlling the microphone unit of the analog microelectromechanical system to periodically turn on to output an analog voice signal, at least two consecutive cycles of energy count values are obtained based on the signal amplitude of the analog voice signal, and at least two consecutive cycles of zero-crossing feature count values are obtained based on the number of zero-crossings of the analog voice signal.
[0064] The voice activity detection unit can controllably drive the analog microelectromechanical system (MEMS) microphone unit to periodically turn on, so that the analog MEMS microphone unit works periodically. When the analog MEMS microphone unit receives valid voice or valid sound wave input during the working phase, it outputs a voice analog signal.
[0065] A period refers to multiple time periods of equal duration determined from the starting point of the analog microelectromechanical system microphone unit's output of the analog voice signal during a single wake-up activity.
[0066] When the microphone unit of the simulated microelectromechanical system outputs a voice analog signal, the voice activity detection unit can connect to extract at least two cycles of voice analog signal as the processing object. Based on the signal amplitude of the voice analog signal in each cycle, the unit performs sampling and processing on the voice analog signal in each cycle, calculates the signal energy in each cycle, and then obtains the energy count value of at least two consecutive cycles.
[0067] When the microphone unit of the simulated microelectromechanical system outputs a voice analog signal, the voice activity detection unit can connect to and extract at least two cycles of voice analog signal as the processing object. It can count the number of times the waveform of the voice analog signal in each cycle crosses the zero level in a unit of time, count the number of zero crossings in each cycle, and then obtain the zero crossing feature count value of at least two consecutive cycles.
[0068] Step S402: Based on the energy count value of the current cycle and the energy count value of the previous cycle, obtain the energy time series change characteristics, and based on the zero-crossing characteristic count value of the current cycle and the zero-crossing characteristic count value of the previous cycle, obtain the zero-crossing time series change characteristics.
[0069] The characteristics of energy temporal variation can be obtained based on the difference and the total continuity between the energy count value of the current period and the energy count value of the previous period.
[0070] The zero-crossing time series change characteristics are obtained based on the difference and the total continuity between the zero-crossing characteristic count value of the current period and the zero-crossing characteristic count value of the previous period.
[0071] Step S403: Based on the energy timing change characteristics, a first speech activity determination result is obtained, and based on the zero-crossing timing change characteristics, a second speech activity determination result is obtained. The first speech activity determination result and the second speech activity determination result are used to characterize whether speech activity exists.
[0072] Energy time-series variation characteristics can include the sum of energy counts and the difference in energy counts.
[0073] The presence of speech activity in the current environment can be determined based on the relative magnitude of the energy count and value in the energy temporal variation characteristics with a first preset threshold. Fixed noise in the natural environment can be eliminated based on the relative magnitude of the energy count difference in the energy temporal variation characteristics with a first preset difference threshold.
[0074] If there is speech activity in the current environment and the speech activity is not fixed noise in the natural environment, then the first speech activity determination result is that there is speech activity; otherwise, the first speech activity determination result is that there is no speech activity.
[0075] Zero-crossing time-series change characteristics can include the sum of zero-crossing feature counts and the difference between zero-crossing feature counts.
[0076] The presence of speech activity in the current environment can be determined by whether the zero-crossing feature count and value in the zero-crossing timing change features are within a second preset threshold range. Noise interference outside the human voice frequency range can be eliminated by the relative magnitude relationship between the difference in the zero-crossing feature counts in the zero-crossing timing change features and the second preset difference threshold.
[0077] If there is voice activity in the current environment and the voice activity is not noise interference outside the human voice frequency range, then the second voice activity determination result is that there is voice activity; otherwise, the second voice activity determination result is that there is no voice activity.
[0078] In step S404, if both the first voice activity determination result and the second voice activity determination result indicate the presence of voice activity, a wake-up signal is output to the wake-up control unit to wake up the wake-up control unit.
[0079] When both the first and second voice activity determination results indicate the presence of voice activity, and the presence of voice activity can be confirmed from both the energy and frequency dimensions of the voice analog signal, a wake-up signal can be output to the wake-up control unit to wake up the wake-up control unit to perform pre-set functional operations, such as performing security anomaly response operations or voice wake-up recognition operations, thereby realizing security monitoring and evidence collection or voice interaction control functions.
[0080] Specifically, when performing a security anomaly response, actions such as taking photos and recording videos can be initiated to preserve on-site data, enabling real-time monitoring and event tracing of risks such as intrusion and abnormal activity, and completing security early warning and on-site recording. When performing a voice wake-up recognition operation, analog-to-digital conversion, human voice detection, and keyword detection can be performed on the voice analog signal. Upon recognizing a preset wake-up word, a valid wake-up command is output to instruct the electronic device to enter working state, realizing voice interactive control functions.
[0081] In the above-mentioned voice activity detection method, the first and second voice activity determination results are obtained based on the temporal change characteristics of the energy count value and zero-crossing feature count value of the current cycle and the previous cycle, so as to determine whether to output a wake-up signal to wake up the wake-up control unit. This can avoid the misjudgment of voice activity caused by judging whether to output a wake-up signal based solely on the energy count value and zero-crossing feature count value of a single cycle. It can improve the accuracy of voice activity detection, reduce the number of false wake-ups, and avoid the power waste caused by frequently activating subsequent high-power units.
[0082] In one embodiment, based on the signal amplitude of the voice analog signal, at least two consecutive cycles of energy count values are obtained. The specific steps are as follows: a comparator is used to perform positive threshold discrimination and negative threshold discrimination on the voice analog signal of each cycle to obtain the positive discrimination result and the negative discrimination result of each cycle; the positive discrimination result and the negative discrimination result of each cycle are logically synthesized to determine the level state of the energy indicator signal of each cycle; the high level state duration of the energy indicator signal of each cycle is counted to obtain the energy count value of each cycle.
[0083] For example, such as Figure 5 As shown, at least two cycles of the voice analog signal VO output by the microphone unit of the analog microelectromechanical system are connected and extracted as the processing object. For the voice analog signal of the Nth cycle, comparators 1 and 2 can be used to perform positive threshold discrimination and negative threshold discrimination on the positive and negative half-cycle energy of the voice analog signal, respectively.
[0084] When the amplitude of the positive half-cycle signal of the analog speech signal is greater than the positive threshold of comparator 1 (which can be called the preset threshold Vref1), comparator 1 outputs a high-level signal STE_OUTP. That is, the positive discrimination result of the Nth cycle is that the output signal STE_OUTP is set to a high level, indicating that the positive half-cycle energy of the analog speech signal exceeds the limit. When the amplitude of the negative half-cycle signal of the analog speech signal is less than the negative threshold of comparator 2 (which can be called the preset threshold Vref2), comparator 2 outputs a high-level signal STE_OUTN. That is, the negative discrimination result of the Nth cycle is that the output signal STE_OUTN is set to a high level, indicating that the negative half-cycle energy of the analog speech signal exceeds the limit.
[0085] A logical OR operation is performed on the positive and negative discrimination results of the Nth cycle to determine the level state of the energy indicator signal (STE_OUT signal) of the Nth cycle; the high-level state duration of the energy indicator signal of the Nth cycle is counted according to the high-frequency clock (CLK) to obtain the energy count value E of the Nth cycle. The energy count value E can reflect the energy level of the voice analog signal of the Nth cycle.
[0086] In this embodiment, a comparator is used to perform positive and negative threshold discrimination on the voice analog signal of each cycle to determine the level state of the energy indicator signal of each cycle, thereby obtaining the energy count value of each cycle, and preparing data for subsequent determination of whether there is voice activity in the current environment.
[0087] In one embodiment, based on the signal amplitude of the speech analog signal, at least two consecutive cycles of zero-crossing feature count values are obtained. The specific steps are as follows: using a comparator to perform positive and negative zero-crossing feature threshold discrimination on the speech analog signal of each cycle, respectively, to obtain the positive and negative zero-crossing feature discrimination results for each cycle; logically synthesizing the positive and negative zero-crossing feature discrimination results for each cycle to determine the zero-crossing indicator signal for each cycle; counting the rising edges of the zero-crossing indicator signal pulses for each cycle to obtain the zero-crossing feature count value for each cycle.
[0088] For example, such as Figure 6 As shown, at least two cycles of the voice analog signal VO output by the microphone unit of the analog microelectromechanical system are connected and extracted as the processing object. For the voice analog signal of the Nth cycle, comparators 3 and 4 can be used to perform positive threshold discrimination of zero-crossing feature and negative threshold discrimination of zero-crossing feature, respectively.
[0089] When the amplitude of the analog speech signal is greater than the positive zero-crossing threshold of comparator 3 (which can be called the preset threshold Vref3), comparator 3 outputs a high-level signal ZCP_OUTP. This means that the positive zero-crossing characteristic determination result for the Nth cycle is that the output signal ZCP_OUTP is high, indicating that a positive zero-crossing transition feature has been detected. When the amplitude of the analog speech signal is less than the negative zero-crossing threshold of comparator 4 (which can be called the preset threshold Vref4), comparator 4 outputs a high-level signal ZCP_OUTN. This means that the negative zero-crossing characteristic determination result for the Nth cycle is that the output signal ZCP_OUTN is high, indicating that a negative zero-crossing transition feature has been detected.
[0090] A logical OR operation is performed on the positive and negative zero-crossing feature discrimination results of the Nth cycle to determine the level state of the zero-crossing indicator signal (ZCP_OUT signal) of the Nth cycle; the zero-crossing count is performed using the rising edge of the zero-crossing indicator signal pulse of each cycle as the counting clock to obtain the zero-crossing feature count value Z of each cycle. This zero-crossing feature count value Z can reflect the frequency characteristics of the voice analog signal of the Nth cycle.
[0091] In this embodiment, positive and negative threshold discrimination of zero-crossing features are performed on the voice analog signal of each cycle to determine the level state of the zero-crossing indicator signal of each cycle, thereby obtaining the zero-crossing feature count value of each cycle, and preparing data for subsequent determination of whether there is voice activity in the current environment.
[0092] In one embodiment, the method provided by this application further includes: after waking up the wake-up control unit, if the human voice feature detection fails, then the positive threshold is increased and the negative threshold is decreased.
[0093] After waking up the wake-up control unit, the wake-up control unit can perform human voice feature detection on the voice analog signal. Specifically, it can perform human voice feature matching on the voice analog signal. If the matching is successful, it indicates that the human voice feature detection is successful; if the matching fails, it indicates that the human voice feature detection fails.
[0094] If voice feature detection fails, the positive threshold in the energy count calculation process can be increased, and the negative threshold can be decreased. Specifically, the positive threshold Vref1 can be increased by one level, and the negative threshold Vref2 can be decreased by one level, so that the next energy detection in the speech activity detection unit requires a greater sound intensity to output a wake-up signal. By successively increasing the wake-up thresholds (positive and negative thresholds) in energy detection, false wake-ups caused by excessive environmental noise energy can be effectively suppressed, improving the accuracy of speech activity detection.
[0095] In one embodiment, the energy time-series change characteristics are obtained based on the energy count value of the current period and the energy count value of the previous period. The specific steps are as follows: perform a sum operation on the energy count value of the current period and the energy count value of the previous period to obtain the sum of energy counts; perform a difference operation on the energy count value of the current period and the energy count value of the previous period to obtain the energy count difference; and obtain the energy time-series change characteristics based on the sum of energy counts and the energy count difference.
[0096] The energy count value E[N] of the current cycle and the energy count value E[N-1] of the previous cycle can be added together to obtain the energy count sum value SUM_E=E[N]+E[N-1].
[0097] The energy count value E[N] of the current cycle can be subtracted from the energy count value E[N-1] of the previous cycle to obtain the energy count difference DIFF_E=E[N]-E[N-1].
[0098] The sum of energy counts, SUM=E[N]+E[N-1], and the difference in energy counts, DIFF=E[N]-E[N-1], can be used as characteristics of energy temporal variation.
[0099] In this embodiment, the energy count value of the current cycle and the energy count value of the previous cycle are summed and subtracted respectively to obtain the sum and difference of energy counts, so as to obtain the energy time sequence change characteristics. This can avoid the misjudgment of voice activity caused by judging whether to output a wake-up signal based on the energy count value of a single cycle. It can improve the accuracy of voice activity detection, reduce the number of false wake-ups, and avoid the power waste caused by frequent activation of subsequent high-power units.
[0100] In one embodiment, a first voice activity determination result is obtained based on the energy temporal change characteristics. The specific steps are as follows: if the sum of energy counts is greater than a first preset sum threshold and the difference in energy counts is greater than a first preset difference threshold, then the first voice activity determination result is that there is voice activity; if the sum of energy counts is less than or equal to the first preset sum threshold, or the difference in energy counts is less than or equal to the first preset difference threshold, then the first voice activity determination result is that there is no voice activity.
[0101] The first preset and threshold can be determined based on the voice energy requirements of the voice wake-up conditions in the actual scenario.
[0102] Since the energy base of fixed noise in the natural environment is relatively stable and the fluctuation is small, a first preset difference threshold can be determined based on the maximum amplitude of the energy fluctuation of fixed noise in the natural environment. Based on the relative magnitude relationship between the energy count difference in the energy time series change characteristics and the first preset difference threshold, fixed noise in the natural environment can be excluded.
[0103] If the sum of the energy counts is greater than the first preset threshold, it indicates that the energy of the voice simulation signal meets the voice wake-up condition, and it can be determined that there is voice activity in the current environment. If the sum of the energy counts is less than or equal to the first preset threshold, it indicates that the energy of the voice simulation signal does not meet the voice wake-up condition, and it can be determined that there is no voice activity in the current environment.
[0104] If the energy count difference is greater than the first preset difference threshold, it indicates a significant increase in sound energy in the current environment, and it can be determined that the speech activity in the current environment is not fixed noise in the natural environment. If the energy count difference is less than or equal to the first preset difference threshold, it indicates that the sound energy fluctuation in the current environment is small, and it can be determined that the speech activity in the current environment is fixed noise in the natural environment.
[0105] If the sum of the energy counts is greater than the first preset threshold, and the difference in energy counts is greater than the first preset difference threshold, it indicates that there is speech activity in the current environment and the speech activity is not fixed noise in the natural environment. Then, the first speech activity determination result is that there is speech activity.
[0106] If the sum of the energy counts is less than or equal to the first preset sum threshold, or the difference in energy counts is less than or equal to the first preset difference threshold, it indicates that there is no speech activity in the current environment or that the speech activity is fixed noise in the natural environment. In this case, the first speech activity determination result is that there is no speech activity.
[0107] In this embodiment, the first voice activity determination result is obtained based on the relative magnitude relationship between the sum of energy counts and the first preset threshold, and the relative magnitude relationship between the energy count difference and the first preset difference threshold, so as to determine whether to output a wake-up signal to wake up the wake-up control unit. This can avoid the misjudgment of voice activity caused by judging whether to output a wake-up signal based on the energy count value of a single cycle, improve the accuracy of voice activity detection, reduce the number of false wake-ups, and avoid the power waste caused by frequently activating subsequent high-power units.
[0108] In one embodiment, the zero-crossing time series change characteristics are obtained based on the zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period. The specific steps are as follows: perform a sum operation on the zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period to obtain the zero-crossing feature count sum value; perform a difference operation on the zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period to obtain the zero-crossing feature count difference value; and obtain the zero-crossing time series change characteristics based on the zero-crossing feature count sum value and the zero-crossing feature count difference value.
[0109] The zero-crossing feature count value Z[N] of the current period and the zero-crossing feature count value of the previous period can be used. Add them together to get the zero-crossing feature count and sum. .
[0110] The zero-crossing feature count value Z[N] of the current period and the zero-crossing feature count value of the previous period can be used. Subtracting the two values yields the difference in zero-crossing feature counts. .
[0111] Zero-crossing feature counts and values can be used. Difference between zero-crossing feature counts This serves as a characteristic of zero-crossing time-series changes.
[0112] In this embodiment, the zero-crossing feature count value of the current cycle and the zero-crossing feature count value of the previous cycle are summed and subtracted to obtain the sum and difference of the zero-crossing feature counts, thereby obtaining the timing change features of the zero-crossing features. This can avoid misjudging voice activity caused by relying solely on the zero-crossing feature count value of a single cycle to determine whether to output a wake-up signal, thus improving the accuracy of voice activity detection, reducing the number of false wake-ups, and avoiding power waste caused by frequently activating subsequent high-power units.
[0113] In one embodiment, a second voice activity determination result is obtained based on the zero-crossing timing change characteristics. The specific steps are as follows: if the sum of the zero-crossing feature counts is within the second preset threshold range and the difference between the zero-crossing feature counts is greater than the second preset difference threshold, then the second voice activity determination result is that there is voice activity; if the sum of the zero-crossing feature counts is outside the second preset threshold range, or the difference between the zero-crossing feature counts is less than or equal to the second preset difference threshold, then the second voice activity determination result is that there is no voice activity.
[0114] The second preset threshold interval [TH_zero_min, TH_zero_max] and the second preset difference threshold TH_zero_diff can be determined based on the human voice frequency range in the actual scene.
[0115] If the sum of the zero-crossing feature counts is within the second preset threshold range, and the difference in the zero-crossing feature counts is greater than the second preset difference threshold, it indicates that the speech activity in the current environment is noise interference outside the human voice frequency range, and the second speech activity determination result is that speech activity exists; if the sum of the zero-crossing feature counts is outside the second preset threshold range, it indicates that there is no speech activity in the current environment, or if the difference in the zero-crossing feature counts is less than or equal to the second preset difference threshold, it indicates that the speech activity in the current environment is noise interference outside the human voice frequency range, and the second speech activity determination result is that there is no speech activity.
[0116] In this embodiment, a second voice activity determination result is obtained based on whether the zero-crossing feature count and value are within the second preset threshold range, and the relative magnitude relationship between the energy count difference and the first preset threshold difference. This result is used to determine whether to output a wake-up signal to wake up the wake-up control unit. This avoids misjudging voice activity caused by relying solely on the zero-crossing feature count value of a single cycle to determine whether to output a wake-up signal. It can improve the accuracy of voice activity detection, reduce the number of false wake-ups, and avoid power waste caused by frequently activating subsequent high-power units.
[0117] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0118] Based on the same inventive concept, this application also provides a voice activity detection device for implementing the aforementioned voice activity detection method. This device is applied to a voice activity detection unit in a voice wake-up system. The voice wake-up system further includes an analog microelectromechanical system (MEMS) microphone unit and a wake-up control unit. The voice activity detection unit is connected to both the analog MEMS microphone unit and the wake-up control unit. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in one or more voice activity detection device embodiments provided below can be found in the limitations of the voice activity detection method described above, and will not be repeated here.
[0119] In one exemplary embodiment, such as Figure 7 As shown, a voice activity detection device is provided, wherein:
[0120] The counting value acquisition module 701 is used to control the microphone unit of the analog microelectromechanical system to periodically turn on to output a voice analog signal, obtain energy count values for at least two consecutive cycles based on the signal amplitude of the voice analog signal, and obtain zero-crossing feature count values for at least two consecutive cycles based on the number of zero-crossings of the voice analog signal.
[0121] The time-series change feature acquisition module 702 is used to obtain energy time-series change features based on the energy count value of the current cycle and the energy count value of the previous cycle, and to obtain zero-crossing time-series change features based on the zero-crossing feature count value of the current cycle and the zero-crossing feature count value of the previous cycle.
[0122] The determination result acquisition module 703 is used to obtain a first voice activity determination result based on the energy timing change characteristics, and to obtain a second voice activity determination result based on the zero-crossing timing change characteristics. The first voice activity determination result and the second voice activity determination result are used to characterize whether there is voice activity.
[0123] The wake-up module 704 of the wake-up control unit is used to output a wake-up signal to the wake-up control unit when both the first voice activity determination result and the second voice activity determination result indicate the presence of voice activity, so as to wake up the wake-up control unit.
[0124] In one embodiment, the count value acquisition module 701 is further configured to: use a comparator to perform positive threshold discrimination and negative threshold discrimination on the voice analog signal of each cycle, respectively, to obtain the positive discrimination result and the negative discrimination result of each cycle; perform logical synthesis on the positive discrimination result and the negative discrimination result of each cycle to determine the level state of the energy indicator signal of each cycle; and count the high level state duration of the energy indicator signal of each cycle to obtain the energy count value of each cycle.
[0125] In one embodiment, the time-series change feature acquisition module 702 is further configured to: perform a sum operation on the energy count value of the current period and the energy count value of the previous period to obtain an energy count sum value; perform a difference operation on the energy count value of the current period and the energy count value of the previous period to obtain an energy count difference value; and obtain energy time-series change features based on the energy count sum value and the energy count difference value.
[0126] In one embodiment, the determination result acquisition module 703 is used to: if the sum of the energy counts is greater than a first preset sum threshold and the difference in energy counts is greater than a first preset difference threshold, then the first voice activity determination result is that there is voice activity; if the sum of the energy counts is less than or equal to the first preset sum threshold, or the difference in energy counts is less than or equal to the first preset difference threshold, then the first voice activity determination result is that there is no voice activity.
[0127] In one embodiment, the time-series change feature acquisition module 702 is further configured to: perform a sum operation on the zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period to obtain a zero-crossing feature count sum value; perform a difference operation on the zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period to obtain a zero-crossing feature count difference value; and obtain the zero-crossing time-series change feature based on the zero-crossing feature count sum value and the zero-crossing feature count difference value.
[0128] In one embodiment, the determination result acquisition module 703 is used to: if the sum of the zero-crossing feature counts is within a second preset threshold range and the difference in the zero-crossing feature counts is greater than a second preset difference threshold, then the second voice activity determination result is that there is voice activity; if the sum of the zero-crossing feature counts is outside the second preset threshold range, or the difference in the zero-crossing feature counts is less than or equal to the second preset difference threshold, then the second voice activity determination result is that there is no voice activity.
[0129] Each module in the aforementioned voice activity detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0130] In one exemplary embodiment, a voice wake-up system is provided. The system includes an analog microelectromechanical system (MEMS) microphone unit, a voice activity detection unit, and a wake-up control unit. The voice activity detection unit is connected to both the analog MEMS microphone unit and the wake-up control unit. The analog MEMS microphone unit is used to output analog voice signals. The voice activity detection unit is used to execute the steps of the method described in the above embodiment. The wake-up control unit is used to execute a pre-set functional operation when a wake-up signal is received.
[0131] The voice activity detection unit can controllably drive the analog microelectromechanical system (MEMS) microphone unit to periodically turn on, so that the analog MEMS microphone unit works periodically. When the analog MEMS microphone unit receives valid voice or valid sound wave input during the working phase, it outputs a voice analog signal.
[0132] The voice activity detection unit, when simulating a voice analog signal output by a microphone unit of a microelectromechanical system (MEMS), obtains at least two consecutive cycles of energy count values based on the signal amplitude of the voice analog signal, and at least two consecutive cycles of zero-crossing feature count values based on the number of zero-crossings of the voice analog signal. Based on the energy count values of the current cycle and the previous cycle, it obtains energy timing change characteristics, and based on the zero-crossing feature count values of the current cycle and the previous cycle, it obtains zero-crossing timing change characteristics. Based on the energy timing change characteristics, it obtains a first voice activity determination result, and based on the zero-crossing timing change characteristics, it obtains a second voice activity determination result. The first and second voice activity determination results are used to characterize whether voice activity exists. If both the first and second voice activity determination results indicate the presence of voice activity, a wake-up signal is output to the wake-up control unit to wake it up.
[0133] The period refers to multiple time periods of equal duration determined from the starting point of the analog microelectromechanical system microphone unit outputting the analog voice signal during a single wake-up activity.
[0134] When a microphone unit of an analog microelectromechanical system receives a valid voice or sound wave input, it outputs an analog voice signal.
[0135] When the microphone unit of the simulated microelectromechanical system outputs a voice analog signal, the voice activity detection unit can connect to extract at least two cycles of voice analog signal as the processing object. Based on the signal amplitude of the voice analog signal in each cycle, the unit performs sampling and processing on the voice analog signal in each cycle, calculates the signal energy in each cycle, and then obtains the energy count value of at least two consecutive cycles.
[0136] When the microphone unit of the simulated microelectromechanical system outputs a voice analog signal, the voice activity detection unit can connect to and extract at least two cycles of voice analog signal as the processing object. It can count the number of times the waveform of the voice analog signal in each cycle crosses the zero level in a unit of time, count the number of zero crossings in each cycle, and then obtain the zero crossing feature count value of at least two consecutive cycles.
[0137] The characteristics of energy temporal variation can be obtained based on the difference and the total continuity between the energy count value of the current period and the energy count value of the previous period.
[0138] The zero-crossing time series change characteristics are obtained based on the difference and the total continuity between the zero-crossing characteristic count value of the current period and the zero-crossing characteristic count value of the previous period.
[0139] Energy time-series variation characteristics can include the sum of energy counts and the difference in energy counts.
[0140] The presence of speech activity in the current environment can be determined based on the relative magnitude of the energy count and value in the energy temporal variation characteristics with a first preset threshold. Fixed noise in the natural environment can be eliminated based on the relative magnitude of the energy count difference in the energy temporal variation characteristics with a first preset difference threshold.
[0141] If there is speech activity in the current environment and the speech activity is not fixed noise in the natural environment, then the first speech activity determination result is that there is speech activity; otherwise, the first speech activity determination result is that there is no speech activity.
[0142] Zero-crossing time-series change characteristics can include the sum of zero-crossing feature counts and the difference between zero-crossing feature counts.
[0143] The presence of speech activity in the current environment can be determined by whether the zero-crossing feature count and value in the zero-crossing timing change features are within a second preset threshold range. Noise interference outside the human voice frequency range can be eliminated by the relative magnitude relationship between the difference in the zero-crossing feature counts in the zero-crossing timing change features and the second preset difference threshold.
[0144] If there is voice activity in the current environment and the voice activity is not noise interference outside the human voice frequency range, then the second voice activity determination result is that there is voice activity; otherwise, the second voice activity determination result is that there is no voice activity.
[0145] When both the first and second voice activity determination results indicate the presence of voice activity, and this activity can be confirmed from both the energy and frequency dimensions of the simulated voice signal, a wake-up signal can be output to the wake-up control unit to activate it and execute a pre-set function. Executing the pre-set function may include performing a security anomaly response, or it may include performing a voice wake-up recognition operation.
[0146] When the pre-set function operation includes performing security anomaly response operation, the wake-up control unit can start evidence collection actions such as taking pictures and recording videos, retain on-site data, realize real-time monitoring and event tracing of risks such as intrusion and abnormality, and complete security early warning and on-site recording.
[0147] When the pre-set function operation includes voice wake-up recognition, the wake-up control unit can include an analog-to-digital converter, a voice feature detection unit, and a keyword detection unit. In this case, the wake-up control unit can perform analog-to-digital conversion, voice detection, and keyword detection on the analog voice signal. Upon recognizing a preset wake-up word, it outputs a valid wake-up command, instructing the electronic device to enter working mode and realizing voice interaction control. The voice feature detection unit uses a low-power analog-to-digital converter instead of a high-precision one for analog-to-digital conversion, significantly reducing energy consumption during the conversion process.
[0148] Specifically, upon receiving a wake-up signal, the wake-up control unit converts the analog voice signal into a digital voice signal, then performs voice feature matching on the digital voice signal. If the matching fails, it indicates that the voice feature detection has failed, and the system does not proceed to the next stage, i.e., keyword detection is not performed. If the matching is successful, it indicates that the voice feature detection has been successful, and the system proceeds to the next stage, where keyword detection is performed on the digital voice signal. If a preset wake-up word is recognized, a valid wake-up command is output to instruct the electronic device to enter the working state.
[0149] In this embodiment, the voice activity detection unit in the voice wake-up system obtains the first and second voice activity determination results based on the temporal change characteristics of the energy count value and zero-crossing feature count value of the current cycle and the previous cycle, so as to determine whether to output a wake-up signal to wake up the wake-up control unit. This can avoid the misjudgment of voice activity caused by judging whether to output a wake-up signal based solely on the energy count value and zero-crossing feature count value of a single cycle. It can improve the accuracy of voice activity detection, reduce the number of false wake-ups, and avoid the power consumption waste caused by frequently activating subsequent high-power units.
[0150] In one embodiment, the voice wake-up system further includes a power management chip connected to an analog microelectromechanical system (MEMS) microphone unit. The power management chip is used to intermittently supply power to the analog MEMS microphone unit when the analog MEMS microphone unit does not receive voice, and to synchronously control the on and off of the voice activity detection unit.
[0151] The power management chip can intermittently power the analog microelectromechanical system microphone unit by controlling the periodic switching of a low dropout regulator (LDO) when the microphone unit does not receive voice.
[0152] For example, the power management chip operates on a 150ms cycle. The first 50ms control the low-dropout linear regulator to turn on, supplying power to the analog MEMS microphone unit and putting it into operation. The next 100ms control the low-dropout linear regulator to turn off, depriving the microphone unit of power and putting it into sleep mode. This 1:2 duty cycle timing control reduces the average power consumption of the analog MEMS microphone unit to one-third of its original value.
[0153] The power management chip provides synchronous power to the voice activity detection unit and the analog microelectromechanical system (MEMS) microphone unit, keeping their operating states synchronized. That is, the voice activity detection unit is turned on when the analog MEMS microphone unit is turned on, and the voice activity detection unit is put into sleep mode when the analog MEMS microphone unit is turned off.
[0154] In this embodiment, the power management chip intermittently supplies power to the analog microelectromechanical system (MEMS) microphone unit when it does not receive voice, which can reduce the static power consumption of the analog MEMS microphone unit and synchronously control the opening and closing of the voice activity detection unit, thereby achieving system-level coordinated energy saving.
[0155] In one embodiment, the pre-set functional operations include analog-to-digital conversion, voice feature detection, and keyword detection. The wake-up control unit includes a low-precision analog-to-digital converter and a high-precision analog-to-digital converter. The low-precision analog-to-digital converter is used to convert the voice analog signal into a digital signal during the voice feature detection stage. The high-precision analog-to-digital converter is used to convert the voice analog signal into a digital signal during the keyword detection stage. The analog-to-digital conversion accuracy of the low-precision analog-to-digital converter is lower than a preset accuracy threshold, and the analog-to-digital conversion accuracy of the high-precision analog-to-digital converter is higher than the preset accuracy threshold.
[0156] You can set a preset accuracy threshold according to the actual situation.
[0157] Since the function of human voice feature detection is to identify whether the input sound has basic human voice features, the accuracy requirements of the analog-to-digital converter are relatively relaxed. Therefore, in the human voice feature detection stage, a low-power analog-to-digital converter can be used to convert the voice analog signal into a digital signal, which can significantly reduce the energy consumption in the analog-to-digital conversion process.
[0158] Among them, low-power analog-to-digital converters are, for example, successive approximation register analog-to-digital converters (SAR ADCs) with low sampling rates. When operating at an 8kHz sampling rate, SAR ADCs consume only tens of microamps or even less power, thus significantly reducing energy consumption in the human voice detection stage.
[0159] To better understand the above method, the following describes in detail an application embodiment of the voice activity detection method and voice wake-up system of this application.
[0160] With the rapid development of artificial intelligence technology, voice wake-up and voice interaction functions have become standard features of modern wearable devices. Since these devices generally use lithium batteries for power, effectively extending battery life has become one of the key technical challenges. Therefore, voice wake-up systems need to minimize power consumption while ensuring full functionality.
[0161] Traditional voice wake-up system architecture, such as Figure 1As shown, it includes the following core units: an analog micro-electro-mechanical system (MIC MEMS) microphone unit, a voice activity detection (VAD) unit, a high-precision analog-to-digital converter (e.g., a summation-delta analog-to-digital converter, or sigma-delta ADC for short), and a keyword detection unit.
[0162] Traditional voice wake-up systems typically involve two stages. The first stage involves a voice activity detection unit that monitors the analog audio signal output from the microphone unit of an analog microelectromechanical system (MEMS). When the voice activity detection unit detects potential voice activity based on the analog audio signal, it generates a wake-up signal. This wake-up signal activates the subsequent high-precision analog-to-digital converter (ADC) and keyword detection unit. The second stage involves the activated ADC performing high-fidelity sampling of the input analog audio signal and converting it into a digital signal. This digital signal is then passed to the keyword detection unit for accurate keyword recognition. Upon recognizing a preset wake-up word, the keyword detection unit outputs a valid wake-up command to instruct the wearable device to enter working mode.
[0163] However, traditional voice wake-up system architectures suffer from significant power consumption issues: First, the pre-amplification circuitry and other auxiliary circuits built into the analog MEMS microphone unit typically cause its static power consumption to exceed 100μA, which conflicts with the ultra-low power design goal. Second, the voice activity detection unit can only make a rough judgment on the analog voice signal, resulting in a high false detection rate and frequent false wake-ups of subsequent high-power units. Furthermore, the keyword detection unit relies on a high-precision analog-to-digital converter and a high-performance central processing unit (CPU) or neural processing unit (NPU), resulting in extremely high operating power consumption; a large number of false wake-up events lead to significant power waste.
[0164] To alleviate this problem, the industry has proposed an improvement: adding a voice feature detection step before the keyword detection unit, such as... Figure 2As shown, since the computational complexity required for the voice feature detection unit is much lower than that for keyword recognition, a low-power, low-computing-power central processing unit or neural network processor can be used to implement voice feature detection. Only when the voice feature detection unit confirms that the input signal contains valid voice components will the voice wake-up system enter the more energy-intensive keyword detection stage. Nevertheless, traditional technologies still have significant shortcomings: First, the high static power consumption problem of analog microelectromechanical system microphone units has not been effectively solved; second, the continued use of high-precision analog-to-digital converters in the voice feature detection stage results in unnecessary power consumption; third, the excessively high false wake-up rate of the voice activity detection unit leads to wasted power.
[0165] To address the technical problems of excessive power consumption and false wake-up of voice activity detection units in traditional voice wake-up systems, this embodiment provides a voice activity detection method and a voice activity detection system. The voice activity detection system provided in this embodiment is based on... Figure 2 The system architecture has been optimized and upgraded, mainly in the following three aspects:
[0166] Firstly, optimize the analog-to-digital conversion scheme for voice feature detection. Since the function of voice feature detection is to identify whether the input sound possesses basic human voice characteristics, the accuracy requirements for the analog-to-digital converter are relatively relaxed. Therefore, a successive approximation register-type analog-to-digital converter can be introduced as a dedicated analog-to-digital converter within the voice feature detection module. In contrast, Figure 2 The system still uses a high-precision analog-to-digital converter (ADC) for analog-to-digital conversion during the voice feature detection stage, resulting in unnecessary power consumption waste. The typical power consumption of a high-precision ADC can reach several milliamps, while the successive approximation register-type ADC with a low sampling rate selected in this embodiment consumes only tens of microamps or even less under an 8kHz sampling rate, thus significantly reducing the energy consumption in the voice detection stage.
[0167] Secondly, a periodic power management strategy is implemented. The power management chip can intermittently power the analog MEMS microphone unit by periodically switching a low-dropout regulator (LDO) on and off when the microphone unit is not receiving voice signals. Specifically, the power management chip operates in 150ms cycles. For the first 50ms, the LDO is turned on to power the microphone unit, putting it in operation. For the next 100ms, the LDO is turned off, putting the microphone unit into a sleep state. This 1:2 duty cycle timing control reduces the average power consumption of the analog MEMS microphone unit to one-third of its original level. In addition, the power management chip provides synchronous power to the voice activity detection unit and the analog MEMS microphone unit, so that the working state of the voice activity detection unit and the working state of the analog MEMS microphone unit are synchronized. That is, the voice activity detection unit works when the analog MEMS microphone unit is turned on, and the voice activity detection unit also goes into sleep mode when the analog MEMS microphone unit is turned off, thereby achieving system-level coordinated energy saving.
[0168] Thirdly, a voice activity detection method is provided to improve the detection accuracy of the voice activity detection unit and avoid unnecessary system power consumption waste caused by frequently waking up the next-level unit. Specifically, a short-time energy (STE) detection module and a zero-crossing point (ZCP) detection module are added to the voice activity detection unit to detect and process the energy and frequency of the voice analog signal.
[0169] The operating mechanism of the voice wake-up system provided in this embodiment is as follows: Figure 8 As shown:
[0170] Phase 1: Periodic Speech Activity Detection (Speech Activity Detection Methods):
[0171] The control logic within the analog module of the voice activity detection unit is responsible for controlling the periodic start and stop of the low-dropout linear regulator and the digital module. When the low-dropout linear regulator is turned on, it supplies power to the analog microelectromechanical system (MEMS) microphone unit, enabling it to acquire ambient sound signals and convert them into differential analog electrical signals MIC+ and MIC- (which can be referred to as voice analog signals), which are then output to the programmable gain amplifier (PGA) in the digital module of the voice activity detection unit. Simultaneously with the low-dropout linear regulator's activation, the PGA, short-time energy detection module, and zero-crossing detection module are also activated. The PGA performs bandpass filtering and amplification on the differential signals MIC+ and MIC-, and then transmits the processed signals to the short-time energy detection module and the zero-crossing detection module. Among them, MIC+ represents the microphone positive differential signal, and MIC- represents the microphone negative differential signal; VSS stands for Voltage Source Sink.
[0172] The specific process of short-time energy detection sampling is as follows: Figure 5 As shown, at least two cycles of analog speech signal are extracted from the output signal VO of the programmable operational amplifier as the processing object. For the Nth cycle of the analog speech signal, comparators 1 and 2 can be used to perform positive threshold discrimination and negative threshold discrimination on the positive and negative half-cycle energy of the analog speech signal, respectively.
[0173] When the amplitude of the positive half-cycle signal of the analog speech signal is greater than the positive threshold of comparator 1 (which can be called the preset threshold Vref1), comparator 1 outputs a high-level signal STE_OUTP. That is, the positive discrimination result of the Nth cycle is that the output signal STE_OUTP is set to a high level, indicating that the positive half-cycle energy of the analog speech signal exceeds the limit. When the amplitude of the negative half-cycle signal of the analog speech signal is less than the negative threshold of comparator 2 (which can be called the preset threshold Vref2), comparator 2 outputs a high-level signal STE_OUTN. That is, the negative discrimination result of the Nth cycle is that the output signal STE_OUTN is set to a high level, indicating that the negative half-cycle energy of the analog speech signal exceeds the limit.
[0174] A logical OR operation is performed on the positive and negative discrimination results of the Nth cycle to determine the level state of the energy indicator signal (STE_OUT signal) for the Nth cycle; specifically, a logical OR operation is performed on the output signals STE_OUTP and STE_OUTN to obtain the STE_OUT signal. The duration of the high-level state of the energy indicator signal for the Nth cycle is counted according to the high-frequency clock (CLK) to obtain the energy count value E for the Nth cycle. This energy count value E reflects the energy level of the voice analog signal in the Nth cycle.
[0175] The specific process of zero-crossing detection sampling is as follows: Figure 6 As shown, at least two cycles of the speech analog signal are connected to the output signal VO of the programmable operational amplifier as the processing object. For the speech analog signal of the Nth cycle, comparators 3 and 4 can be used to perform positive threshold discrimination of zero-crossing feature and negative threshold discrimination of zero-crossing feature, respectively.
[0176] When the amplitude of the analog speech signal is greater than the positive zero-crossing threshold of comparator 3 (which can be called the preset threshold Vref3), comparator 3 outputs a high-level signal ZCP_OUTP. This means that the positive zero-crossing characteristic determination result for the Nth cycle is that the output signal ZCP_OUTP is high, indicating that a positive zero-crossing transition feature has been detected. When the amplitude of the analog speech signal is less than the negative zero-crossing threshold of comparator 4 (which can be called the preset threshold Vref4), comparator 4 outputs a high-level signal ZCP_OUTN. This means that the negative zero-crossing characteristic determination result for the Nth cycle is that the output signal ZCP_OUTN is high, indicating that a negative zero-crossing transition feature has been detected.
[0177] A logical OR operation is performed on the positive and negative zero-crossing feature discrimination results of the Nth cycle to determine the level state of the zero-crossing indicator signal (ZCP_OUT signal) of the Nth cycle. Specifically, a logical OR operation is performed on the output signals ZCP_OUTP and ZCP_OUTN to obtain the ZCP_OUT signal. The rising edge of the zero-crossing indicator signal pulse of each cycle is used as the counting clock to count zero crossings, obtaining the zero-crossing feature count value Z for each cycle. This zero-crossing feature count value Z can reflect the frequency characteristics of the voice analog signal of the Nth cycle.
[0178] Wake-up determination mechanism: After short-term energy detection sampling and zero-crossing detection sampling are completed, the energy count value and zero-crossing feature count value are transmitted to the wake-up signal module in the voice activity detection unit for wake-up determination. The determination process is as follows: Figure 9 As shown.
[0179] For short-term energy detection, the energy count value E[N] of the current cycle is added to the energy count value E[N-1] of the previous cycle stored in the register to obtain the energy count sum SUM_E = E[N] + E[N-1]. When the energy count sum SUM_E is greater than the first preset sum threshold TH_energy_sum, the sum judgment output signal ste_sum is determined to be high (ste_sum=1), otherwise it is output low (ste_sum=0). At the same time, the energy count value E[N] of the current cycle is subtracted from the energy count value E[N-1] of the previous cycle to obtain the energy count difference DIFF_E = E[N] - E[N-1]. When the energy count difference DIFF_E is greater than the first preset difference threshold TH_energy_diff, the difference judgment output signal ste_diff is determined to be high (ste_diff=1), otherwise it is output low (ste_diff=0). When both output signals ste_sum and ste_diff are high, the short-time energy detection output is high, indicating that the energy of the input signal has reached the condition to wake up the next level. The summation operation is used to determine whether there is sound activity in the current environment, while the difference operation is used to eliminate fixed noise in the natural environment. Since the energy base of fixed noise in the natural environment is relatively stable with small fluctuations, only when the ambient sound energy shows a significant increase (i.e., a large difference) is it determined to be speech activity, thus effectively improving anti-interference capability.
[0180] For zero-crossing detection, the zero-crossing feature count value Z[N] of the current period and the zero-crossing feature count value of the previous period are also used. Add them together to get the zero-crossing feature count and sum. If the zero-crossing feature count and value SUM_Z are within the second preset and threshold range [TH_zero_min, TH_zero_max], then the zero-crossing point is determined and the judgment signal zcp_sum is output at a high level (zcp_sum=1); otherwise, it is output at a low level (zcp_sum=0). Simultaneously, the zero-crossing feature count value Z[N] of the current cycle and the zero-crossing feature count value of the previous cycle are compared. Subtracting the two values yields the difference in zero-crossing feature counts. If the zero-crossing feature count difference DIFF_Z is greater than the second preset difference threshold TH_zero_diff, then the zero-crossing difference judgment signal zcp_diff is determined to output a high level (zcp_diff=1); otherwise, it is low (zcp_diff=0). When both signals zcp_sum and zcp_diff are high, the zero-crossing detection output is high. By performing summation and difference operations, noise interference outside the human voice frequency range can be effectively eliminated. Assuming the human voice frequency range is 500Hz to 5kHz and the sampling time is 1 second, then the number of zero-crossings should be between 1 / 1 / 500=200 times (lower limit) and 1 / 1 / 5000=1000 times (upper limit). If the sum of the two sampling results falls within the range of 2×200=400 times to 2×1000=2000 times, then human voice activity can be considered to exist.
[0181] When both the short-time energy detection and zero-crossing detection criteria are met, the wake-up signal 4 output by the system's voice activity detection unit is set to a high level to activate the next-level unit. After detection is completed, the energy count and zero-crossing feature count values of the current cycle are updated in the register for use in the next detection.
[0182] Phase Two: Voice Feature Detection
[0183] When wake-up signal 4 goes high, the voice feature detection unit is activated. This unit enables a successive approximation register-type analog-to-digital converter (ADC) to perform analog-to-digital conversion on the output signal VO of the programmable operational amplifier, converting the analog speech signal into a digital speech signal for voice feature matching. If the match is successful, the wake-up signal 5 output by the voice feature detection unit is set to high to activate the next unit; if the match fails, the system shuts down and enters standby mode, waiting for the next detection. Furthermore, if voice feature detection fails, the system dynamically adjusts the reference voltage threshold in short-term energy detection: increasing the positive threshold in the energy count calculation process and decreasing the negative threshold. Specifically, the positive threshold Vref1 can be increased by one level and the negative threshold Vref2 can be decreased by one level. This mechanism requires a greater sound intensity to trigger the successive approximation register-type ADC wake-up for the next energy detection. By successively increasing the wake-up threshold, false wake-ups caused by excessive ambient noise energy can be effectively suppressed, further improving detection accuracy. Among them, wake-up signal 4 is the wake-up signal output by the voice activity detection unit when both the short-time energy detection and zero-crossing detection conditions are met, and wake-up signal 5 is the wake-up signal output by the voice feature detection unit after the voice feature matching is successful.
[0184] Since human voice feature detection only requires the extraction of basic speech feature information and the data volume is relatively small, a low-power small-core CPU or neural network processor can be used for processing to reduce overall power consumption. At the same time, the successive approximation register-type analog-to-digital converter can adopt a low sampling rate (such as 8 kHz) and 12-bit precision specification, so that the power consumption is controlled at the tens of microamps level, meeting the requirements of low-power applications.
[0185] Phase Three: Keyword Detection
[0186] When wake-up signal 5 goes high, the keyword detection unit is activated. This unit enables a high-precision analog-to-digital converter (ADC) to perform high-precision analog-to-digital conversion on the differential signals MIC+ and MIC- of the analog microelectromechanical system (MEMS) microphone unit. The differential signals are filtered and amplified by a microphone programmable gain amplifier (MIC_PGA) before being sent to the ADC to be converted into digital signals, which are then accurately identified by the keyword detection unit. Keyword detection requires processing a large amount of data, therefore a high-core CPU is used for computation. A successful match indicates that the electronic device has entered the working state and the system's voice interaction function has been activated; otherwise, the system is shut down and waits for the next detection cycle.
[0187] The complete implementation process of the voice wake-up system provided in this embodiment is as follows: Figure 10 As shown, a three-level progressive detection mechanism is used to achieve significant power consumption optimization while ensuring detection accuracy.
[0188] Specifically, the voice activity detection unit performs short-time energy detection and zero-crossing detection on the voice analog signal output by the analog microelectromechanical system microphone unit. When the output results of the short-time energy detection and zero-crossing detection are not both 1, that is, when the first voice activity determination result and the second voice activity determination result do not both indicate the presence of voice activity, the short-time energy detection and zero-crossing detection are activated again after 100ms. When the output results of the short-time energy detection and zero-crossing detection are both 1, that is, when the first voice activity determination result and the second voice activity determination result both indicate the presence of voice activity, a wake-up signal is output to the wake-up control unit.
[0189] When the pre-set functions in the wake-up control unit include voice wake-up recognition, the wake-up signal can wake up the successive approximation register-type analog-to-digital converter in the wake-up control unit, converting the voice analog signal into a voice digital signal. Then, the voice feature detection unit in the wake-up control unit performs voice feature matching on the voice digital signal. If the voice feature matching is unsuccessful, it indicates that the voice feature detection has failed. The positive and negative thresholds are then modified, and the successive approximation register-type analog-to-digital converter is turned off. Detection is restarted after 100ms. If the voice feature matching is successful, it indicates that the voice feature detection is successful. The keyword detection unit then performs keyword detection on the voice digital signal. When the keyword is correctly recognized, that is, when the preset wake-up word is recognized, a valid wake-up command is output, waking up the voice wake-up system. The input audio analog signal is filtered and amplified by the microphone programmable gain amplifier and then sent to the high-precision analog-to-digital converter to be converted into an audio digital signal, obtaining recording information (audio analog-to-digital conversion recording). After the command is executed, the voice wake-up system is turned off.
[0190] When the pre-set functions in the wake-up control unit include performing security anomaly response operations, the wake-up signal can wake up the wake-up control unit, enabling it to start taking photos, recording videos, and other evidence collection actions, retaining on-site data, realizing real-time monitoring and event tracing of risks such as intrusion and abnormal activity, and completing security early warning and on-site recording.
[0191] The voice activity detection method and voice wake-up system provided in this embodiment have the following beneficial effects:
[0192] (1) Reduce power consumption in the human voice feature detection stage: By using a low-power analog-to-digital converter to replace a high-precision analog-to-digital converter, the energy consumption in the analog-to-digital conversion process is significantly reduced.
[0193] (2) Reduce the static power consumption of the analog microelectromechanical system microphone unit: By implementing a periodic power management strategy and controlling the on / off timing of the low dropout linear regulator, intermittent power supply to the analog microelectromechanical system microphone unit is achieved, which greatly reduces the average power consumption of the analog microelectromechanical system microphone unit.
[0194] (3) Improve the accuracy of voice activity detection: By introducing short-time energy detection and zero-crossing detection mechanisms, the detection accuracy of the voice activity detection unit is improved, the number of false wake-ups is reduced, and energy waste caused by frequent activation of subsequent high-power units is avoided.
[0195] (4) Achieve system-level collaborative energy saving: Through the timing coordination control of each functional unit, a complete low-power management system is constructed to maximize the overall energy saving effect.
[0196] In one embodiment, as well as such Figure 11The electronic device shown includes a processor and a voice wake-up system as described in the above embodiment. The processor is connected to the wake-up control unit of the voice wake-up system. The processor is used to operate when a valid wake-up command is received. The valid wake-up command is an instruction output when a preset wake-up word is recognized during keyword detection of the voice analog signal.
[0197] Electronic devices can be, for example, wearable devices.
[0198] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0199] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0200] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0201] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0202] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0203] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for detecting speech activity, characterized in that, A voice activity detection unit is applied to a voice wake-up system, the voice wake-up system further comprising an analog microelectromechanical system (MEMS) microphone unit and a wake-up control unit, the voice activity detection unit being connected to the analog MEMS microphone unit and the wake-up control unit respectively, the method comprising: By controlling the microphone unit of the analog microelectromechanical system to periodically turn on to output an analog voice signal, at least two consecutive energy count values are obtained based on the signal amplitude of the analog voice signal, and at least two consecutive zero-crossing feature count values are obtained based on the number of zero-crossings of the analog voice signal. Based on the energy count value of the current cycle and the energy count value of the previous cycle, the energy time series change characteristics are obtained, and based on the zero-crossing characteristic count value of the current cycle and the zero-crossing characteristic count value of the previous cycle, the zero-crossing time series change characteristics are obtained. Based on the energy timing change characteristics, a first voice activity determination result is obtained, and based on the zero-crossing timing change characteristics, a second voice activity determination result is obtained. The first voice activity determination result and the second voice activity determination result are used to characterize whether voice activity exists. If both the first voice activity determination result and the second voice activity determination result indicate the presence of voice activity, a wake-up signal is output to the wake-up control unit to wake up the wake-up control unit.
2. The method according to claim 1, characterized in that, The step of obtaining at least two consecutive energy count values based on the signal amplitude of the simulated speech signal includes: A comparator is used to perform positive and negative threshold discrimination on the speech analog signal of each cycle, so as to obtain the positive and negative discrimination results for each cycle. The positive and negative discrimination results of each cycle are logically synthesized to determine the level state of the energy indication signal for each cycle; The duration of the high-level state of the energy indicator signal in each cycle is counted to obtain the energy count value for each cycle.
3. The method according to claim 1, characterized in that, The process of obtaining energy temporal variation characteristics based on the energy count value of the current period and the energy count value of the previous period includes: The energy count value of the current period is summed with the energy count value of the previous period to obtain the sum of the energy counts. The energy count difference is obtained by subtracting the energy count value of the current cycle from the energy count value of the previous cycle. Based on the sum of the energy counts and the difference between the energy counts, the energy temporal variation characteristics are obtained.
4. The method according to claim 3, characterized in that, The step of obtaining the first speech activity determination result based on the energy temporal change characteristics includes: If the sum of the energy counts is greater than the first preset threshold, and the difference in the energy counts is greater than the first preset difference threshold, then the first voice activity determination result is that there is voice activity. If the sum of the energy counts is less than or equal to a first preset threshold, or the difference in the energy counts is less than or equal to a first preset difference threshold, then the first voice activity determination result is that there is no voice activity.
5. The method according to any one of claims 1-4, characterized in that, The step of obtaining the zero-crossing time series change characteristics based on the zero-crossing characteristic count value of the current period and the zero-crossing characteristic count value of the previous period includes: The zero-crossing feature count value of the current period and the zero-crossing feature count value of the previous period are summed to obtain the zero-crossing feature count sum value. The zero-crossing feature count value of the current period is subtracted from the zero-crossing feature count value of the previous period to obtain the zero-crossing feature count difference value. The zero-crossing time-series change characteristics are obtained based on the sum of the zero-crossing feature counts and the difference between the zero-crossing feature counts.
6. The method according to claim 5, characterized in that, The step of obtaining the second speech activity determination result based on the zero-crossing timing change characteristics includes: If the sum of the zero-crossing feature counts is within the second preset threshold range, and the difference between the zero-crossing feature counts is greater than the second preset threshold, then the second speech activity determination result is that speech activity exists. If the zero-crossing feature count and value are outside the second preset threshold range, or if the difference in the zero-crossing feature count is less than or equal to the second preset difference threshold, then the second voice activity determination result is that there is no voice activity.
7. A voice wake-up system, characterized in that, The system includes an analog microelectromechanical system (MEMS) microphone unit, a voice activity detection unit, and a wake-up control unit; the voice activity detection unit is connected to both the analog MEMS microphone unit and the wake-up control unit. The analog microelectromechanical system microphone unit is used to output analog voice signals; The voice activity detection unit is configured to perform the steps of the method according to any one of claims 1 to 6; The wake-up control unit is used to execute pre-set functional operations upon receiving a wake-up signal.
8. The voice wake-up system according to claim 7, characterized in that, The voice wake-up system also includes a power management chip, which is connected to the analog microelectromechanical system microphone unit. The power management chip is used to intermittently supply power to the analog microelectromechanical system microphone unit when the microphone unit does not receive voice, and to synchronously control the opening and closing of the voice activity detection unit.
9. The voice wake-up system according to claim 8, characterized in that, The preset functional operations include analog-to-digital conversion, voice feature detection, and keyword detection. The wake-up control unit includes a low-precision analog-to-digital converter and a high-precision analog-to-digital converter. The low-precision analog-to-digital converter is used to convert the voice analog signal into a digital signal during the voice feature detection stage. The high-precision analog-to-digital converter is used to convert the voice analog signal into a digital signal during the keyword detection stage. The analog-to-digital conversion accuracy of the low-precision analog-to-digital converter is lower than a preset accuracy threshold, and the analog-to-digital conversion accuracy of the high-precision analog-to-digital converter is higher than the preset accuracy threshold.
10. An electronic device, characterized in that, The system includes a processor and a voice wake-up system as described in any one of claims 7 to 9, wherein the processor is connected to the wake-up control unit of the voice wake-up system, and the processor is configured to operate upon receiving a valid wake-up command; the valid wake-up command is an instruction output when a preset wake-up word is identified during keyword detection of the voice analog signal.