A method and system for detecting one-dimensional discrete-time signal activity using adaptive sensing
Through the adaptively perceived one-dimensional discrete time signal activity detection method, pre-emphasis, frame division, windowing and dynamic energy threshold adjustment, combined with the feedback of the later-stage neural network, the problems of low signal detection accuracy and high power consumption caused by fixed thresholds are solved, and efficient and low-power signal activity detection is achieved.
Patent Information
- Application Number
- CN202411520261.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-29
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2044-10-29
AI Technical Summary
In the existing end-side devices for identifying one-dimensional discrete time signals based on fixed thresholds, signal detection accuracy is low and power consumption is limited, so it is impossible to effectively filter invalid signals to reduce the power consumption of neural network inference.
Adaptive sensing of one-dimensional discrete time signal activity detection method is adopted, through pre-emphasis, signal framing, windowing and dynamic energy threshold adjustment, combined with the feedback of the later-stage identification neural network, the energy threshold is dynamically adjusted to improve detection accuracy and reduce invalid inference power consumption.
It improves the accuracy of signal activity detection, reduces equipment power consumption, and realizes efficient signal activity detection in an adaptive environment.
Smart Images

Figure CN119400202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech signal processing, and in particular to a method and system for detecting one-dimensional discrete-time signal activity using adaptive perception. Background Art
[0002] A one-dimensional time signal is a one-dimensional data sequence that varies over time. It has widespread applications in many fields, such as audio signal processing and data transmission in communications, electrocardiograms and electroencephalograms in medicine and biomedical engineering, and stock price analysis in finance and economics. Neural networks have become a key pillar of artificial intelligence, achieving significant success in computer vision, natural language processing, and speech recognition. They are driving innovation and progress in science and technology, and have demonstrated broad application potential in healthcare, finance, transportation, and other fields. They are considered a key technology driving the development of an intelligent society. Key issues for on-device neural network processors include balancing performance and power consumption, model compatibility, low-latency real-time inference capabilities, security and privacy protection, and flexibility to adapt to diverse application scenarios. While achieving high performance, processors must effectively support models of varying scale and complexity, while offering low power consumption and latency to meet the real-time responsiveness and long-term use requirements of on-device devices. Furthermore, given the diversity and security requirements of on-device environments, processors must also incorporate robust security and privacy protection mechanisms to mitigate potential security threats and privacy breaches.
[0003] Reducing the power consumption of network inference is crucial for on-device AI deployment. Devices are often limited by battery life and heat dissipation capabilities, and high-power inference processes can accelerate battery drain and cause device overheating. By reducing power consumption, device-side devices can extend their usability, improve performance stability, and be more suitable for deployment in mobile environments and wireless networks, providing a better user experience and a wider range of application scenarios. Furthermore, low-power AI processors help reduce energy consumption, lower environmental impact, and align with sustainable development trends. Several methods exist to reduce power consumption associated with network inference on the device side: model compression and lightweighting, for example, using techniques such as pruning, quantization, distillation, and sparse representation to reduce model parameters and computational complexity, thereby reducing computational requirements during inference and thus reducing power consumption. Alternatively, dynamic allocation of computing resources can be implemented, dynamically adjusting resource allocation based on the actual inference load to reduce power consumption during idle or low-load conditions.
[0004] The aforementioned technical means of reducing end-side device power consumption are all implemented during the neural network inference phase. However, in real-world applications, the significant portion of end-side device power consumption is driven by the large amount of invalid inferences triggered by a continuous stream of input signals. This is the always-on pre-processing portion of the neural network inference system. Existing end-side devices that use neural networks to identify one-dimensional discrete-time signals mostly use experience-based, fixed-threshold activity detection methods and systems that have poor environmental adaptability. These methods use a manually set fixed threshold, which significantly impacts filtering accuracy, to filter out invalid signals whose signal and amplitude-related features fall below the threshold and are useless for subsequent neural network inference. This allows for data processing prior to neural network inference, thus reducing power consumption. However, setting a fixed threshold reduces signal detection accuracy and reduces power consumption to a limited extent. Summary of the Invention
[0005] In response to the above problems, the present invention proposes an adaptively perceiving one-dimensional discrete-time signal activity detection method and system, the purpose of which is to filter out the invalid parts of the continuous input signal before neural network inference, reduce the activation frequency of the subsequent neural network, and thus reduce the power consumption of the terminal device due to invalid inference. At the same time, the present invention proposes an environment-adaptive threshold adjustment mechanism to improve the efficiency and accuracy of signal activity detection.
[0006] In one aspect, the present invention provides a method for detecting activity in a one-dimensional discrete-time signal using noise adaptive sensing, the method comprising:
[0007] Step S1: Pre-emphasis. Pre-emphasis compensates for the spectrum attenuation of a one-dimensional time signal by enhancing high-frequency components, thereby making the high-frequency portion of the signal more prominent. The pre-emphasis processing method is: y[n] = x[n] - αx[n-1], where x[n] is the original signal at the current moment, x[n-1] is the original signal at the previous moment, y[n] is the signal after pre-emphasis at the current moment, and α is the pre-emphasis coefficient.
[0008] Step S2: signal framing, dividing the one-dimensional time signal stream after pre-emphasis processing into multiple short time frames of fixed length;
[0009] Step S3: Windowing the frame signal. Windowing reduces the influence of both ends of the frame on energy calculation by applying a weight to the edge of the frame. The window functions used include Hamming window and Hanning window. The expression of Hamming window is: n=0,1,...N-1, where N is the number of sampling points in each frame. After windowing the signal frame x[n], the windowed signal x is obtained. ω [n],x ω [n]=x[n]·ω[n];
[0010] Step S4: Calculate the short-time energy integral of the single-frame signal. For each frame signal x ω [n], its short-time energy integral E is expressed as: Where N is the number of sampling points contained in each frame;
[0011] Step S5: Dynamic energy threshold adjustment. The energy threshold adjustment includes two parallel sub-processes: one is adaptive following threshold adjustment, in which the energy threshold is always updated in the direction of following the signal amplitude; the other is post-stage recognition neural network assisted adjustment. After the post-stage neural network recognizes an active frame or an inactive frame, the threshold is updated in the direction that is more likely to produce an active frame.
[0012] Step S6: Signal activity detection. After the above-mentioned energy threshold adjustment, a decision threshold is set. A larger threshold is obtained by adding a constant to the energy threshold obtained after the above-mentioned threshold adjustment as the decision threshold. Then, the decision threshold is used to perform signal activity detection to obtain the real-time classification result of the frame, that is, the short-time energy integral of the frame is compared with the decision threshold. A frame with a short-time energy integral greater than or equal to the decision threshold is classified as an active frame, and a frame with a short-time energy integral less than the decision threshold is classified as an inactive frame.
[0013] Furthermore, the adaptive following threshold adjustment includes:
[0014] Determining the initial threshold: Adaptive following determines whether and how to adjust the threshold based on the results of frame judgments over a period of time. This period, measured in frames, is called the adjustment window. At the beginning of the process, an initial threshold is manually set based on experience.
[0015] Cumulative judgment result: When the number of frames does not reach the adjustment window length, the short-time energy integral of the frame is compared with the current energy threshold to determine whether the frame is an active frame or an inactive frame. At the same time, the number of active frames is continuously accumulated. When the short-time energy integral of the frame is greater than the current energy threshold, the frame is an active frame, and the active frame count is increased by 1.
[0016] Determine the direction of threshold adjustment: After the number of frames reaches the adjustment window length, the number of active frames in the accumulated number of frames of the past adjustment window length is used to determine whether the current energy threshold setting is appropriate. Two values N are set based on actual experience. less and N more , N less <N more , N less To lower the threshold, N more To increase the threshold, when the number of active frames is less than N less When the number of active frames is greater than N, it means that the number of active frames is too small, that is, the current energy threshold is too high and the threshold needs to be lowered. moreWhen the number of active frames is N, it means that there are too many active frames, that is, the current energy threshold is too low and the threshold needs to be raised. less and N more When it is between , it means the number of active frames is appropriate, that is, the current energy threshold is appropriate and does not need to be adjusted;
[0017] Threshold adjustment: After determining the threshold adjustment direction, select a fixed threshold adjustment range or a variable threshold adjustment range bound to the current energy threshold to adjust the threshold.
[0018] Furthermore, the auxiliary adjustment of the post-stage recognition neural network includes:
[0019] Determine the direction of threshold adjustment: Frame signals that are considered active frames by SAD will be sent to the subsequent recognition neural network. After accumulating active frames that meet the minimum inference frame length, the subsequent recognition neural network will perform an inference. The inference result will be fed back to SAD to determine whether the frames sent to the subsequent stage before are active frames. If the inference feedback of the subsequent neural network is: SAD classification is correct, SAD will lower the threshold to ensure that subsequent signals are more easily sent to the subsequent neural network and avoid the loss of valid signals; if the feedback is: SAD classification is wrong, SAD will increase the threshold to ensure that subsequent invalid signals will not be sent to the subsequent neural network again, so as to reduce the power consumption of false inference;
[0020] Threshold adjustment: After determining the threshold adjustment direction, SAD will immediately correct the energy threshold. The threshold adjustment amplitude can be fixed or non-fixed.
[0021] On the other hand, the present invention also provides an adaptive sensing one-dimensional discrete time signal activity detection system, the system comprising a pre-emphasis module, a data buffer module, a windowing module, a short-time energy integral extraction module, a following threshold update module and a decision module;
[0022] Pre-emphasis module: The pre-emphasis module compensates for the spectrum attenuation of the one-dimensional time signal by enhancing the high-frequency components, making the high-frequency part of the signal more prominent. The pre-emphasis processing method is: y[n] = x[n] - αx[n-1], where x[n] is the original signal at the current moment, x[n-1] is the original signal at the previous moment, y[n] is the signal after pre-emphasis at the current moment, and α is the pre-emphasis coefficient;
[0023] Data cache module: The signal stream processed by the pre-emphasis module is continuously fed into the data cache module. This module uses different memories to ping-pong store data of different half frames of the same channel, and stores data of different channels in different address segments of the same memory. This caching process will also complete the signal framing operation;
[0024] Windowing module: The windowing module uses a window function to apply a weight to the edge of the frame, so that the influence of the two ends of the frame on the energy calculation is reduced. The window function includes Hamming window and Hanning window. The expression of Hamming window is: n=0,1,...N-1, where N is the number of sampling points in each frame. After windowing the signal frame x[n], the windowed signal x is obtained. ω [n],x ω [n]=x[n]·ω[n];
[0025] Short-time energy integral extraction module: The short-time energy integral extraction module calculates the short-time energy integral of the frame signal after being framed by the data cache module. For each frame signal x ω [n], its short-time energy integral E is expressed as: Where N is the number of sampling points contained in each frame;
[0026] Following threshold update module: This module completes the adaptive following threshold adjustment and the adaptive threshold adjustment assisted by the inference result feedback of the post-stage recognition neural network through the short-term energy integration provided by the energy integration extraction module;
[0027] Decision module: A decision threshold is set according to the following energy threshold provided by the following threshold update module. A larger threshold is obtained by adding a constant to the above-mentioned following energy threshold as the decision threshold. The decision threshold is then used to detect signal activity and obtain the real-time classification result of the frame. That is, the short-time energy integral of the frame is compared with the decision threshold. If the short-time energy integral is greater than or equal to the decision threshold, it is classified as an active frame. If the short-time energy integral is less than the decision threshold, it is classified as an inactive frame. The classification result will be used to control whether the frame data is output or not.
[0028] Furthermore, the adaptive following threshold adjustment includes:
[0029] Determining the initial threshold: Adaptive following determines whether and how to adjust the threshold based on the results of frame judgments over a period of time. This period, measured in frames, is called the adjustment window. At the beginning of the process, an initial threshold is manually set based on experience.
[0030] Cumulative judgment result: When the number of frames does not reach the adjustment window length, the short-time energy integral of the frame is compared with the current energy threshold to determine whether the frame is an active frame or an inactive frame. At the same time, the number of active frames is continuously accumulated. When the short-time energy integral of the frame is greater than the current energy threshold, the frame is an active frame, and the active frame count is increased by 1.
[0031] Determine the direction of threshold adjustment: After the number of frames reaches the adjustment window length, the number of active frames in the accumulated number of frames of the past adjustment window length is used to determine whether the current energy threshold setting is appropriate. Two values N are set based on actual experience. less and N more , N less <N more , N less To lower the threshold, N more To increase the threshold, when the number of active frames is less than N less When the number of active frames is greater than N, it means that the number of active frames is too small, that is, the current energy threshold is too high and the threshold needs to be lowered. more When the number of active frames is N, it means that there are too many active frames, that is, the current energy threshold is too low and the threshold needs to be raised. less and N more When it is between , it means the number of active frames is appropriate, that is, the current energy threshold is appropriate and does not need to be adjusted;
[0032] Threshold adjustment: After determining the threshold adjustment direction, select a fixed threshold adjustment range or a variable threshold adjustment range bound to the current energy threshold to adjust the threshold.
[0033] Furthermore, the auxiliary adjustment of the post-stage recognition neural network includes:
[0034] Determine the direction of threshold adjustment: Frame signals that are considered active frames by SAD will be sent to the subsequent recognition neural network. After accumulating active frames that meet the minimum inference frame length, the subsequent recognition neural network will perform an inference. The inference result will be fed back to SAD to determine whether the frames sent to the subsequent stage before are active frames. If the inference feedback of the subsequent neural network is: SAD classification is correct, SAD will lower the threshold to ensure that subsequent signals are more easily sent to the subsequent neural network and avoid the loss of valid signals; if the feedback is: SAD classification is wrong, SAD will increase the threshold to ensure that subsequent invalid signals will not be sent to the subsequent neural network again, so as to reduce the power consumption of false inference;
[0035] Threshold adjustment: After determining the threshold adjustment direction, SAD will immediately correct the energy threshold. The threshold adjustment amplitude can be fixed or non-fixed.
[0036] Furthermore, the short-time energy integral extraction module is divided into two branches to extract features of the frame signal. The left branch includes an absolute value sum module, which calculates the absolute value of all sampling values of a frame signal and adds them up as the short-time energy integral of the frame signal; the right branch squares all sampling values of a frame signal and adds them up as the short-time energy integral of the frame signal. The right branch includes a zero-jumping multiplication module and a dynamic quantization module. The zero-jumping multiplication module regards the square of the sampling point with a very small value as zero. The dynamic quantization module decides whether to perform rounding calculations based on the size of the shifted data to reduce the loss generated during the quantization process.
[0037] The beneficial technical effects of the present invention are as follows:
[0038] (1) The present invention proposes an adaptive sensing one-dimensional discrete-time signal activity detection method, which adopts adaptive following energy threshold adjustment and integrates the inference feedback of the post-stage recognition neural network. The energy threshold is adjusted according to the inference feedback. Through the adaptive following energy threshold adjustment and the post-stage recognition neural network assisted energy threshold adjustment, the final energy threshold will follow the signal amplitude, achieving the purpose of adaptive environment, making the accuracy of signal activity detection much higher than the traditional fixed threshold based on short-time energy integration, and reducing the power consumption of the device;
[0039] (2) The present invention also proposes a low-power and low-cost signal activity detection system based on a configurable architecture. By adopting ping-pong operation, zero-skipping multiplication and dynamic quantization, circuit multiplexing and integration of approximate operations are achieved to reduce circuit power consumption, thereby realizing a multi-mode configurable signal activity detection system with relatively low additional hardware overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 This is a flow chart of a method for detecting activity of a one-dimensional discrete-time signal using adaptive sensing provided by an embodiment of the present invention;
[0042] Figure 2 This is a structural diagram of an adaptive sensing one-dimensional discrete-time signal activity detection system provided by an embodiment of the present invention;
[0043] Figure 3 It is a structural diagram of the short-time energy integration extraction module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0045] Signal Activity Detection (SAD) is a process designed to detect activity from a signal stream. SAD has a wide range of applications, covering various fields from communication systems to sensor networks. Its significance and advantages are also obvious. For example, SAD can be used to distinguish between speech and non-speech parts in the speech signal processing process, thereby reducing the system's storage overhead and power consumption when there is no speech.
[0046] The features commonly used to implement SAD include time domain features, frequency domain features, time-frequency joint features, statistical features and machine learning features. The present invention selects time domain short-time energy, which is easy to extract and easy to build in systems with high real-time requirements, as the detection feature used to distinguish active and inactive parts in the signal stream.
[0047] An adaptive sensing method for one-dimensional discrete-time signal activity detection, such as Figure 1 As shown, the method includes:
[0048] Step S1: Pre-emphasis. The spectrum of a natural speech signal typically exhibits a rapid attenuation of high-frequency components, i.e., low-frequency components dominate the energy. This phenomenon is known as spectral attenuation. Pre-emphasis can compensate for this attenuation by enhancing high-frequency components, thereby making the high-frequency portion of the signal (the rapidly changing portion) more prominent. In speech signals, pre-emphasis increases the energy of the edge portions of the speech signal (such as consonants and plosives), making them easier to detect. The essence of pre-emphasis is a high-pass filter:
[0049] y[n]=x[n]-αx[n-1]
[0050] Where x[n] is the original signal at the current moment, x[n-1] is the original signal at the previous moment, y[n] is the signal after pre-emphasis at the current moment, and α is the pre-emphasis coefficient, which ranges from 0.9 to 1.0.
[0051] Pre-emphasis is optional and may not be used when resources are limited.
[0052] Step S2: Signal framing, dividing the pre-emphasized one-dimensional time signal stream into multiple short time frames of fixed length (usually between 10 milliseconds and 30 milliseconds). Because the characteristics of the one-dimensional time signal are relatively stable in a short time, subsequent processes are further processed on the basis of frames.
[0053] Step S3: Windowing the frame signal. After the signal is framed, it is equivalent to truncating each frame signal, resulting in discontinuity of the signal at the truncation point. This phenomenon is called the boundary effect. The influence of the frame boundary may cause inaccurate calculation of the short-term energy of the frame. Windowing reduces the influence of the two ends of the frame on the energy calculation by applying a smaller weight to the edge of the frame, thereby more accurately reflecting the energy of the central part of the signal. Common window functions include Hamming window, Hanning window, etc. This embodiment uses the Hamming window to window each frame signal. The mathematical expression of the Hamming window is:
[0054]
[0055] Where N is the number of sampling points contained in each frame. After windowing the signal frame x[n], the windowed signal x is obtained. ω [n]:
[0056] x ω [n]=x[n]·ω[n]
[0057] Step S4: Calculate the short-time energy integral of the single-frame signal. For each frame signal x ω [n], its short-time energy integral E is expressed as:
[0058]
[0059] Where N is the number of sampling points contained in each frame;
[0060] Frame signal windowing is also optional and may not be used when resources are limited.
[0061] Step S5: Dynamic energy threshold adjustment. The energy threshold adjustment includes two parallel sub-processes: one is adaptive following threshold adjustment, in which the energy threshold is always updated in the direction of following the signal amplitude; the other is auxiliary adjustment by the post-stage recognition neural network. Signal activity detection often does not work independently, but often serves as a pre-process of the post-stage recognition neural network. After the post-stage neural network recognizes an active frame or an inactive frame, it will update the threshold in the direction that is more likely to produce an active frame.
[0062] According to the traditional mechanism, the short-term energy integral of a frame is compared with a fixed empirical energy threshold to obtain the classification result of the frame, that is, active frame or inactive frame. However, in the mechanism we proposed, the energy threshold used to obtain the classification result is not fixed, but can be adaptively adjusted as the environment changes;
[0063] The adaptive following threshold adjustment includes:
[0064] Determining the initial threshold: Adaptive following determines whether and how to adjust the threshold based on the results of frame judgments over a period of time. This period is manually configured in frames and is called the adjustment window. At the beginning of the process, an initial threshold is manually determined based on experience. Different initial threshold settings affect the time it takes for the threshold to reach the environmental level.
[0065] Cumulative judgment result: When the number of frames does not reach the adjustment window length, the short-time energy integral of the frame is compared with the current energy threshold to determine whether the frame is an active frame or an inactive frame. At the same time, the number of active frames is continuously accumulated. When the short-time energy integral of the frame is greater than the current energy threshold, the frame is an active frame, and the active frame count is increased by 1.
[0066] Determine the direction of threshold adjustment: After the number of frames reaches the adjustment window length, the number of active frames in the accumulated number of frames of the past adjustment window length is used to determine whether the current energy threshold setting is appropriate. Two values N are set based on actual experience. less and N more , N less <N more , N less To lower the threshold, N more To increase the threshold, the typical value is N less =Adjustment window length*0.3, N more =Adjust window length*0.8, when the number of active frames is less than N less When the number of active frames is greater than N, it is considered that the number of active frames is too small, that is, the current energy threshold is too high, and the threshold needs to be lowered; when the number of active frames is greater than N more When the number of active frames is N, it is considered that there are too many active frames, that is, the current energy threshold is too low, and the threshold needs to be raised; when the number of active frames is N less and N more When the number of active frames is between , it is considered that the current energy threshold is appropriate and no adjustment is required;
[0067] Threshold adjustment: After determining the threshold adjustment direction, the threshold is adjusted. You can choose a fixed threshold adjustment range or a variable threshold adjustment range that is always bound to the current energy threshold. The fixed threshold adjustment range is more stable during adjustment than the variable threshold adjustment range. The variable threshold adjustment range can follow the environmental signal more quickly.
[0068] The auxiliary adjustment of the post-stage recognition neural network includes:
[0069] Determining the direction of threshold adjustment: In common application scenarios, frame signals that are considered active frames by the SAD will be sent to the subsequent recognition neural network. The subsequent recognition neural network does not perform inference at the frame level. It often performs inference after accumulating active frames that meet the minimum inference frame length. The inference result will be fed back to the SAD to determine whether the frames sent to the subsequent stage are active frames, thereby establishing a feedback mechanism to let the SAD know whether the current threshold is appropriate. If the inference feedback of the subsequent recognition neural network is: SAD classification is correct, SAD will lower the energy threshold to ensure that subsequent signals are more easily sent to the subsequent neural network and avoid the loss of valid signals. If the feedback is: SAD classification is incorrect, SAD will increase the energy threshold to ensure that subsequent invalid signals are not sent to the subsequent neural network again, thereby reducing the power consumption of incorrect inference.
[0070] Threshold adjustment: After determining the inference direction, SAD will immediately adjust the energy threshold. Unlike adaptive follow-up threshold adjustment, it will not wait for an adjustment window length before adjusting. The adjustment range can also be fixed or variable. The energy threshold adjustment range in this process is larger than that in adaptive follow-up threshold adjustment.
[0071] Step S6: Signal activity detection. After the above-mentioned threshold adjustment, the energy threshold will follow the signal amplitude to achieve the purpose of adaptive environment. The goal of SAD is to remove the inactive part and only obtain the active part. Therefore, a decision threshold is set. The decision threshold is directly related to the following energy threshold obtained in the above process. A larger threshold can be obtained by adding a constant to the above-mentioned following energy threshold as the decision threshold. Then, the decision threshold is used to perform signal activity detection to obtain the real-time classification result of the frame, that is, the short-time energy integral of the frame is compared with the decision threshold. If the short-time energy integral is greater than or equal to the decision threshold, it is classified as an active frame, and if the short-time energy integral is less than the decision threshold, it is classified as an inactive frame.
[0072] On the other hand, the present invention proposes an adaptive sensing one-dimensional discrete time signal activity detection system, such as Figure 2 As shown in the figure, the functions implemented by the signal activity detection system are: signal stream preprocessing, dividing the continuous signal stream into signal segments in frames, extracting the characteristics of the signal segments and obtaining the classification results of the signal segments based on these characteristics - whether they belong to active signal segments or inactive signal segments, and deciding whether the corresponding signal segments should be sent to the subsequent recognition neural network based on the classification results.
[0073] The signal activity detection system comprises a pre-emphasis module, a data buffer module, a windowing module, a short-time energy integral extraction module, a follow-up threshold updating module and a judgment module.
[0074] Pre-emphasis module: The pre-emphasis module compensates for the spectral attenuation of the signal by enhancing the high-frequency components, making the high-frequency part of the signal more prominent; the pre-emphasis processing method is: y[n] = x[n]-αx[n-1], where x[n] is the original signal at the current moment, x[n-1] is the original signal at the previous moment, y[n] is the signal after pre-emphasis at the current moment, and α is the pre-emphasis coefficient; in actual applications, the pre-emphasis module can be replaced by various types of filters, noise elimination, echo cancellation and other pre-processing. The main purpose is to increase the SAD accuracy and provide higher quality signals for different subsequent processing algorithms / hardware. During SAD work, the pre-emphasis module can be configured to switch different pre-processing methods or skip to adapt to the requirements of different subsequent processing algorithms / hardware.
[0075] Data cache module: The signal stream processed by the pre-emphasis module will be continuously fed into the data cache module. This module uses different memories to ping-pong store data of different half frames of the same channel, and stores data of different channels in different address segments of the same memory. The caching process will complete the framing operation at the same time.
[0076] Windowing module: The windowing module uses a window function to apply a smaller weight to the edge of the frame, reducing the influence of the two ends of the frame on the energy calculation, thereby more accurately reflecting the energy of the central signal. The window functions include Hamming window and Hanning window. The mathematical expression of the Hamming window is: n=0,1,...N-1, where N is the number of sampling points in each frame. After windowing the signal frame x[n], the windowed signal x is obtained. ω [n],x ω [n]=x[n]·ω[n];
[0077] Short-time energy integral extraction module: The short-time energy integral extraction module calculates the short-time energy integral of the frame signal after being framed by the data cache module. For each frame signal x ω [n], its short-time energy integral E is expressed as: Where N is the number of sampling points contained in each frame;
[0078] The short-time energy integral extraction module is divided into two branches to extract features of the frame signal, such as Figure 3As shown, the left branch includes an absolute value sum module, which calculates the absolute value of all sampling values of a frame signal and adds them up as the short-time energy integral of the frame signal. The right branch squares all sampling values of a frame signal and adds them up as the short-time energy integral of the frame signal. The right branch includes a zero-jumping multiplication module and a dynamic quantization module. The zero-jumping multiplication module regards the square of a sampling point with a very small value as zero; the dynamic quantization module decides whether to perform rounding calculations based on the size of the shifted data to minimize the loss in the quantization process. Because fixed-point numbers are operated in hardware, the position of the decimal point will change after the fixed-point numbers are multiplied, and the decimal point needs to be returned to its original position by shifting. This process is called quantization. The zero-jumping multiplication module and the dynamic quantization module are both technical means to reduce hardware computing overhead in the process of squaring the sampling value.
[0079] Following threshold update module: This module completes the adjustment of the adaptive following threshold and the adaptive energy threshold of the feedback adjustment of the reasoning result of the post-stage recognition neural network through the short-time energy integral provided by the short-time energy integral extraction module;
[0080] Adaptive follow threshold adjustment includes:
[0081] Determining the initial threshold: Adaptive following determines whether and how to adjust the threshold based on the results of frame judgments over a period of time. This period is manually configured in frames and is called the adjustment window. At the beginning of the process, an initial threshold is manually determined based on experience. Different initial threshold settings affect the time it takes for the threshold to reach the environmental level.
[0082] Cumulative judgment result: When the number of frames does not reach the adjustment window length, the short-time energy integral of the frame is compared with the current energy threshold to determine whether the frame is an active frame or an inactive frame. At the same time, the number of active frames is continuously accumulated. When the short-time energy integral of the frame is greater than the current energy threshold, the frame is an active frame, and the active frame count is increased by 1.
[0083] Determine the direction of threshold adjustment: After the number of frames reaches the adjustment window length, the number of active frames in the accumulated number of frames of the past adjustment window length is used to determine whether the current energy threshold setting is appropriate. Two values N are set based on actual experience. less and N more , N less <N more , N less To lower the threshold, N more To increase the threshold, the typical value is N less =Adjustment window length*0.3, N more =Adjust window length*0.8, when the number of active frames is less than N lessWhen the number of active frames is greater than N, it is considered that the number of active frames is too small, that is, the current energy threshold is too high, and the threshold needs to be lowered; when the number of active frames is greater than N more When the number of active frames is N, it is considered that there are too many active frames, that is, the current energy threshold is too low, and the threshold needs to be raised; when the number of active frames is N less and N more When the number of active frames is between , it is considered that the current energy threshold is appropriate and no adjustment is required;
[0084] Threshold adjustment: After determining the threshold adjustment direction, the threshold is adjusted. You can choose a fixed threshold adjustment range or a variable threshold adjustment range that is always bound to the current energy threshold. The fixed threshold adjustment range is more stable during adjustment than the variable threshold adjustment range. The variable threshold adjustment range can follow the environmental signal more quickly.
[0085] The auxiliary adjustment of the post-stage recognition neural network includes:
[0086] Determining the direction of threshold adjustment: In common application scenarios, frame signals that are considered active frames by the SAD will be sent to the subsequent recognition neural network. The subsequent recognition neural network does not perform inference at the frame level. It often performs inference after accumulating active frames that meet the minimum inference frame length. The inference result will be fed back to the SAD to determine whether the frames sent to the subsequent stage are active frames, thereby establishing a feedback mechanism to let the SAD know whether the current threshold is appropriate. If the inference feedback of the subsequent recognition neural network is: SAD classification is correct, SAD will lower the energy threshold to ensure that subsequent signals are more easily sent to the subsequent neural network and avoid the loss of valid signals. If the feedback is: SAD classification is incorrect, SAD will increase the energy threshold to ensure that subsequent invalid signals are not sent to the subsequent neural network again, thereby reducing the power consumption of incorrect inference.
[0087] Threshold adjustment: After determining the inference direction, SAD will immediately adjust the energy threshold. Unlike adaptive follow-up threshold adjustment, it will not wait for an adjustment window length before adjusting. The adjustment range can also be fixed or variable. The energy threshold adjustment range in this process is larger than that in adaptive follow-up threshold adjustment.
[0088] Decision module: After the above threshold adjustment, the energy threshold will follow the signal amplitude to achieve the purpose of adaptive environment, and the goal of SAD is to remove the inactive part and only obtain the active part. Therefore, a decision threshold is set according to the following energy threshold provided by the following threshold update module. The decision threshold is directly related to the following energy threshold obtained in the above process. A larger threshold can be obtained by adding a constant to the above following energy threshold as the decision threshold, and then the decision threshold is used to perform signal activity detection to obtain the real-time classification result of the frame, that is, the short-time energy integral of the frame is compared with the decision threshold. The short-time energy integral is greater than or equal to the decision threshold and is classified as an active frame. The short-time energy integral is less than the decision threshold and is classified as an inactive frame. The classification result will be used to control whether the frame data is output or not.
[0089] The above-mentioned adaptive perception one-dimensional discrete-time signal activity detection system was designed with multi-mode application scenarios in mind. Therefore, in order to meet diverse processing requirements, most of the functional modules of the circuit are configurable. By configuring the circuit modules during initialization, the required functional modules can be selected to achieve different signal processing. The system supports one-dimensional discrete-time signals with a maximum number of channels equivalent to the number of memory channels (allowing switching between single-channel and multi-channel modes during processing). Each additional channel requires expanding the size of half a frame on each memory. Considering that the energy values of multiple channels are not much different, the channel data is separated in dual-channel and multi-channel modes, and only the signal of one of the channels is framed and calculated. The system also supports configuring whether there is overlap between frames. The overlapping mode allows the module to better extract the connection information between frames.
[0090] To further reduce computational overhead while streamlining the architecture, the signal activity detection system employs the following two approaches:
[0091] (1) Near-zero Ignore: Considering that in some scenarios, there are mostly no active signals, a zero-skipping mechanism is added during the calculation process to filter out near-zero data in the sampled signal that is less than the preset threshold, further reducing power consumption and operations while ensuring accuracy;
[0092] (2) Dynamic quantization: After performing square sum operations on the data, in order to maintain the consistency of data accuracy, quantization control is performed through external configuration to achieve dynamic quantization processing.
[0093] Through the solution of the present invention, taking the hit rate of the binary classification index as the evaluation standard, in a random signal-to-noise ratio environment of -4dB to 20dB, the dynamically adjusted energy threshold of the present invention has a hit rate increase of nearly 30% compared with the traditional fixed threshold SAD, thereby improving the accuracy of signal activity detection and reducing cost and power consumption.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting activity of one-dimensional discrete-time signals with adaptive sensing, characterized in that: The method comprises: Step S1: Pre-emphasis. Pre-emphasis compensates for the spectrum attenuation of a one-dimensional time signal by enhancing high-frequency components, thereby making the high-frequency portion of the signal more prominent. The pre-emphasis processing method is: y[n] = x[n] - αx[n-1], where x[n] is the original signal at the current moment, x[n-1] is the original signal at the previous moment, y[n] is the signal after pre-emphasis at the current moment, and α is the pre-emphasis coefficient. Step S2: signal framing, dividing the one-dimensional time signal stream after pre-emphasis processing into multiple short time frames of fixed length; Step S3: Windowing the frame signal. Windowing reduces the influence of both ends of the frame on energy calculation by applying a weight to the edge of the frame. The window functions used include Hamming window and Hanning window. The expression of Hamming window is: Where N is the number of sampling points contained in each frame. After windowing the signal frame x[n], the windowed signal x is obtained. ω [n],x ω [n]=x[n]·ω[n]; Step S4: Calculate the short-time energy integral of the single-frame signal. For each frame signal x ω [n], its short-time energy integral E is expressed as: Where N is the number of sampling points contained in each frame; Step S5: Dynamic energy threshold adjustment. The energy threshold adjustment includes two parallel sub-processes: one is adaptive following threshold adjustment, in which the energy threshold is always updated in the direction of following the signal amplitude; the other is post-stage recognition neural network assisted adjustment. After the post-stage neural network recognizes an active frame or an inactive frame, the threshold is updated in the direction that is more likely to produce an active frame. Step S6: Signal activity detection. After the above-mentioned energy threshold adjustment, a decision threshold is set. A larger threshold is obtained by adding a constant to the energy threshold obtained after the above-mentioned threshold adjustment as the decision threshold. Then, the decision threshold is used to perform signal activity detection to obtain the real-time classification result of the frame, that is, the short-time energy integral of the frame is compared with the decision threshold. A frame with a short-time energy integral greater than or equal to the decision threshold is classified as an active frame, and a frame with a short-time energy integral less than the decision threshold is classified as an inactive frame.
2. The method according to claim 1, characterized in that In step S5, the adaptive following threshold adjustment includes: Determining the initial threshold: Adaptive following determines whether and how to adjust the threshold based on the results of frame judgments over a period of time. This period, measured in frames, is called the adjustment window. At the beginning of the process, an initial threshold is manually set based on experience. Cumulative judgment result: When the number of frames does not reach the adjustment window length, the short-time energy integral of the frame is compared with the current energy threshold to determine whether the frame is an active frame or an inactive frame. At the same time, the number of active frames is continuously accumulated. When the short-time energy integral of the frame is greater than the current energy threshold, the frame is an active frame, and the active frame count is increased by 1. Determine the direction of threshold adjustment: After the number of frames reaches the adjustment window length, the number of active frames in the accumulated number of frames of the past adjustment window length is used to determine whether the current energy threshold setting is appropriate. Two values N are set based on actual experience. less and N more , N less <N more , N less To lower the threshold, N more To increase the threshold, when the number of active frames is less than N less When the number of active frames is greater than N, it means that the number of active frames is too small, that is, the current energy threshold is too high and the threshold needs to be lowered. more When the number of active frames is N, it means that there are too many active frames, that is, the current energy threshold is too low and the threshold needs to be raised. less and N more When it is between , it means the number of active frames is appropriate, that is, the current energy threshold is appropriate and does not need to be adjusted; Threshold adjustment: After determining the threshold adjustment direction, select a fixed threshold adjustment range or a variable threshold adjustment range bound to the current energy threshold to adjust the threshold.
3. The method according to claim 1, characterized in that In step S5, the post-stage recognition neural network auxiliary adjustment includes: Determine the direction of threshold adjustment: Frame signals that are considered active frames by SAD will be sent to the subsequent recognition neural network. After accumulating active frames that meet the minimum inference frame length, the subsequent recognition neural network will perform an inference. The inference result will be fed back to SAD to determine whether the frames sent to the subsequent stage before are active frames. If the inference feedback of the subsequent neural network is: SAD classification is correct, SAD will lower the threshold to ensure that subsequent signals are more easily sent to the subsequent neural network and avoid the loss of valid signals; if the feedback is: SAD classification is wrong, SAD will increase the threshold to ensure that subsequent invalid signals will not be sent to the subsequent neural network again, so as to reduce the power consumption of false inference; Threshold adjustment: After determining the threshold adjustment direction, SAD will immediately correct the energy threshold. The threshold adjustment amplitude can be fixed or non-fixed.
4. An adaptive sensing one-dimensional discrete time signal activity detection system, characterized in that: The system includes a pre-emphasis module, a data buffer module, a windowing module, a short-time energy integral extraction module, a follow-up threshold update module and a judgment module; Pre-emphasis module: The pre-emphasis module compensates for the spectrum attenuation of the one-dimensional time signal by enhancing the high-frequency components, making the high-frequency part of the signal more prominent. The pre-emphasis processing method is: y[n] = x[n] - αx[n-1], where x[n] is the original signal at the current moment, x[n-1] is the original signal at the previous moment, y[n] is the signal after pre-emphasis at the current moment, and α is the pre-emphasis coefficient; Data cache module: The signal stream processed by the pre-emphasis module is continuously fed into the data cache module. This module uses different memories to ping-pong store data of different half frames of the same channel, and stores data of different channels in different address segments of the same memory. This caching process will also complete the signal framing operation; Windowing module: The windowing module uses a window function to apply a weight to the edge of the frame, so that the influence of the two ends of the frame on the energy calculation is reduced. The window function includes Hamming window and Hanning window. The expression of Hamming window is: Where N is the number of sampling points contained in each frame. After windowing the signal frame x[n], the windowed signal x is obtained. ω [n],x ω [n]=x[n]·ω[n]; Short-time energy integral extraction module: The short-time energy integral extraction module calculates the short-time energy integral of the frame signal after the frame is divided by the data buffer module. For each frame signal x ω [n], its short-time energy integral E is expressed as: Where N is the number of sampling points contained in each frame; Following threshold update module: This module completes the adaptive following threshold adjustment and the adaptive threshold adjustment assisted by the inference result feedback of the post-stage recognition neural network through the short-term energy integration provided by the energy integration extraction module; Decision module: A decision threshold is set according to the following energy threshold provided by the following threshold update module. A larger threshold is obtained by adding a constant to the above-mentioned following energy threshold as the decision threshold. The decision threshold is then used to detect signal activity and obtain the real-time classification result of the frame. That is, the short-time energy integral of the frame is compared with the decision threshold. If the short-time energy integral is greater than or equal to the decision threshold, it is classified as an active frame. If the short-time energy integral is less than the decision threshold, it is classified as an inactive frame. The classification result will be used to control whether the frame data is output or not.
5. The system according to claim 4, characterized in that Adaptive follow threshold adjustment includes: Determining the initial threshold: Adaptive following determines whether and how to adjust the threshold based on the results of frame judgments over a period of time. This period, measured in frames, is called the adjustment window. At the beginning of the process, an initial threshold is manually set based on experience. Cumulative judgment result: When the number of frames does not reach the adjustment window length, the short-time energy integral of the frame is compared with the current energy threshold to determine whether the frame is an active frame or an inactive frame. At the same time, the number of active frames is continuously accumulated. When the short-time energy integral of the frame is greater than the current energy threshold, the frame is an active frame, and the active frame count is increased by 1. Determine the direction of threshold adjustment: After the number of frames reaches the adjustment window length, the number of active frames in the accumulated number of frames of the past adjustment window length is used to determine whether the current energy threshold setting is appropriate. Two values N are set based on actual experience. less and N more , N less <N more , N less To lower the threshold, N more To increase the threshold, when the number of active frames is less than N less When the number of active frames is greater than N, it means that the number of active frames is too small, that is, the current energy threshold is too high and the threshold needs to be lowered. more When the number of active frames is N, it means that there are too many active frames, that is, the current energy threshold is too low and the threshold needs to be raised. less and N more When it is between , it means the number of active frames is appropriate, that is, the current energy threshold is appropriate and does not need to be adjusted; Threshold adjustment: After determining the threshold adjustment direction, select a fixed threshold adjustment range or a variable threshold adjustment range bound to the current energy threshold to adjust the threshold.
6. The system according to claim 4, characterized in that The auxiliary adjustment of the post-stage recognition neural network includes: Determine the direction of threshold adjustment: Frame signals that are considered active frames by SAD will be sent to the subsequent recognition neural network. After accumulating active frames that meet the minimum inference frame length, the subsequent recognition neural network will perform an inference. The inference result will be fed back to SAD to determine whether the frames sent to the subsequent stage before are active frames. If the inference feedback of the subsequent neural network is: SAD classification is correct, SAD will lower the threshold to ensure that subsequent signals are more easily sent to the subsequent neural network and avoid the loss of valid signals; if the feedback is: SAD classification is wrong, SAD will increase the threshold to ensure that subsequent invalid signals will not be sent to the subsequent neural network again, so as to reduce the power consumption of false inference; Threshold adjustment: After determining the threshold adjustment direction, SAD will immediately correct the energy threshold. The threshold adjustment amplitude can be fixed or non-fixed.
7. The system according to claim 4, wherein: The short-time energy integral extraction module is divided into two branches to extract features from the frame signal. The left branch includes an absolute value sum module, which calculates the absolute value of all sampling values of a frame signal and adds them up as the short-time energy integral of the frame signal; the right branch squares all sampling values of a frame signal and adds them up as the short-time energy integral of the frame signal. The right branch includes a zero-jumping multiplication module and a dynamic quantization module. The zero-jumping multiplication module regards the square of the sampling point with a very small value as zero. The dynamic quantization module decides whether to perform rounding calculations based on the size of the shifted data to reduce the loss generated during the quantization process.
Citation Information
Patent Citations
Voice activity detection method and device and voice recognition method and device
CN108346425A
Speech activity detection equipment and method
CN110556131A