Processing device and voice activity detection system and method
By configuring the feature extraction module and the voice activity detection module in the processing device, the digital signal processor is activated only when necessary during the voice activity detection process, which solves the problem of power consumption waste in the prior art and realizes more efficient power consumption management.
Patent Information
- Application Number
- CN202510330501.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, voice activity detection requires that the digital signal processor always remain powered on, resulting in waste of power consumption.
A processing device is designed, including a feature extraction module and a voice activity detection module. The voice activity detection module detects the received audio signal in a low power consumption mode and controls the digital signal processor to power on for processing when voice activity is detected.
By configuring the feature extraction module and the voice activity detection module in the processing device, the digital signal processor is powered on only when voice activity is detected, thereby maintaining the power off state at other times, significantly saving power consumption.
Smart Images

Figure CN120108433A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a processing device and a voice activity detection system and method. Background Art
[0002] Voice Activity Detection (VAD) is mainly used to distinguish speech segments from non-speech segments (such as background noise, silence, etc.) from audio signals. By accurately detecting the presence or absence of voice activity, it can help optimize resource usage, improve processing efficiency, and improve user experience.
[0003] In the related art, voice activity detection requires a digital signal processor (DSP) to be always powered on, resulting in a waste of power consumption. Summary of the invention
[0004] In view of this, the present application provides a processing device and a voice activity detection system and method to solve related technical problems.
[0005] In a first aspect, the present application provides a processing device, including a feature extraction module and a voice activity detection module,
[0006] The voice activity detection module is connected to the feature extraction module;
[0007] The feature extraction module is used to continuously receive an audio signal, extract feature information from the audio signal, and send the feature information and the audio signal to the voice activity detection module;
[0008] The voice activity detection module is used to perform voice activity detection on the audio signal based on the feature information, and send the audio signal to the digital signal processor for processing when voice activity is detected in the audio signal.
[0009] Optionally, the processing device operates when the power consumption is less than a predetermined power consumption threshold.
[0010] Optionally, the voice activity detection module is specifically configured to mark a voice segment in the audio signal with a voice activity detection identifier based on a voice activity detection result.
[0011] Optionally, the voice activity detection module is further configured to determine, when the audio signal contains the voice activity detection identifier, whether there is voice activity in the audio signal, and control the digital signal processor to power on and send the audio signal to the digital signal processor.
[0012] In a second aspect, the present application provides a voice activity detection system, comprising: a digital signal processor and a processing device according to any one of claims 1 to 4 connected to each other;
[0013] The processing device is used to continuously receive an audio signal, extract feature information from the audio signal, perform voice activity detection on the audio signal based on the feature information, and send the audio signal to a digital signal processor for processing when voice activity is detected in the audio signal.
[0014] Optionally, it also includes: an application processor connected to the digital signal processor; the digital signal processor is used to detect the operation instruction contained in the audio signal when recognizing that there is a target wake-up word in the audio signal, and send the operation instruction to the application processor or execute the operation instruction;
[0015] The application processor is configured to execute the operation instruction upon receiving the operation instruction.
[0016] Optionally, the digital signal processor is further used to recognize a wake-up word in the audio signal, and remain powered on when it is recognized that there is a target wake-up word in the audio signal;
[0017] The digital signal processor is further configured to detect the audio signal in a power-on state, and when the operation instruction is detected in the audio signal, control the application processor to power on and send the operation instruction to the application processor or execute the operation instruction.
[0018] Optionally, the digital signal processor is further used to power off when it is identified that the audio signal does not contain the target wake-up word, and / or power off when no operation instructions are detected in the audio signal within a predetermined time period.
[0019] Optionally, the digital signal processor and the application processor each include at least one voice processing module;
[0020] The at least one voice processing module is used to execute the operation instruction when receiving the operation instruction.
[0021] In a third aspect, the present application provides a voice activity detection method, which is applied to the processing device described in the first aspect, comprising:
[0022] Continuously receiving an audio signal and extracting feature information from the audio signal;
[0023] Voice activity detection is performed on the audio signal based on the feature information, and when voice activity is detected in the audio signal, the audio signal is sent to a digital signal processor for processing.
[0024] Optionally, after performing voice activity detection on the audio signal based on the feature information, the method further includes:
[0025] Based on the voice activity detection result, the voice segment in the audio signal is marked with a voice activity detection identifier to obtain an audio signal containing the voice activity detection identifier.
[0026] Optionally, the processing the audio signal when voice activity is detected in the audio signal includes:
[0027] If the audio signal includes the voice activity detection flag, it is determined that there is voice activity in the audio signal, and the digital signal processor is controlled to power on and the audio signal is sent to the digital signal processor.
[0028] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the voice activity detection method as described in the third aspect.
[0029] In a fifth aspect, the present application provides an electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the voice activity detection method as described in the second aspect when executing the computer program.
[0030] In a sixth aspect, the present application provides a chip comprising one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send the signal to the processor, wherein the signal includes a computer instruction stored in the memory; when the processor executes the computer instruction, the electronic device executes the voice activity detection method as described in the second aspect.
[0031] By means of the above technical solution, the present application provides a processing device and a voice activity detection system and method. Compared with the current related technologies, the present application configures a feature extraction module and a voice activity detection module in a processing device, and connects the voice activity detection module with the feature extraction module and a digital signal processor, so that the voice activity detection module performs voice activity detection on the received audio signal in a low power consumption mode, and controls the digital signal processor to power on when voice activity is detected in the audio signal, so as to process the audio signal through the digital signal processor; since the voice activity detection module and the feature extraction module need to be kept in a normally open state to detect the voice signal, the present application can configure the voice activity detection module and the feature extraction module in the processing device, so that the voice activity detection module and the feature extraction module can be detected in a low power consumption mode, and control the digital signal processor to power on when voice activity is detected in the audio signal, so that the audio signal can be processed by the digital signal processor, so that during the detection process of the voice activity detection module and the feature extraction module, the digital signal processor can be kept in a power-off state, and only powered on when voice activity is detected, which can save power consumption to a large extent and avoid power consumption waste.
[0032] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0035] Figure 1 A schematic diagram of the structure of a processing device provided in an embodiment of the present application is shown;
[0036] Figure 2 A schematic diagram of the structure of a voice activity detection system provided in an embodiment of the present application is shown;
[0037] Figure 3 A schematic diagram showing an example provided by an embodiment of the present application;
[0038] Figure 4 A schematic diagram showing an example provided by an embodiment of the present application;
[0039] Figure 5 A flow chart of a voice activity detection method provided in an embodiment of the present application is shown;
[0040] Figure 1 middle:
[0041] 1-processing device, 11-feature extraction module, 12-speech activity detection module;
[0042] Figure 2 middle:
[0043] 2- Digital signal processor;
[0044] 3- Application processor. DETAILED DESCRIPTION
[0045] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present application.
[0046] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0047] In this application, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0048] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.
[0049] Combine the following Figure 1 A processing device according to some embodiments of the present application is described.
[0050] The present application provides a processing device 1, such as Figure 1 As shown, it includes: a feature extraction module 11 and a voice activity detection module 12; the voice activity detection module 12 is connected to the feature extraction module 11; the feature extraction module 11 is used to continuously receive the audio signal, and extract feature information in the audio signal, and send the feature information and the audio signal to the voice activity detection module 12; the voice activity detection module 12 is used to perform voice activity detection on the audio signal based on the feature information, and send the audio signal to the digital signal processor for processing when voice activity is detected in the audio signal.
[0051] Optionally, the processing device 1 operates when the power consumption is less than a predetermined power consumption threshold.
[0052] In the embodiment of the present application, the processing device 1 may be a low-power always-on processor (AON Processor) or a low-power always-on processing module, which is a microprocessor or processing unit that is used to keep the device active most of the time and consume very little power. Such a processor is usually integrated into devices that require long standby time and have instant response capabilities, such as smart watches, smart home devices, and in-vehicle smart cockpit systems.
[0053] In some examples, the processing device 1 can remain online all the time, even if the main system is in sleep or low-power mode, the processing device 1 remains active, able to monitor specific events or trigger conditions in real time, and wake up the main processor when necessary. The design uses low-power components and optimized power management strategies to ensure that the overall battery life of the device is not significantly affected even if it runs for a long time. It can quickly process and respond to specific types of input (for example, specific voice commands), and then wake up the main processor to complete more complex tasks, thereby achieving instant response.
[0054] For this embodiment, in speech processing, the main goal of the feature extraction module 11 is to extract features that can effectively represent the speech content from the audio signal for further processing, such as voice activity detection (VAD), automatic speech recognition (ASR) and speaker recognition.
[0055] As an optional method, the Voice Activity Detection module 12 (VAD) is a key component in the speech processing system, which is used to distinguish between speech segments and non-speech segments (such as background noise, silence, etc.) in the audio stream. The VAD module is very important in many application scenarios, such as automatic speech recognition (ASR), voice wake-up system, telephone call enhancement, etc. It determines whether the current signal contains user voice by analyzing the input audio signal and extracting features.
[0056] It should be noted that the feature extraction module 11 and the voice activity detection module 12 in the embodiment of the present application are configured in the processing device 1. Specifically, since the voice activity detection module 12 needs to be always on, in the related art, the voice activity detection module 12 is mounted on a digital signal processor for operation, which causes the digital signal processor to also be always on. However, the total computing power of a digital signal processor is usually redundant compared to the computing power requirements of the voice activity detection algorithm. This means that most of the computing power is idle and most of the power consumption is wasted; therefore, in the embodiment of the present application, configuring the feature extraction module 11 and the voice activity detection module 12 in the processing device 1 can keep the feature extraction module 11 and the voice activity detection module 12 in a low-power always-on state, thereby saving power consumption while ensuring performance.
[0057] In an embodiment of the present application, the audio signal received by the feature extraction module 11 may be an audio signal generated based on the audio in the environment; the feature information may include but is not limited to: short-time energy (Short-Time Energy), zero crossing rate (Zero Crossing Rate, ZCR), Mel-Frequency Cepstral Coefficients (Mel-Frequency Cepstral Coefficients, MFCCs), zero crossing rate (Zero Crossing Rate, ZCR), linear predictive coding coefficient (Linear Predictive Coding, LPC), fundamental frequency (Fundamental Frequency, F0), spectral centroid (Spectral Centroid), etc.
[0058] In some examples, the extraction process may include preprocessing, feature calculation, and feature normalization. Specifically, preprocessing is to divide the continuous audio stream into short time segments (frames), each frame is usually 20-40 milliseconds long, and has 50%-75% overlap, and a window function (such as a Hamming window) is applied to each frame to reduce spectrum leakage and improve the accuracy of spectrum analysis. Feature calculation is to select a suitable feature extraction algorithm according to the application scenario and calculate the feature vector of each frame. For example, when calculating MFCC features, steps such as fast Fourier transform (FFT), Mel filter bank, logarithmic compression, and discrete cosine transform (DCT) are required. Furthermore, in order to eliminate the impact of different recording environments, the extracted features are usually normalized, such as mean variance normalization (MVN).
[0059] In some examples, the voice activity detection module 12 can perform voice activity detection on the received audio signal in the low power consumption mode, and control the digital signal processor to power on when voice activity is detected in the audio signal, so as to process the audio signal through the digital signal processor. In this way, the digital signal processor can remain powered off even when no voice activity is detected, thereby saving system power consumption.
[0060] For this embodiment, a digital signal processor (DSP) is a microprocessor specially designed to perform digital signal processing tasks. Compared with general-purpose microprocessors, DSP chips are optimized for high-speed numerical calculations and are particularly suitable for application scenarios that require rapid processing of large amounts of data. DSPs are widely used in audio, video, communications, image processing and other fields.
[0061] Compared with the current related art, this embodiment configures the feature extraction module and the voice activity detection module in the processing device, and connects the voice activity detection module with the feature extraction module and the digital signal processor, so that the voice activity detection module performs voice activity detection on the received audio signal in the low power consumption mode, and controls the digital signal processor to power on when the voice activity is detected in the audio signal, so as to process the audio signal through the digital signal processor; since the voice activity detection module and the feature extraction module need to be kept in a normally open state to detect the voice signal, this embodiment can configure the voice activity detection module and the feature extraction module in the processing device, so that the voice activity detection module and the feature extraction module can be detected in the low power consumption mode, and the digital signal processor is controlled to power on when the voice activity is detected in the audio signal, and the audio signal can be processed by the digital signal processor, so that during the detection process of the voice activity detection module and the feature extraction module, the digital signal processor can be kept in a power-off state, and is only powered on when the voice activity is detected, which can save power consumption to a large extent and avoid power consumption waste.
[0062] Optionally, the voice activity detection module 12 is further configured to mark a voice segment in the audio signal with a voice activity detection identifier based on the voice activity detection result.
[0063] In the embodiment of the present application, the speech segment may be a portion containing human speech sounds, with obvious intonation, rhythm and spectral characteristics; correspondingly, the non-speech segment may include background noise, silence or other non-speech sounds (such as wind noise, traffic noise, machine running sound, etc.).
[0064] In some examples, the voice activity detection identifier can be an identifier corresponding to a speech segment, and the non-speech segment can also be marked with a corresponding identifier. For example, the voice activity detection identifier can be 1, and the identifier corresponding to the non-speech segment can be 0, or the voice activity detection identifier can be 0, and the identifier corresponding to the non-speech segment can be 1, and so on. The voice activity detection identifier can be set according to needs and is not limited here.
[0065] Optionally, the voice activity detection module 12 is further configured to determine whether there is voice activity in the audio signal when the audio signal includes a voice activity detection flag, and control the digital signal processor to power on and send the audio signal to the digital signal processor.
[0066] For this embodiment, if there is a voice activity detection flag in the audio signal, it means there is voice activity in the audio signal, and it is necessary to control the digital signal processor to power on and send the audio signal to the digital signal processor.
[0067] Compared with the current related art, this embodiment configures the feature extraction module and the voice activity detection module in the processing device, and connects the voice activity detection module with the feature extraction module and the digital signal processor, so that the voice activity detection module performs voice activity detection on the received audio signal in the low power consumption mode, and controls the digital signal processor to power on when the voice activity is detected in the audio signal, so as to process the audio signal through the digital signal processor; since the voice activity detection module and the feature extraction module need to be kept in a normally open state to detect the voice signal, this embodiment can configure the voice activity detection module and the feature extraction module in the processing device, so that the voice activity detection module and the feature extraction module can be detected in the low power consumption mode, and the digital signal processor is controlled to power on when the voice activity is detected in the audio signal, and the audio signal can be processed by the digital signal processor, so that during the detection process of the voice activity detection module and the feature extraction module, the digital signal processor can be kept in a power-off state, and is only powered on when the voice activity is detected, which can save power consumption to a large extent and avoid power consumption waste.
[0068] This embodiment also provides a voice activity detection system, such as Figure 2 As shown, the system includes:
[0069] A voice activity detection system provided by the present application is characterized in that it includes: a digital signal processor 2 and the processing device 1 in the above embodiment connected to each other; the processing device 1 is used to continuously receive an audio signal, extract feature information in the audio signal, perform voice activity detection on the audio signal based on the feature information, and send the audio signal to the digital signal processor 2 for processing when voice activity is detected in the audio signal.
[0070] In an embodiment of the present application, the processing device 1 can remain online all the time. Even if the main system is in sleep or low power mode, the processing device 1 remains active, can monitor specific events or trigger conditions in real time, and wake up the main processor, that is, the digital processor in the embodiment of the present application, when necessary.
[0071] In some examples, such as Figure 3 As shown in (b), the embodiment of the present application performs voice activity detection on the received audio signal through the processing device 1, that is, the processing device needs to remain in a normally open state, and control the digital signal processor to power on when voice activity is detected in the audio signal. Compared with the related art, Figure 3 In the case shown in (a), voice activity detection needs to be performed by a digital signal processor, so that the digital signal processor is always turned on. The embodiment of the present application can save power consumption to a greater extent.
[0072] Optionally, the digital signal processor 2 is used to detect the operation instructions contained in the audio signal when the target wake-up word exists in the identified audio signal, and send the operation instructions to the application processor 3 or execute the operation instructions; the application processor 3 is used to execute the operation instructions when the operation instructions are received.
[0073] In the embodiments of the present application, the application processor (AP) generally integrates multiple functional modules such as a central processing unit (CPU), a graphics processing unit (GPU), a memory management unit (MMU), a multimedia codec, etc., and supports multiple external interfaces.
[0074] In some examples, wake-up word detection is a key function in the voice interaction system, allowing the device to continuously listen for a specific wake-up word in low-power mode, and activate a more powerful processing unit to execute subsequent commands when the wake-up word is recognized. The target wake-up word can be a wake-up word set by the user according to needs. For example, the target wake-up word can be "AA", "BBB", etc. The specific content of the target wake-up word is not limited in the embodiments of this application.
[0075] For example, if the user is driving a car equipped with an advanced smart cockpit system, when the user says: "Hello", even in a noisy urban traffic environment, the processing device 1 in the car can accurately extract features from the background noise and identify valid voice activities through the VAD module. Once the wake-up word is detected, the digital signal processor 2 activates the main processor to prepare to receive and process the user's voice command. Next, the user can directly give an instruction: "Navigate to the nearest gas station." At this time, the processing device 1 continues to monitor the audio stream to extract features, and passes these features to the automatic speech recognition (ASR) module through the digital signal processor 2 to parse the specific commands and perform corresponding operations. If there is an abnormal sound in the car (such as a crying child), the processing device 1 can analyze the sound features through the digital signal processor 2 and its built-in algorithm, and take action according to pre-set conditions, such as sending a notification to remind the user to pay attention.
[0076] Optionally, the digital signal processor 2 is also used to recognize wake-up words in the audio signal and remain in the powered-on state when the target wake-up word is recognized in the audio signal; the digital signal processor 2 is also used to detect the audio signal in the powered-on state, and when an operation instruction is detected in the audio signal, control the application processor 3 to power on and send the operation instruction to the application processor 3 or execute the operation instruction.
[0077] For example, if the target wake-up words include A, B and C, if any one of the wake-up words A, B and C is detected during the detection of the current audio signal, it can be determined that the target wake-up word exists in the audio signal, and the digital signal processor 2 can be controlled to remain powered on; further, the digital signal processor 2 also needs to detect the operation instructions in the audio signal, and when the actual operation instructions are detected, the operation instructions are completed through the digital signal processor 2 or the application processor 3.
[0078] Optionally, the digital signal processor is also used to power off when it is recognized that the audio signal does not contain the target wake-up word, and / or power off when no operation instructions are detected in the audio signal within a predetermined time period.
[0079] In an embodiment of the present application, the predetermined time period can be a time period set according to demand; for example, if the target wake-up words include A, B and C, if the wake-up word D is detected or the wake-up word is not recognized during the detection of the current audio signal, it is determined that the audio signal does not contain the target wake-up word, and the digital signal processor 2 is controlled to power off; if the predetermined time period is 3 minutes, then if any one of the wake-up words A, B and C is detected, within 3 minutes after the detection, if no operation instruction is detected, the digital signal processor 2 is controlled to power off.
[0080] In some examples, such as Figure 4 As shown in (b), in the embodiment of the present application, the processing device 1 performs voice activity detection on the received audio signal, that is, the processing device 1 needs to remain in a normally open state, and control the digital signal processor 3 to power on when voice activity is detected in the audio signal, and the digital signal processor 3 powers off when it is recognized that the audio signal does not contain the target wake-up word, and powers off when no operation instruction is detected in the audio signal within a predetermined time period; compared with the related art such as Figure 4 In the case shown in (a), voice activity detection needs to be performed by a digital signal processor, so that the digital signal processor is always turned on. The embodiment of the present application can save power consumption to a greater extent.
[0081] Optionally, the digital signal processor 2 and the application processor 3 each include at least one voice processing module; the at least one voice processing module is used to execute an operation instruction when an operation instruction is received.
[0082] In some examples, both the digital signal processor (DSP) and the application processor (AP) typically contain specially designed voice processing modules to support various voice-related functions. These modules can exist in the DSP or AP alone, or both can work together to achieve efficient and low-power voice processing capabilities.
[0083] As an optional method, the voice processing module may include but is not limited to an automatic speech recognition (ASR) module, a natural language processing (NLP) module, a speech synthesis (TTS, Text-to-Speech) module, a speaker verification / recognition module, a multimodal fusion module, etc. The voice processing module can execute operation instructions when the operation instructions are received.
[0084] Compared with the current related art, this embodiment configures the feature extraction module and the voice activity detection module in the processing device, and connects the voice activity detection module with the feature extraction module and the digital signal processor, so that the voice activity detection module performs voice activity detection on the received audio signal in the low power consumption mode, and controls the digital signal processor to power on when the voice activity is detected in the audio signal, so as to process the audio signal through the digital signal processor; since the voice activity detection module and the feature extraction module need to be kept in a normally open state to detect the voice signal, this embodiment can configure the voice activity detection module and the feature extraction module in the processing device, so that the voice activity detection module and the feature extraction module can be detected in the low power consumption mode, and the digital signal processor is controlled to power on when the voice activity is detected in the audio signal, and the audio signal can be processed by the digital signal processor, so that during the detection process of the voice activity detection module and the feature extraction module, the digital signal processor can be kept in a power-off state, and is only powered on when the voice activity is detected, which can save power consumption to a large extent and avoid power consumption waste.
[0085] This embodiment provides a voice activity detection method, such as Figure 5 As shown, it is applied to Figure 1 The processing device shown, the method includes:
[0086] Step 101: Continuously receive audio signals and extract feature information from the audio signals.
[0087] It should be noted that the execution subject of the embodiments of the present application may be the processing device in the above embodiments, or may be other processors that can achieve the same effect as the processing device in the above embodiments.
[0088] Step 102: Perform voice activity detection on the audio signal based on the feature information, and send the audio signal to a digital signal processor for processing when voice activity is detected in the audio signal.
[0089] Optionally, step 102 specifically further includes: marking a speech segment in the audio signal with a speech activity detection identifier based on the speech activity detection result, to obtain an audio signal including the speech activity detection identifier.
[0090] In an embodiment of the present application, the voice activity detection identifier can be an identifier corresponding to a speech segment, and the non-speech segment can also be marked with a corresponding identifier. For example, the voice activity detection identifier can be 1, and the identifier corresponding to the non-speech segment can be 0, or the voice activity detection identifier can be 0, and the identifier corresponding to the non-speech segment can be 1, and so on. The voice activity detection identifier can be set according to needs and is not limited here.
[0091] Optionally, step 102 may specifically include: determining that there is voice activity in the audio signal when the audio signal includes a voice activity detection flag, and controlling the digital signal processor to power on and sending the audio signal to the digital signal processor.
[0092] For this embodiment, if there is a voice activity detection flag in the audio signal, it means there is voice activity in the audio signal, and it is necessary to control the digital signal processor to power on and send the audio signal to the digital signal processor.
[0093] Compared with the current related art, this embodiment configures the feature extraction module and the voice activity detection module in the processing device, and connects the voice activity detection module with the feature extraction module and the digital signal processor, so that the voice activity detection module performs voice activity detection on the received audio signal in the low power consumption mode, and controls the digital signal processor to power on when the voice activity is detected in the audio signal, so as to process the audio signal through the digital signal processor; since the voice activity detection module and the feature extraction module need to be kept in a normally open state to detect the voice signal, this embodiment can configure the voice activity detection module and the feature extraction module in the processing device, so that the voice activity detection module and the feature extraction module can be detected in the low power consumption mode, and the digital signal processor is controlled to power on when the voice activity is detected in the audio signal, and the audio signal can be processed by the digital signal processor, so that during the detection process of the voice activity detection module and the feature extraction module, the digital signal processor can be kept in a power-off state, and is only powered on when the voice activity is detected, which can save power consumption to a large extent and avoid power consumption waste.
[0094] Based on the above Figure 5 The method shown in the embodiment also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned Figure 5 The method shown.
[0095] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0096] Based on the above Figure 5 In order to achieve the above-mentioned purpose, the present application also provides an electronic device including a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figure 5 The method shown.
[0097] Optionally, the above-mentioned physical device may also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0098] Those skilled in the art will appreciate that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different arrangements of components.
[0099] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned physical device, and supports the operation of the information processing program and other software and / or programs. The network communication module is used to realize the communication between the components inside the storage medium, and the communication with other hardware and software in the information processing physical device.
[0100] Through the description of the above implementation methods, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or by hardware. By applying the solution of this embodiment, compared with the current related art, this embodiment configures the feature extraction module and the voice activity detection module in the processing device, and connects the voice activity detection module with the feature extraction module and the digital signal processor, so that the voice activity detection module performs voice activity detection on the received audio signal in the low power consumption mode, and controls the digital signal processor to power on when the voice activity is detected in the audio signal, so that the audio signal is processed by the digital signal processor; because the voice activity detection module and the feature extraction module need to be kept in a normally open state to detect the voice signal, this embodiment can configure the voice activity detection module and the feature extraction module in the processing device, so that the voice activity detection module and the feature extraction module can be detected in the low power consumption mode, and the digital signal processor is controlled to power on when the voice activity is detected in the audio signal, and the audio signal can be processed by the digital signal processor, so that during the detection process of the voice activity detection module and the feature extraction module, the digital signal processor can be kept in a power-off state, and only powered on when the voice activity is detected, which can save power consumption to a large extent and avoid power consumption waste.
[0101] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0102] The above is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features applied for herein.
Claims
1. A processing device, characterized in that: include: Feature extraction module and voice activity detection module; The voice activity detection module is connected to the feature extraction module; The feature extraction module is used to continuously receive an audio signal, extract feature information from the audio signal, and send the feature information and the audio signal to the voice activity detection module; The voice activity detection module is used to perform voice activity detection on the audio signal based on the feature information, and send the audio signal to the digital signal processor for processing when voice activity is detected in the audio signal.
2. The processing device according to claim 1, characterized in that The processing device operates with power consumption less than a predetermined power consumption threshold.
3. The processor according to claim 1, wherein: The voice activity detection module is specifically configured to mark a voice segment in the audio signal with a voice activity detection identifier based on a voice activity detection result.
4. The processing device according to claim 3, characterized in that The voice activity detection module is further specifically configured to determine whether there is voice activity in the audio signal when the audio signal contains the voice activity detection identifier, and control the digital signal processor to power on and send the audio signal to the digital signal processor.
5. A voice activity detection system, characterized in that: include: A digital signal processor and a processing device as claimed in any one of claims 1 to 4 connected to each other; The processing device is used to continuously receive an audio signal, extract feature information from the audio signal, perform voice activity detection on the audio signal based on the feature information, and send the audio signal to a digital signal processor for processing when voice activity is detected in the audio signal.
6. The system according to claim 5, characterized in that Also includes: an application processor connected to the digital signal processor; The digital signal processor is used to detect an operation instruction contained in the audio signal when recognizing that there is a target wake-up word in the audio signal, and send the operation instruction to the application processor or execute the operation instruction; The application processor is configured to execute the operation instruction upon receiving the operation instruction.
7. The system according to claim 6, characterized in that The digital signal processor is further used to recognize a wake-up word in the audio signal, and remain powered on when it is recognized that there is a target wake-up word in the audio signal; The digital signal processor is further configured to detect the audio signal in a power-on state, and when the operation instruction is detected in the audio signal, control the application processor to power on and send the operation instruction to the application processor or execute the operation instruction.
8. The system according to claim 7, characterized in that The digital signal processor is also used to power off when it is recognized that the audio signal does not contain the target wake-up word, and / or power off when no operation instruction is detected in the audio signal within a predetermined time period.
9. The system according to claim 6, characterized in that The digital signal processor and the application processor each include at least one voice processing module; The at least one voice processing module is used to execute the operation instruction when receiving the operation instruction.
10. A voice activity detection method, characterized in that: The processing device according to any one of claims 1 to 4, wherein the method comprises: Continuously receiving an audio signal and extracting feature information from the audio signal; Voice activity detection is performed on the audio signal based on the feature information, and when voice activity is detected in the audio signal, the audio signal is sent to a digital signal processor for processing.
11. The method according to claim 10, characterized in that After performing voice activity detection on the audio signal based on the feature information, the method further includes: Based on the voice activity detection result, the voice segment in the audio signal is marked with a voice activity detection identifier to obtain an audio signal containing the voice activity detection identifier.
12. The method according to claim 11, characterized in that The processing of the audio signal when voice activity is detected in the audio signal comprises: If the audio signal includes the voice activity detection flag, it is determined that there is voice activity in the audio signal, and the digital signal processor is controlled to power on and the audio signal is sent to the digital signal processor.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 10 to 12 is implemented.
14. An electronic device, characterized in that: The method comprises a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the method according to any one of claims 10 to 12 when executing the computer program.
15. A chip, characterized in that: It comprises one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send the signal to the processor, the signal comprising a computer instruction stored in the memory; when the processor executes the computer instruction, the electronic device executes the method described in any one of claims 10 to 12.
Citation Information
Patent Citations
Multistage recognition voice awakening method and device, computer storage medium and equipment
CN111199733A
Voice wake-up template acquisition method and device, electronic equipment and computer readable storage medium
CN111326146A
Voice acquisition device, controller, control method and voice acquisition control system
CN113990311A
Voice wake-up method, system, device and equipment
CN116386643A
Low-power-consumption voice wake-up system and method and electronic equipment
CN118865961A
Cited By
Voice activity signal detection method and voice activity signal detection system
CN120452428A
Voice activity signal detection method and voice activity signal detection system
CN120452428B