Voice wake-up method and system based on heterogeneous chip collaboration
Through a heterogeneous chip collaborative architecture, the low-power dedicated voice processing unit performs audio monitoring and wake-up word detection in standby mode, while the main processing unit handles complex tasks after wake-up. This solves the problem of balancing power consumption and performance in existing technologies and achieves a balance between low power consumption and high performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-03
AI Technical Summary
In existing voice wake-up solutions, devices based on a single main control chip have high standby power consumption. Homogeneous or similar dual-chip solutions are difficult to achieve high-precision front-end acoustic processing and wake-up word recognition in a low-power state, and cannot efficiently complete complex calculations or network interactions after wake-up.
It adopts a heterogeneous chip collaborative architecture, utilizing a low-power dedicated voice processing unit to independently perform audio monitoring and wake word detection in standby mode, while the main processing unit is responsible for handling complex tasks after wake-up, including acoustic front-end processing and network interaction.
It achieves a balance between low-power listening in standby mode and high-performance processing after wake-up, reduces the overall power consumption of the system, ensures the efficient execution of complex tasks, and improves the overall performance of the voice wake-up system.
Smart Images

Figure CN121789673A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech recognition technology, specifically relating to a speech wake-up method and system based on heterogeneous chip collaboration. Background Technology
[0002] Intelligent voice wake-up is a key function of IoT devices. Existing voice wake-up solutions mainly face two types of problems: First, solutions based on a single main control chip require the chip to be continuously operational to perform audio monitoring and wake-up word recognition, resulting in high standby power consumption and severely limiting device battery life. Second, solutions using homogeneous or similarly performing dual-chip architectures can partially reduce power consumption, but due to limitations in chip computing power and architecture, it is difficult to achieve high-precision front-end acoustic processing and wake-up word recognition in a low-power state. Furthermore, after wake-up, it cannot efficiently complete voice command processing requiring complex calculations or network interactions. Therefore, existing technologies suffer from the problem of balancing device standby power consumption with the overall performance of the voice wake-up system. Summary of the Invention
[0003] The purpose of this invention is to provide a voice wake-up method and system based on heterogeneous chip collaboration to solve the problems mentioned in the background art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: a voice wake-up method based on heterogeneous chip collaboration, comprising:
[0005] In standby mode, the main processing unit enters sleep mode, and a dedicated voice processing unit with lower power consumption than the main processing unit runs independently to continuously collect and analyze audio signals.
[0006] When the dedicated voice processing unit identifies a preset wake-up word in the audio signal, it sends a wake-up signal to the main processing unit.
[0007] The main processing unit is activated in response to the wake-up signal and obtains subsequent voice commands from the dedicated voice processing unit.
[0008] The main processing unit processes the subsequent voice commands and controls the execution of corresponding operations based on the processing results.
[0009] Preferably, the main processing unit processes the subsequent voice commands, including:
[0010] Determine the complexity of the subsequent voice commands;
[0011] If the instruction is determined to be a low-complexity local instruction, the main processing unit will invoke local resources to execute it.
[0012] If the instruction is determined to be a highly complex cloud-based command, the main processing unit will send the subsequent voice instruction or its features to the cloud server via the wireless communication module and receive the response from the cloud server.
[0013] Preferably, after the main processing unit completes processing the subsequent voice commands, the method further includes:
[0014] Start timeout counting;
[0015] If no new valid instruction is received before the timeout reaches the preset threshold, the main processing unit saves the current context state and re-enters sleep mode, allowing the dedicated voice processing unit to resume independent operation, thereby restoring the entire system to the standby state.
[0016] Preferably, at the same time or after the dedicated voice processing unit sends the wake-up signal to the main processing unit, it also transmits audio data processed by the acoustic front end to the main processing unit, wherein the acoustic front end processing includes at least echo cancellation.
[0017] A voice wake-up system based on heterogeneous chip collaboration includes a main processing unit, a dedicated voice processing unit, an audio input unit, an audio output unit, and a power management unit.
[0018] The power consumption of the dedicated voice processing unit is lower than that of the main processing unit;
[0019] The audio input unit is electrically connected to the dedicated voice processing unit for collecting ambient sound.
[0020] The dedicated voice processing unit is configured to work independently when the system is in standby mode, continuously analyze the audio signals collected by the audio input unit, and wake up the main processing unit when a wake-up word is recognized.
[0021] The main processing unit is configured to: after being awakened from sleep mode, obtain voice commands from the dedicated voice processing unit and execute them;
[0022] The audio output unit is controlled by the main processing unit;
[0023] The power management unit provides power to the main processing unit and the dedicated voice processing unit, and supports independent control of the power supply to the main processing unit.
[0024] Preferably, the system further includes an audio amplification unit, the input of which is connected to the first audio output of the main processing unit, and the output of which is connected to the audio output unit; the feedback input of the audio amplification unit is connected to the reference signal output of the dedicated speech processing unit, forming an acoustic echo cancellation loop.
[0025] Preferably, the dedicated voice processing unit is connected to the main processing unit via at least one control signal line and one data bus; wherein the control signal line is used to transmit the wake-up signal, and the data bus is used to transmit audio data or control commands.
[0026] Preferably, the power management unit includes:
[0027] The charging management circuit connects the external power source and the internal battery.
[0028] A controlled switch circuit is connected in series between the charging management circuit and the power input terminal of the main processing unit;
[0029] The first control path is that the control terminal of the controlled switch circuit is connected to a general-purpose input / output pin of the main processing unit, which is used to perform software shutdown control on the main processing unit after it is powered on.
[0030] The second control path includes a physical wake-up button connected to the control terminal of the controlled switch circuit, forming a hardware enable path in parallel with the first control path. This enable path is used to directly control the controlled switch circuit to power on the main processing unit when the system is completely powered off.
[0031] Preferably, the main processing unit is a microcontroller or application processor with integrated wireless communication function; the dedicated voice processing unit is a low-power voice chip that integrates an analog microphone interface, a digital signal processor and a wake-word recognition algorithm.
[0032] Preferably, the audio input unit is a multi-microphone array, and the dedicated voice processing unit also integrates a sound source localization or beamforming algorithm based on the multi-microphone array to improve the robustness of wake word recognition.
[0033] Compared with the prior art, the beneficial effects of the present invention are:
[0034] This invention employs a heterogeneous collaborative architecture of a main processing unit and a dedicated voice processing unit to achieve a balance between standby power consumption and overall performance in a voice wake-up system. In standby mode, the low-power dedicated voice processing unit independently performs audio monitoring and wake-up word detection, significantly reducing overall system power consumption. After wake-up, the high-performance main processing unit takes over the processing and response to subsequent voice commands, ensuring efficient execution of complex tasks. This solution addresses the challenge of balancing high power consumption and high performance at the system architecture level, creating a unified system architecture that combines low-power continuous monitoring with high-performance on-demand processing. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating the voice wake-up method of the present invention.
[0036] Figure 2 This is a schematic diagram of the first processing method of the main processing unit of the present invention.
[0037] Figure 3 This is a schematic diagram of the second processing method of the main processing unit of the present invention.
[0038] Figure 4 This is a schematic diagram of the speech signal processing flow of the present invention.
[0039] Figure 5 This is a flowchart illustrating the voice wake-up system of the present invention.
[0040] Figure 6 This is a schematic diagram of the echo cancellation process of the present invention.
[0041] Figure 7 This is a schematic diagram of the power management process of the present invention.
[0042] Figure 8 This is a schematic diagram of the microphone signal processing flow of the present invention. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Example 1:
[0045] A voice wake-up method based on heterogeneous chip collaboration includes:
[0046] In standby mode, the main processing unit enters sleep mode, and a dedicated voice processing unit with lower power consumption than the main processing unit runs independently to continuously collect and analyze audio signals.
[0047] When the dedicated voice processing unit recognizes the preset wake-up word in the audio signal, it sends a wake-up signal to the main processing unit.
[0048] The main processing unit is activated in response to the wake-up signal and obtains subsequent voice commands from the dedicated voice processing unit.
[0049] The main processing unit processes subsequent voice commands and controls the execution of corresponding operations based on the processing results. The main processing unit processes subsequent voice commands, including:
[0050] Determine the complexity of subsequent voice commands;
[0051] If the instruction is determined to be a low-complexity local instruction, the main processing unit will call local resources to execute it.
[0052] If the command is determined to be a high-complexity cloud instruction, the main processing unit sends the subsequent voice command or its characteristics to the cloud server via the wireless communication module and receives the response from the cloud server. After the main processing unit completes the processing of the subsequent voice command, the method further includes:
[0053] Start timeout counting;
[0054] If no new valid instruction is received before the timeout reaches a preset threshold, the main processing unit saves its current context state and re-enters sleep mode, allowing the dedicated voice processing unit to resume independent operation, thus restoring the entire system to standby mode. Simultaneously or subsequently, the dedicated voice processing unit sends a wake-up signal to the main processing unit, and also transmits audio data processed by the acoustic front-end to the main processing unit. The acoustic front-end processing includes at least echo cancellation.
[0055] A voice wake-up system based on heterogeneous chip collaboration includes a main processing unit, a dedicated voice processing unit, an audio input unit, an audio output unit, and a power management unit.
[0056] The dedicated voice processing unit consumes less power than the main processing unit;
[0057] The audio input unit is electrically connected to a dedicated voice processing unit for collecting ambient sound.
[0058] The dedicated voice processing unit is configured to work independently when the system is in standby mode, continuously analyze the audio signals collected by the audio input unit, and wake up the main processing unit when a wake word is recognized.
[0059] The main processing unit is configured to: after being awakened from sleep mode, retrieve voice commands from the dedicated voice processing unit and execute them;
[0060] The audio output unit is controlled by the main processing unit;
[0061] The power management unit provides power to the main processing unit and the dedicated voice processing unit, and supports independent control of the power supply to the main processing unit. The system also includes an audio amplification unit; the input of the audio amplification unit is connected to the first audio output of the main processing unit, and the output is connected to an audio output unit; the feedback input of the audio amplification unit is connected to the reference signal output of the dedicated voice processing unit, forming an acoustic echo cancellation loop. The dedicated voice processing unit is connected to the main processing unit via at least one control signal line and one data bus; the control signal line is used to transmit wake-up signals, and the data bus is used to transmit audio data or control commands. The power management unit includes:
[0062] The charging management circuit connects the external power source and the internal battery.
[0063] A controlled switch circuit is connected in series between the charging management circuit and the power input terminal of the main processing unit;
[0064] The first control path connects the control terminal of the controlled switch circuit to a general-purpose input / output pin of the main processing unit, which is used to perform software shutdown control on the main processing unit after it is powered on.
[0065] The second control path connects the control terminal of the controlled switch circuit to a physical wake-up button, forming a hardware enable path in parallel with the first control path. This enable path directly controls the controlled switch circuit to power on the main processing unit when the system is completely powered off. The main processing unit is a microcontroller or application processor with integrated wireless communication functionality. The dedicated voice processing unit is a low-power voice chip integrating an analog microphone interface, a digital signal processor, and a wake-up word recognition algorithm. The audio input unit is a multi-microphone array, and the dedicated voice processing unit also integrates a sound source localization or beamforming algorithm based on the multi-microphone array to improve the robustness of wake-up word recognition.
[0066] Through the above technical solution, this invention adopts a heterogeneous collaborative architecture of a main processing unit and a dedicated voice processing unit, achieving a balance between standby power consumption and overall performance in the voice wake-up system. In standby mode, the low-power dedicated voice processing unit independently performs audio monitoring and wake-up word detection, significantly reducing the overall system power consumption. After wake-up, the high-performance main processing unit takes over the processing and response to subsequent voice commands, ensuring the efficient execution of complex tasks. This solution addresses the problem of the difficulty in coexisting high power consumption and high performance at the system architecture level, forming a unified system architecture that combines low-power continuous monitoring with high-performance on-demand processing.
[0067] Example 2:
[0068] The method in this embodiment is applied to a smart voice device, which includes a main processing unit and a dedicated voice processing unit. The dedicated voice processing unit is a chip specifically designed for audio signal processing, and its operating power consumption is lower than that of the main processing unit.
[0069] When the device is in standby mode, the main processing unit enters sleep mode to reduce power consumption. During this time, the dedicated voice processing unit operates independently. This unit continuously acquires audio signals from the environment and analyzes and processes them. The analysis process includes front-end acoustic processing of the audio signals, such as noise reduction and feature extraction, and then executing a wake-word detection algorithm. Due to the dedicated voice processing unit's specialized architecture and low power consumption, the device can perform continuous audio monitoring and preliminary analysis with low power consumption even in standby mode.
[0070] When the dedicated voice processing unit identifies an acoustic feature pattern matching a preset wake-up word in the continuously analyzed audio signal, it determines that the wake-up word has appeared. At this time, the dedicated voice processing unit generates a wake-up signal and sends the signal to the main processing unit, which is in sleep mode, via a hardware interrupt or communication interface.
[0071] Upon receiving a wake-up signal from the dedicated voice processing unit, the main processing unit is triggered and activated from sleep mode, entering normal operation. Once activated, the main processing unit establishes communication with the dedicated voice processing unit to obtain subsequent voice command data collected after the wake-up word is recognized. This data consists of the specific voice commands issued by the user.
[0072] After acquiring subsequent voice command data, the main processing unit begins processing it. The main processing unit possesses enhanced general-purpose computing capabilities, enabling it to perform complex speech recognition and semantic understanding tasks, and can connect to the network for cloud-based interaction. Once the main processing unit has processed the command, it generates control commands based on the semantic understanding results and sends them to the corresponding actuators in the device to control the device to perform the user-required operations, such as playing music, querying information, or controlling home appliances.
[0073] Example 3:
[0074] In this embodiment, after the main processing unit is awakened and obtains subsequent voice commands, it needs to process the commands. The core of the processing lies in judging the complexity of the command and selecting different execution paths based on the judgment result.
[0075] The main processing unit first analyzes the acquired subsequent voice commands to determine their complexity. This determination is primarily based on the computational resource requirements and interaction patterns involved in the command content. For example, if a command only involves simple local device control or status queries, and its processing does not rely on external networks or extensive computing, it is classified as a low-complexity local command. Conversely, if a command requires access to a large-scale knowledge base, complex semantic understanding, or access to external network services, it is classified as a high-complexity cloud command.
[0076] When the main processing unit determines that the subsequent voice command is a low-complexity local command, the processing will be completed internally within the device. The main processing unit will utilize the device's local software and hardware resources to execute the command. For example, for commands like "turn up the volume" or "turn on the lights," the main processing unit directly parses the command content, generates the corresponding control signals, and drives the corresponding functional modules to perform the operation. The entire processing flow does not require data interaction with an external network, has a fast response speed, and is completed entirely within the device.
[0077] When the main processing unit determines that the subsequent voice command is a highly complex cloud-based command, processing requires the cooperation of the cloud server. The main processing unit sends the raw audio data of the subsequent voice command, or the feature parameters extracted from the audio, to the remote cloud server via the device's integrated wireless communication module. Leveraging its powerful computing capabilities and stored data resources, the cloud server performs in-depth processing of the command, such as natural language understanding, contextual dialogue management, or information retrieval. After processing, the cloud server sends the generated response back via the network. The main processing unit receives this response via the wireless communication module and controls the device to perform corresponding operations based on the result, such as broadcasting the retrieved weather information or executing a complex smart home scene linkage.
[0078] Through the aforementioned judgment and traffic distribution mechanisms, the system achieves a rational allocation of computing resources. Simple instructions are responded to quickly by the local main processing unit, ensuring the real-time performance of basic operations; complex instructions are completed with the help of cloud computing power, breaking through the performance limitations of the local chip and enabling the device to handle diverse voice interaction tasks. This processing mode of local and cloud collaboration constitutes a key component of the main processing unit's function in the heterogeneous chip collaborative architecture.
[0079] Example 4:
[0080] In this embodiment, after the main processing unit completes the corresponding operation based on the acquired subsequent voice command, the system does not immediately return to standby mode. At this time, the system will activate a timeout function. This timeout function is used to monitor whether the user will input new voice commands within a short period of time. Its core purpose is that after the main processing unit completes the current task, if the user does not have any new interaction needs for the time being, the system can automatically restore to a low-power standby mode, thereby saving overall energy consumption.
[0081] After the timeout period begins, the system continuously monitors the audio input channel. If the system does not receive any new, valid voice commands before the timeout reaches a preset threshold, it will execute a state saving and switching process. At this time, the main processing unit will first save the current context state. This context state may include the current device operating mode, session information from the user's last interaction, or temporary data for unfinished tasks. Saving these states ensures that when the user wakes up the device again, the main processing unit can quickly resume its previous working state, providing a continuous interactive experience and avoiding the need for the user to repeat previous commands or settings.
[0082] After saving the context state, the main processing unit will re-enter sleep mode. Sleep mode for the main processing unit means that most of its functional modules cease operation, retaining only the extremely low-power basic monitoring circuitry to receive wake-up signals from the dedicated voice processing unit. With the main processing unit entering sleep mode, system control is once again handed over to the dedicated voice processing unit. The dedicated voice processing unit resumes independent operation, continuously acquiring and analyzing ambient audio signals, focusing on detecting the preset wake-up word.
[0083] At this point, the entire system has returned from its high-performance operating state after being woken up to its initial standby state. In standby mode, the low-power dedicated voice processing unit handles the continuous monitoring task, while the high-performance, high-power main processing unit is in a sleep-energy-saving state. This automatic state rollback triggered by a timeout mechanism constitutes a complete, energy-efficient interactive loop. It ensures that after responding to user commands, the device can intelligently determine the user's intent, promptly reduce power consumption when there are no subsequent commands, and be quickly woken up to provide high-performance services when needed by the user. This achieves an effective combination of long-term low-power monitoring and instantaneous high-performance processing at the system level.
[0084] Example 5:
[0085] In standby mode, the main processing unit enters sleep mode, and a dedicated, low-power voice processing unit operates independently. This dedicated voice processing unit continuously collects and analyzes audio signals from the environment to detect whether they contain a preset wake-up word. When the dedicated voice processing unit identifies the preset wake-up word in the audio signal, it sends a wake-up signal to the main processing unit.
[0086] Simultaneously or subsequently, the dedicated voice processing unit sends a wake-up signal to the main processing unit, and also transmits audio data processed by the acoustic front-end to the main processing unit. This acoustic front-end processing includes at least echo cancellation. Echo cancellation effectively suppresses echoes generated by the device's own speakers, thereby improving the clarity and recognition accuracy of subsequent voice commands. The dedicated voice processing unit operates continuously in a low-power state, integrating hardware modules optimized for audio signal processing, thus enabling it to complete acoustic front-end processing tasks, including echo cancellation, while maintaining low power consumption.
[0087] The main processing unit is activated from sleep mode in response to the received wake-up signal. Upon activation, the main processing unit retrieves subsequent voice commands from the dedicated voice processing unit. These subsequent voice commands are based on the audio data processed by the aforementioned acoustic front-end. Because the audio data has undergone preliminary processing, particularly echo cancellation, the quality of the voice commands received by the main processing unit is improved, creating favorable conditions for subsequent accurate recognition and processing.
[0088] Once activated, the main processing unit utilizes its enhanced general-purpose computing power to process subsequent voice commands. This processing can include more complex speech recognition, semantic understanding, and potential network interactions. Based on the processing results, the main processing unit controls the device to perform corresponding operations, such as playing music, querying information, or controlling smart home devices. Through this collaborative approach, the system maintains low-power listening and initial processing by a dedicated voice processing unit during standby, while the main processing unit provides high-performance, complex task processing capabilities upon wake-up, thus achieving a balance between low power consumption and high performance overall.
[0089] Example 6:
[0090] The system in this embodiment includes a main processing unit, a dedicated voice processing unit, an audio input unit, an audio output unit, and a power management unit. The dedicated voice processing unit consumes less power than the main processing unit. The audio input unit is electrically connected to the dedicated voice processing unit and its function is to collect sound signals from the environment.
[0091] When the system is in standby mode, the main processing unit enters sleep mode to reduce power consumption. At this time, the dedicated voice processing unit operates independently, continuously receiving and analyzing audio signals from the audio input unit. This unit integrates a wake-up word recognition algorithm, enabling real-time processing of the continuous audio stream. When the analysis confirms that the acquired audio signal contains a preset wake-up word, the dedicated voice processing unit generates a wake-up signal.
[0092] The wake-up signal is sent to the power management unit and the main processing unit. The power management unit provides power to the entire system and has the ability to independently control the power supply to the main processing unit. Upon receiving the wake-up signal, the power management unit restores full power supply to the main processing unit. The main processing unit is then awakened from hibernation and enters normal operating mode.
[0093] After the main processing unit is woken up, it communicates with the dedicated speech processing unit to retrieve audio data containing complete speech commands or corresponding processing results collected before and after the wake-up time from its internal buffer or registers. Subsequently, the main processing unit executes subsequent speech command processing tasks. These tasks may include more complex natural language understanding, semantic parsing, or queries and operations requiring network connection for cloud interaction.
[0094] The audio output unit, such as a speaker, is controlled by the main processing unit. Based on the processing results of the voice command, the main processing unit generates corresponding response content and plays it through the audio output unit, thus completing a full voice interaction. Throughout this process, the heterogeneous architecture achieves division of labor: a low-power dedicated voice processing unit is responsible for long-term, uninterrupted monitoring and initial recognition; the high-performance main processing unit is only activated when needed, responsible for computationally intensive subsequent processing and system control. This design allows the system to maintain extremely low standby power consumption without sacrificing post-wake-up processing performance and responsiveness.
[0095] Example 7:
[0096] The system in this embodiment includes an audio amplification unit, which works in conjunction with the main processing unit and a dedicated speech processing unit to achieve acoustic echo cancellation.
[0097] In the system architecture, the input of the audio amplification unit is connected to the first audio output of the main processing unit. The output of the audio amplification unit is connected to the audio output unit. The feedback input of the audio amplification unit is connected to the reference signal output of the dedicated speech processing unit. This connection method constitutes an acoustic echo cancellation loop.
[0098] When the system is in operation, the main processing unit is in sleep mode while the dedicated voice processing unit operates independently. The dedicated voice processing unit continuously collects and analyzes the audio signals sent from the audio input unit. When the dedicated voice processing unit recognizes a preset wake-up word in the audio signal, it sends a wake-up signal to the main processing unit. The main processing unit is activated in response to this signal and retrieves subsequent voice commands from the dedicated voice processing unit for processing.
[0099] When the main processing unit is woken up and needs to play audio, such as to announce a wake-up response or feedback on an executed command, the main processing unit generates an audio signal from its first audio output. This signal is sent to the input of the audio amplification unit. The audio amplification unit amplifies the signal and then drives the audio output unit to emit sound.
[0100] During this process, the sound emitted by the audio output unit may be re-acquired by the audio input unit, creating an acoustic echo. If this echo is not eliminated, it may interfere with the system's accurate recognition of subsequent user voice commands. To address this issue, when the dedicated voice processing unit transmits the acoustically processed audio data to the main processing unit, it outputs a reference signal generated internally for echo cancellation from its reference signal output terminal. This reference signal essentially reflects the audio content that is about to be played or is currently being played by the audio output unit.
[0101] The feedback input of the audio amplification unit receives the reference signal from the reference signal output of the dedicated speech processing unit. The audio amplification unit integrates echo cancellation processing circuitry or logic. This circuit compares and calculates the received reference signal with the audio signal fed back from the audio input path, which may contain echoes. Through algorithms such as adaptive filtering, the audio amplification unit can predict and generate a cancellation signal that is the opposite of the echo signal.
[0102] The cancellation signal is superimposed on the main audio signal to be amplified inside the audio amplification unit, thereby pre-canceling any potential acoustic echo components before the signal is amplified and sent to the audio output unit. In this way, the portion of the sound emitted from the audio output unit that might cause echo interference is pre-attenuated. Even if residual sound is picked up by the audio input unit, its echo effect is greatly reduced, allowing the dedicated voice processing unit or main processing unit to more clearly capture and process the user's valid speech, improving the reliability and experience of voice interaction.
[0103] By adding an audio amplification unit and forming the aforementioned loop, the system achieves efficient acoustic echo cancellation at the hardware level, eliminating the need for the main processing unit to perform complex real-time software calculations, reducing its processing load, and simultaneously improving the anti-interference capability of the entire voice interaction link.
[0104] Example 8:
[0105] In the system hardware architecture, a dedicated voice processing unit is connected to the main processing unit via physical lines. These lines include at least one control signal line and one data bus. The control signal line can be unidirectional or bidirectional, and its core function is to transmit a wake-up signal. When the dedicated voice processing unit successfully identifies a preset wake-up word in the continuously analyzed audio signal, it generates a digital level signal or a specific pulse sequence as a wake-up signal. This wake-up signal is directly sent to a specific wake-up pin of the main processing unit via this dedicated control signal line. This point-to-point signal transmission method has a clear path, low latency, and can reliably trigger the main processing unit to switch from sleep mode to active working mode.
[0106] The data bus is a high-speed communication channel, either parallel or serial, used to transmit data between the main processing unit and the main processing unit after the main processing unit is awakened. The data bus has two main functions. First, it transmits audio data. Depending on the system design, simultaneously with or after sending the wake-up signal, the dedicated voice processing unit can send audio data processed by its internal acoustic front-end to the main processing unit via the data bus. This audio data contains subsequent voice commands following the wake-up word, providing raw material for the main processing unit to perform command recognition and semantic understanding. Acoustic front-end processing may include echo cancellation, noise suppression, etc., to improve audio quality. Second, the data bus also transmits control commands. After the main processing unit is awakened, it can send control commands to the dedicated voice processing unit via the data bus, such as requesting audio data of a specific format, adjusting the operating parameters of the dedicated voice processing unit, or querying its status. Conversely, the dedicated voice processing unit can also provide status information or respond to queries from the main processing unit via the data bus.
[0107] This connection design, which separates the transmission of wake-up signals from the data stream, offers clear advantages. The control signal line is dedicated to transmitting critical wake-up event signals, ensuring the immediacy and determinism of the wake-up action, unaffected by potential communication congestion or protocol processing delays on the data bus. The data bus, on the other hand, focuses on transmitting large volumes of audio data and flexible bidirectional command interaction, meeting the needs of the main processing unit to obtain the necessary information for processing after wake-up and to coordinate the work of both units. The control signal line and the data bus work together to form the physical foundation and communication guarantee for efficient and reliable collaboration between heterogeneous chips, enabling seamless integration of low-power monitoring and high-performance processing.
[0108] Example 9:
[0109] In this embodiment, the charging management circuit is connected to the external power supply and the internal battery, and is responsible for managing the battery's charging. A controlled switch circuit is connected in series between the charging management circuit and the power input terminal of the main processing unit; its function is to control the on / off state of the main processing unit's power supply circuit.
[0110] The control terminal of the controlled switch circuit is connected to a general-purpose input / output pin of the main processing unit through a first control path. After the main processing unit powers on and completes initialization, its software can output a specific level signal to this pin to control the shutdown of the controlled switch circuit. This allows the main processing unit to actively cut off its own power supply via software commands after completing voice command processing tasks, thus entering a completely power-off state, rather than merely entering a low-power sleep mode. This helps to further reduce the system's total power consumption during non-working periods.
[0111] The second control path is connected in parallel with the first control path, linking the control terminal of the controlled switch circuit to a physical wake-up button. This path constitutes the hardware enable path. When the system is completely powered off, i.e., the main processing unit has no power supply, the first control path fails due to the power loss of the main processing unit. At this time, when the user presses the physical wake-up button, a valid voltage level can be applied directly to the control terminal of the controlled switch circuit through the second control path, forcibly turning on the controlled switch circuit, thereby restoring power supply to the power input terminal of the main processing unit and realizing hardware power-on wake-up.
[0112] The power management unit, designed in conjunction with the methods in Embodiments 1 and 2, achieves refined management of the power supply to the main processing unit. In standby mode, the dedicated voice processing unit is continuously powered by the power management unit to maintain operation, while the main processing unit can be completely powered off via software control, significantly reducing standby power consumption. When the main processing unit needs to operate, it can be triggered by the dedicated voice processing unit recognizing a wake-up word and using a signal connection (such as an interrupt pin) between itself and the main processing unit to wake it from sleep mode (if the main processing unit is in sleep mode); alternatively, when the system is completely powered off, the user can directly power it on via a physical button. This dual control mechanism combining hardware and software enhances system reliability and user interaction flexibility, ensuring that the advantages of the heterogeneous chip collaborative architecture in terms of power consumption and performance are fully realized.
[0113] Example 10:
[0114] The system in this embodiment includes a main processing unit and a dedicated voice processing unit. The main processing unit employs a microcontroller or application processor with integrated wireless communication capabilities. The dedicated voice processing unit employs a low-power voice chip that integrates an analog microphone interface, a digital signal processor, and a wake-word recognition algorithm.
[0115] In system standby mode, the main processing unit is in hibernation mode. At this time, the dedicated voice processing unit operates independently due to its low power consumption. Ambient sound signals collected by the audio input unit are input to the dedicated voice processing unit via an analog microphone interface. The unit's internal digital signal processor preprocesses the raw audio signal, performing tasks such as noise reduction and gain control. Subsequently, its built-in wake-up word recognition algorithm continuously analyzes the processed audio stream to determine if it contains a preset wake-up word. Because the dedicated voice processing unit has a streamlined structure and is optimized for voice wake-up tasks, its power consumption is significantly lower than that of the main processing unit, thus maintaining an extremely low power consumption level for the entire system in standby mode.
[0116] When the dedicated voice processing unit successfully identifies a preset wake-up word in the audio signal, it sends a wake-up signal to the main processing unit via an interrupt or a specific communication interface. Upon receiving this signal, the main processing unit is activated from its sleep state and enters normal operation mode. After being awakened, the main processing unit issues instructions to the dedicated voice processing unit to obtain subsequent voice command data. The dedicated voice processing unit can transmit cached or real-time acquired subsequent audio data to the main processing unit.
[0117] Once activated, the main processing unit (MSU) is responsible for in-depth processing of subsequent voice commands. As the MSU is a microcontroller or application processor, it boasts greater computing power and richer resources. It can execute more complex speech recognition algorithms or perform semantic understanding of commands. Depending on processing needs, the MSU can utilize local resources to perform corresponding operations, such as controlling the audio output unit to play sound. For commands requiring network connectivity or complex calculations, the MSU can leverage its integrated wireless communication capabilities to send the voice command or extracted features to a cloud server, and receive and process the responses returned from the cloud, thereby completing higher-level interactive tasks.
[0118] Through the above methods, the system in this embodiment achieves the cooperation of heterogeneous chips. During most of the standby time, the highly integrated low-power voice chip independently handles listening and wake-up, significantly reducing system power consumption. Only when complex tasks need to be processed is the more powerful and feature-rich main processing unit activated. This architecture effectively solves the problems of high standby power consumption in single-chip solutions and the difficulty in optimizing performance and power consumption in homogeneous dual-chip solutions, achieving a balance between low-power continuous listening and high-performance on-demand processing.
[0119] Example 11:
[0120] The system in this embodiment includes a main processing unit, a dedicated voice processing unit, an audio input unit, an audio output unit, and a power management unit. The audio input unit employs a multi-microphone array. The dedicated voice processing unit integrates a sound source localization algorithm and a beamforming algorithm based on this multi-microphone array to improve the robustness of wake-word recognition.
[0121] When the system is in standby mode, the main processing unit enters sleep mode under the control of the power management unit to reduce power consumption. At this time, the low-power dedicated voice processing unit continues to operate independently. A multi-microphone array continuously collects audio signals from the environment and transmits the raw audio data to the dedicated voice processing unit. The dedicated voice processing unit first calls its integrated sound source localization algorithm to analyze the multi-microphone signals and calculate the direction or location of the target sound source. Subsequently, the system uses a beamforming algorithm to weight and synthesize the multi-microphone signals based on the sound source localization results, forming an enhanced beam pointing towards the target sound source while suppressing noise and interference from other directions. The signal-to-noise ratio of the audio signal after beamforming is improved.
[0122] A dedicated speech processing unit continuously detects and analyzes wake words on the audio signal after the acoustic front-end enhancement process described above. Because beamforming effectively focuses the speech in the target direction and suppresses environmental noise, the features of the wake word are clearer and more prominent in the audio stream. When the dedicated speech processing unit confirms and identifies a preset wake word during analysis, a valid wake-up event is determined to have occurred. At this time, the dedicated speech processing unit sends a wake-up signal to the main processing unit.
[0123] Upon receiving a wake-up signal, the main processing unit is activated by the power management unit, transitioning from sleep mode to normal operation. After waking, the main processing unit retrieves subsequent user voice command data from the dedicated voice processing unit. This command data is also pre-processed audio data using sound source localization and beamforming algorithms, resulting in good voice quality. Leveraging its superior computing power, the main processing unit performs recognition and semantic understanding on these subsequent voice commands, and controls the audio output unit to produce sound or execute other corresponding device operations based on the processing results.
[0124] By integrating sound source localization and beamforming algorithms into a low-power dedicated voice processing unit based on a multi-microphone array, this embodiment enables the system to not only monitor ambient sounds with extremely low standby power consumption, but also actively focus on and enhance target speech that may contain wake words, significantly reducing noise interference. This design improves the robustness of the system in accurately triggering wake-up in complex acoustic environments (such as the presence of background music, multiple conversations, or ambient noise), avoiding false or missed wake-ups, and achieving a combination of low-power monitoring and high-reliability wake-up.
[0125] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0126] The above description is only used to illustrate the technical solution of the present invention and is not intended to limit it. Any other modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention, as long as they do not depart from the spirit and scope of the technical solution of the present invention, should be covered within the scope of the claims of the present invention.
Claims
1. A voice wake-up method based on heterogeneous chip collaboration, characterized in that, include: In standby mode, the main processing unit enters sleep mode, and a dedicated voice processing unit with lower power consumption than the main processing unit runs independently to continuously collect and analyze audio signals. When the dedicated voice processing unit identifies a preset wake-up word in the audio signal, it sends a wake-up signal to the main processing unit. The main processing unit is activated in response to the wake-up signal and obtains subsequent voice commands from the dedicated voice processing unit. The main processing unit processes the subsequent voice commands and controls the execution of corresponding operations based on the processing results.
2. The voice wake-up method based on heterogeneous chip collaboration according to claim 1, characterized in that, The main processing unit processes the subsequent voice commands, including: Determine the complexity of the subsequent voice commands; If the instruction is determined to be a low-complexity local instruction, the main processing unit will invoke local resources to execute it. If the instruction is determined to be a highly complex cloud-based command, the main processing unit will send the subsequent voice instruction or its features to the cloud server via the wireless communication module and receive the response from the cloud server.
3. The voice wake-up method based on heterogeneous chip collaboration according to claim 1, characterized in that, After the main processing unit completes processing the subsequent voice commands, the method further includes: Start timeout counting; If no new valid instruction is received before the timeout reaches the preset threshold, the main processing unit saves the current context state and re-enters sleep mode, allowing the dedicated voice processing unit to resume independent operation, thereby restoring the entire system to the standby state.
4. The voice wake-up method based on heterogeneous chip collaboration according to claim 1, characterized in that, At the same time or after the dedicated voice processing unit sends the wake-up signal to the main processing unit, it also transmits audio data processed by the acoustic front end to the main processing unit, wherein the acoustic front end processing includes at least echo cancellation.
5. A voice wake-up system based on heterogeneous chip collaboration, characterized in that, It includes a main processing unit, a dedicated voice processing unit, an audio input unit, an audio output unit, and a power management unit; The power consumption of the dedicated voice processing unit is lower than that of the main processing unit; The audio input unit is electrically connected to the dedicated voice processing unit for collecting ambient sound. The dedicated voice processing unit is configured to work independently when the system is in standby mode, continuously analyze the audio signals collected by the audio input unit, and wake up the main processing unit when a wake-up word is recognized. The main processing unit is configured to: after being awakened from sleep mode, obtain voice commands from the dedicated voice processing unit and execute them; The audio output unit is controlled by the main processing unit; The power management unit provides power to the main processing unit and the dedicated voice processing unit, and supports independent control of the power supply to the main processing unit.
6. A voice wake-up system based on heterogeneous chip collaboration according to claim 5, characterized in that, The system also includes an audio amplification unit, the input of which is connected to the first audio output of the main processing unit, and the output of which is connected to the audio output unit; the feedback input of the audio amplification unit is connected to the reference signal output of the dedicated speech processing unit, forming an acoustic echo cancellation loop.
7. A voice wake-up system based on heterogeneous chip collaboration according to claim 5, characterized in that, The dedicated voice processing unit is connected to the main processing unit via at least one control signal line and one data bus; wherein, the control signal line is used to transmit the wake-up signal, and the data bus is used to transmit audio data or control commands.
8. A voice wake-up system based on heterogeneous chip collaboration according to claim 5, characterized in that, The power management unit includes: The charging management circuit connects the external power source and the internal battery. A controlled switch circuit is connected in series between the charging management circuit and the power input terminal of the main processing unit; The first control path is that the control terminal of the controlled switch circuit is connected to a general-purpose input / output pin of the main processing unit, which is used to perform software shutdown control on the main processing unit after it is powered on. The second control path includes a physical wake-up button connected to the control terminal of the controlled switch circuit, forming a hardware enable path in parallel with the first control path. This enable path is used to directly control the controlled switch circuit to power on the main processing unit when the system is completely powered off.
9. A voice wake-up system based on heterogeneous chip collaboration according to claim 5, characterized in that, The main processing unit is a microcontroller or application processor with integrated wireless communication function; the dedicated voice processing unit is a low-power voice chip that integrates an analog microphone interface, a digital signal processor and a wake-word recognition algorithm.
10. A voice wake-up system based on heterogeneous chip collaboration according to claim 5, characterized in that, The audio input unit is a multi-microphone array, and the dedicated voice processing unit also integrates a sound source localization or beamforming algorithm based on the multi-microphone array to improve the robustness of wake word recognition.