Processing method and electronic equipment
By introducing a collaborative processing unit architecture and an independent speech recognition module into electronic devices, the problem of ensuring voice interaction functionality while reducing power consumption is solved. Stable and comprehensive voice interaction is achieved under different conditions, providing a convenient and high-quality user experience, and improving security and privacy.
Patent Information
- Application Number
- CN202511786533.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-30
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies struggle to maintain stable voice interaction functionality in electronic devices while reducing power consumption.
The system adopts a collaborative processing unit architecture. The first processing unit performs preliminary voice interaction processing when not in operation, while the second processing unit performs complex voice interaction processing when in operation. By utilizing independent speech recognition modules and audio input devices, the system achieves a division of labor and cooperation between voice wake-up, simple command interaction, and complex voice processing.
While reducing power consumption, it ensures the stability and comprehensiveness of voice interaction functions of electronic devices in different states, provides a convenient and high-quality voice interaction experience, and improves security and privacy through voiceprint recognition and keyword wake-up.
Smart Images

Figure CN121600927A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly to a processing method and an electronic device. Background Technology
[0002] In the usage scenarios of electronic devices, users have high expectations for the interactive and audio experience, pursuing convenient and smooth interaction as well as high-quality audio effects. However, with current technology, it is difficult for electronic devices to stably guarantee voice interaction functionality while reducing power consumption. Summary of the Invention
[0003] The technical solution provided in this application is as follows:
[0004] The first aspect of this application provides a processing method, including:
[0005] Obtain the first parameter; the first parameter is used to identify the system status of the electronic device;
[0006] If the first parameter indicates that the system state of the electronic device is in a non-working state, the first processing unit of the electronic device performs voice interaction processing on the user based on the audio signal of the audio device, so that after the system state of the electronic device is switched to a working state, the second processing unit of the electronic device can perform voice interaction processing on the audio signal of the audio device; wherein, the first processing unit can identify the user based on the audio signal.
[0007] In one possible implementation, the audio device includes: an audio input device;
[0008] The first processing unit of the electronic device performs user voice interaction processing based on the audio signal from the audio device, including:
[0009] The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module to obtain a first recognition result, and performs a first operation based on the first recognition result;
[0010] The second processing unit of the electronic device is capable of performing voice interaction processing based on the audio signal from the audio device, including:
[0011] The second processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the second speech recognition module.
[0012] In one possible implementation, the audio input device includes: a first audio input device and a second audio input device; the first audio input device is connected to the first processing unit; the second audio input device is connected to the second processing unit; the first audio input device operates independently of the second audio input device.
[0013] The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module to obtain a first recognition result, including:
[0014] The first processing unit of the electronic device performs speech recognition based on the audio signal from the first audio input device using the first speech recognition module to obtain a first recognition result.
[0015] The second processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the second speech recognition module, including:
[0016] The second processing unit of the electronic device performs speech recognition based on the second speech recognition module, at least according to the audio signal from the second audio input device.
[0017] In one possible implementation, the audio device further includes: an audio output device;
[0018] The processing method further includes:
[0019] After the electronic device switches its system state to the working state, the first processing unit determines that the electronic device has entered the user login state and identifies whether the audio output device is in the playback state.
[0020] If so, the first processing unit stops performing speech recognition based on the audio signal from the first audio input device and sends the audio signal from the first audio input device to the second processing unit;
[0021] The second processing unit of the electronic device performs speech recognition based on the second speech recognition module, at least according to the audio signal from the second audio input device, including:
[0022] The second processing unit of the electronic device performs speech recognition based on the audio signals from the first audio input device and the second audio input device using the second speech recognition module.
[0023] In one possible implementation, the non-working state includes a power-off state; the step of the first processing unit of the electronic device performing speech recognition based on the audio signal from the audio input device using a first speech recognition module to obtain a first recognition result, and performing a first operation based on the first recognition result, includes at least one of the following:
[0024] Based on the first identification result, it is determined that the operating system needs to be woken up. The first processing unit of the electronic device sends a wake-up command to the control unit, so that the control unit responds to the wake-up command and triggers the electronic device to switch from the power-off state to the power-on state.
[0025] Based on the first identification result, it is determined that basic input / output system configuration needs to be performed. The first processing unit of the electronic device obtains the basic input / output system configuration instruction from the first identification result and sends the configuration instruction to the control unit, so that the control unit stores the configuration instruction and transmits the configuration instruction to the basic input / output system, so that the basic input / output system performs initialization settings based on the configuration instruction during power-on.
[0026] In one possible implementation, the audio device further includes: an audio output device;
[0027] The processing method further includes:
[0028] After the electronic device switches its system state to working state, it is determined that the electronic device has not entered the user login state. The first processing unit continues to perform voice recognition based on the audio signal of the audio input device by the first voice recognition module to obtain the second recognition result.
[0029] Based on the second recognition result, it is determined that the user has a need for voice interaction. The first processing unit generates voice feedback content based on the second recognition result and plays it through the audio output device.
[0030] In one possible implementation, the audio device includes: an audio input device;
[0031] The user's voice interaction processing performed by the first processing unit of the electronic device based on the audio signal from the audio device includes at least one of the following:
[0032] The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module, and determines the user's direction based on the first recognition result obtained from the speech recognition.
[0033] Based on the user's orientation, drive the display screen to rotate so that the rotated display screen faces the user;
[0034] During the rotation of the display screen, the indicator lights on the display screen are driven to flash to indicate the direction of rotation of the display screen.
[0035] In one possible implementation, the audio device includes: an audio output device; the audio output device includes: an audio power amplification unit connected to both the first processing unit and the second processing unit, and an audio playback unit connected to the audio power amplification unit;
[0036] The first processing unit of the electronic device performs user voice interaction processing based on the audio signal from the audio device, including:
[0037] The audio power amplification unit is controlled to enter a first mode; in the first mode, the audio power amplification unit is only allowed to process the audio feedback signal output by the first processing unit;
[0038] After the electronic device switches from the non-working state to the working state, the audio power amplification unit is controlled to switch from the first mode to the second mode; in the second mode, the audio power amplification unit is allowed to process the audio feedback signals output by the first processing unit and the second processing unit.
[0039] In one possible implementation, the audio device includes: an audio input device;
[0040] The processing method further includes:
[0041] The ambient sound collected by the audio input device is detected by the voice activity detection unit inside the first processing unit;
[0042] If the voice activity detection unit does not detect a valid sound signal, then the other functional units in the first processing unit, excluding the voice activity detection unit, are put into a sleep state.
[0043] If the voice activity detection unit detects a valid sound signal, it switches the other functional units from the sleep state to the working state.
[0044] Another aspect of this application provides an electronic device, comprising:
[0045] A first audio input device is connected to a first processing unit;
[0046] A second audio input device is connected to a second processing unit; the first audio input device operates independently of the second audio input device; the power consumption of the first processing unit is lower than that of the second processing unit.
[0047] The first processing unit is configured to:
[0048] If the first parameter represents the system state of the electronic device as being in a non-working state, the user's voice interaction processing is performed based on the audio signal from the audio device.
[0049] The second processing unit is used to perform voice interaction processing based on the audio signal from the audio device after the system state of the electronic device is switched to the working state; wherein, the first processing unit is able to identify the user based on the audio signal. Attached Figure Description
[0050] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0051] Figure 1 A flowchart illustrating a processing method provided in Embodiment 1 of this application;
[0052] Figure 2 A schematic diagram of the structure of an electronic device provided in this application;
[0053] Figure 3 This is a flowchart illustrating a processing method provided in Embodiment 4 of this application;
[0054] Figure 4 A connection diagram of an audio input device provided in this application;
[0055] Figure 5 A schematic diagram of a BIOS configuration scenario provided in this application;
[0056] Figure 6 This application provides a schematic diagram of a scenario where a user is not logged in.
[0057] Figure 7 This application provides a schematic diagram of a scenario for a non-working state.
[0058] Figure 8 A schematic diagram of a screen control-related structure provided in this application;
[0059] Figure 9 This application provides a schematic diagram of a working state scenario.
[0060] Figure 10 Another structural schematic diagram of an electronic device provided in this application;
[0061] Figure 11 A schematic diagram of the structure of a first processing unit provided in this application;
[0062] Figure 12 This is a functional diagram of an AI chip provided in this application. Detailed Implementation
[0063] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0064] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0065] The terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0066] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0067] Reference Figure 1 This is a flowchart illustrating a processing method provided in Embodiment 1 of this application, as shown below. Figure 1 As shown, the method may include the following steps:
[0068] Step S101: Obtain the first parameter; the first parameter is used to identify the system state of the electronic device.
[0069] The first processing unit can obtain the first parameter from the embedded controller (EC).
[0070] There are two possible ways for the EC to obtain the first parameter. On the one hand, it can obtain it from the Basic Input / Output System (BIOS) via eSPI. As the firmware program that runs first when an electronic device boots up, the BIOS stores a large amount of basic information about the device's hardware configuration and initial state. The EC can extract data related to the system state from this information as the first parameter.
[0071] On the other hand, the EC can also determine the first parameter on its own. As one of the core control units of an electronic device, the EC can monitor the operation of various hardware components in the device in real time, such as power status, connection and working status of hardware devices, etc. By analyzing and judging this real-time data, the EC can determine the current system status of the electronic device on its own and use it as the first parameter.
[0072] Step S102: If the first parameter represents the system state of the electronic device as being in a non-working state, the first processing unit of the electronic device performs voice interaction processing on the user based on the audio signal of the audio device, so that after the system state of the electronic device is switched to a working state, the second processing unit of the electronic device can perform voice interaction processing on the audio signal of the audio device; wherein, the first processing unit can identify the user based on the audio signal.
[0073] Non-working states can include, but are not limited to, MS, S4, or S5.
[0074] MS can represent a specific low-power mode, in which most of the hardware of an electronic device operates at low power or even stops working, retaining only some basic functions to maintain a minimum level of operation.
[0075] S4 can represent a hibernation state, in which the electronic device saves the currently running data to the hard drive, then shuts down most of the hardware devices and enters a low-power standby state, which can be quickly restored to the previous working state.
[0076] S5 indicates the power-off state. In the power-off state, the main processing unit (SOC, System-on-a-Chip) and core supporting components inside the electronic device will completely stop operating. Specifically, this includes the main CPU, memory, discrete graphics card / integrated graphics, hard drive, etc. integrated in the SoC. The main power rail (+12V / +5V / +3.3V) is cut off, and only the +5VSB standby voltage is retained to maintain the basic circuits. The entire device is in an extremely low-power standby mode.
[0077] The first processing unit can integrate a voice activity detection unit (e.g., VAD), a digital signal processing unit (e.g., DSP), a dedicated CPU (different from the main CPU in a SOC that is responsible for complex calculations) and a local storage unit. The collaborative work of these components enables the first processing unit to have voice processing capabilities.
[0078] The voice activity detection unit can monitor audio signals in real time and accurately determine whether there is voice input. When a voice signal is detected, it quickly triggers subsequent processing flows, avoiding unnecessary computation and power consumption waste.
[0079] The digital signal processing unit can perform fast noise reduction, filtering and other preprocessing operations on audio signals to improve the quality of speech signals and provide clearer and more accurate input for subsequent speech recognition and interactive processing.
[0080] The local storage unit can store some commonly used speech models and parameters, so that the first processing unit does not need to frequently interact with external storage devices when processing speech, thereby further improving processing efficiency and reducing power consumption.
[0081] As the core computing component of the first processing unit, the dedicated CPU is responsible for coordinating the workflow between various units. For example, the dedicated CPU can receive signals from the speech activity detection unit, determine that there is speech input, and then instruct the digital signal processing unit to preprocess the audio signal. Afterward, the dedicated CPU can call upon commonly used speech models and parameters stored in the local storage unit to perform in-depth analysis and processing on the preprocessed speech signal.
[0082] Furthermore, the first processing unit can support low-voltage (e.g., 1.8V) power supply, which results in extremely low power consumption during operation. Moreover, it can operate independently when core components such as the SOC (System-on-a-Chip) and memory are in sleep or powered off, without relying on the startup of other high-power components, thus enabling it to independently complete preliminary voice interaction processing tasks in non-working states.
[0083] When the first parameter characterizes the system state of the electronic device in these non-working states, the first processing unit can perform user voice interaction processing based on the audio signals collected by the audio device (such as a microphone).
[0084] Next, using a specific scenario, we will explain how the first processing unit processes user voice interactions based on the audio signals collected by the audio device when it is not in operation.
[0085] For example, in a voice wake-up scenario, when a user issues a voice wake-up command, the voice activity detection unit of the first processing unit first detects the voice signal and transmits it to a dedicated CPU. After determining that there is voice input, the dedicated CPU instructs the digital signal processing unit to preprocess the audio signal and extract voice features. Then, the dedicated CPU calls upon the voice model and algorithm stored in the local storage unit to recognize and analyze the voice features, determining whether it is a valid wake-up command. If it is a valid wake-up command, the dedicated CPU sends a wake-up signal to the power management module of the electronic device, causing the device to switch to working mode.
[0086] In a battery level query scenario, when an electronic device is in a non-working state (such as sleep mode), the user may want to know the remaining battery level. In this case, the user issues a voice command such as "How much battery does the device have left?" The voice activity detection unit of the first processing unit detects the voice signal and transmits it to a dedicated CPU. The dedicated CPU triggers a digital signal processing unit to perform preprocessing, and then uses a locally stored voice model to recognize that the user's intention is a battery level query. Subsequently, the dedicated CPU communicates with the power management module to obtain the current battery level information, and uses voice synthesis technology (if the first processing unit has simple synthesis capabilities or collaborates with other synthesis modules within the device) to feed back the battery information to the user, such as "The device currently has 60% battery remaining."
[0087] In a network connectivity status query scenario, if the electronic device is not in operation but is connected to the network (e.g., a laptop maintains a Wi-Fi connection even when in sleep mode), the user issues a voice command asking "Is the device connected to the network?" The voice activity detection unit of the first processing unit detects the voice signal and transmits it to a dedicated CPU. After recognizing the command, the dedicated CPU communicates with the device's network module to obtain network connectivity status information, and then provides feedback to the user through voice synthesis or other methods, such as "The device is connected to Wi-Fi, and the signal strength is good."
[0088] In a volume setting scenario, when the device is not in operation, if the user wants to pre-set the volume level after powering on, they issue a voice command saying "Set the volume to 50%". The voice activity detection unit of the first processing unit detects the voice signal and transmits it to the dedicated CPU. After recognizing the command, the dedicated CPU records the volume setting information. Once the device is woken up and enters operating mode, it sends a command to the device's volume control module to set the volume to the user-specified 50%.
[0089] In scenarios involving switching working modes, some devices with multiple working modes, such as smart cameras with monitoring and privacy modes, can switch to privacy mode when the device is not in operation (e.g., in sleep mode). The voice activity detection unit of the first processing unit detects the voice signal and transmits it to a dedicated CPU. The dedicated CPU recognizes the voice signal, records the mode switching command, and sends it to the camera's control module when the device wakes up, controlling the camera to switch to privacy mode, turn off the camera lens, or perform physical blocking operations.
[0090] Once the electronic device's system state switches to active mode, the second processing unit will take over the voice interaction processing tasks. This second processing unit may include the electronic device's main processor, such as the main CPU, which possesses more powerful computing capabilities and richer functions, enabling more complex and comprehensive voice interaction processing.
[0091] For example, after the first processing unit wakes up the device, the user may issue a command to open a specific application. At this point, the second processing unit can use more advanced natural language processing algorithms to understand the user's specific command based on the audio signal collected by the audio device. It will then interact with the operating system and other applications to perform corresponding operations, such as launching a browser or opening a document.
[0092] In this embodiment, when the electronic device is in a non-working state, the first processing unit can operate independently. Its integrated high-efficiency components can directly perform preliminary voice interaction processing on the audio signal, such as recognizing voice wake-up commands and simple command interactions, without having to wake up high-power components such as the main processor. While realizing voice interaction functions in a non-working state, it effectively reduces power consumption.
[0093] When electronic devices switch to working mode, and are faced with more complex and diverse voice commands from users, such as opening specific applications or performing complex queries, the second processing unit can take over the task with its powerful computing capabilities and rich functions. It can accurately understand and execute complex user commands, ensuring the stability and comprehensiveness of voice interaction functions.
[0094] This division of labor and collaboration achieves complementary advantages, taking into account both the need for low power consumption and ensuring the stability and comprehensiveness of voice interaction functions under different conditions.
[0095] As another optional embodiment of this application, which is a processing method provided in Embodiment 2 of this application, this embodiment is mainly an implementation of the above step S102. In this embodiment, the audio device may include: an audio input device.
[0096] Audio input devices can capture the voice signals emitted by users, providing raw data for subsequent voice interaction processing. Whether the electronic device is in a non-operating or operating state, the audio input device remains continuously capable of receiving voice signals, ensuring that any voice commands issued by the user can be captured.
[0097] The first processing unit of the electronic device performing user voice interaction processing based on the audio signal from the audio device may include:
[0098] Step S11: The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module to obtain a first recognition result, and performs a first operation based on the first recognition result.
[0099] The first speech recognition module may, but is not limited to, having voiceprint recognition (VOICE ID) and keyword wake-up functions.
[0100] Voiceprint recognition is a function that identifies a speaker by analyzing the unique biometric features in a speech signal. Like a fingerprint, each person's voiceprint is unique, determined by factors such as the physiological structure of the vocal organs and pronunciation habits. The voiceprint recognition function in the first speech recognition module can perform in-depth analysis of the collected speech when the electronic device is not in operation, extracting the speaker's voiceprint features and comparing them with pre-stored legitimate user voiceprint models.
[0101] Users can record their own voice samples according to system prompts. The first voice recognition module can extract and analyze the features of these samples, construct a unique voiceprint model for each user, and store it. When subsequent voice input occurs, the first voice recognition module can quickly extract the voiceprint features of the new voice and accurately match it with the stored model. If the matching degree reaches a preset threshold, it can be determined that the voice was spoken by a legitimate user; if the matching degree is low, it may be determined as an unauthorized user or environmental noise, thus effectively preventing others from waking up the device or performing operations without authorization, ensuring the security and privacy of device use.
[0102] For example, in a home setting, only electronic devices with pre-recorded voiceprints by family members can accurately recognize and respond to voice commands issued by family members; however, when a stranger tries to wake up the device, it will not respond due to the mismatch in voiceprints, thus avoiding the risk of the device being misoperated.
[0103] Keyword wake-up functionality focuses on recognizing specific wake words or phrases to activate electronic devices from a non-working state. These wake words can be simple, easy-to-remember words or phrases with specific meanings, such as "Xiao X Assistant" or "Wake up the device."
[0104] The first speech recognition module can store a speech model, which can recognize different pronunciations, intonations, speech rates, and wake words under various environmental noise conditions. When the audio input device (such as a microphone) collects the user's speech signal, the first speech recognition module can first use the voiceprint recognition function to determine whether the user is legitimate. If the user is legitimate, the module will further perform keyword wake-up recognition on the speech signal.
[0105] This method, combining voiceprint recognition and keyword wake-up, ensures both the accuracy and security of device wake-up and improves the success rate of recognition in complex environments. For example, in noisy public places, even with high ambient noise, the device can still accurately identify and wake up the user as long as the legitimate user utters the correct wake-up word, providing convenient service.
[0106] Once the first voice recognition module completes voiceprint recognition and keyword wake-up recognition and obtains the first recognition result, the first processing unit can quickly execute the corresponding first operation based on the result to achieve the initial interaction between the user and the device.
[0107] For example, in a voice wake-up scenario, if the first voice recognition module determines that a legitimate user has issued a valid wake-up command, the first processing unit can immediately send a wake-up signal to the power management module of the electronic device. After receiving the signal, the power management module will gradually wake up the various hardware components in the device according to a preset process.
[0108] For example, the power management module can first restore power to the main processor (SOC), bringing it back from a low-power or hibernation state to normal operation. Then, it can sequentially wake up core components such as memory, hard drive, and graphics card, and load necessary system programs and drivers. Once the electronic device is fully awake, the user can issue more complex voice commands, which will be processed and responded to by the second processing unit.
[0109] Meanwhile, since the user's identity has been confirmed through voiceprint recognition during the wake-up process, electronic devices can automatically load personalized user interfaces and applications based on different users' preferences and settings, providing a more personalized user experience. For example, for users who frequently use the music playback function, the device can directly open the music playback application and play the song the user last listened to after waking up.
[0110] In simple command interaction scenarios, after the first voice recognition module recognizes a simple command issued by a legitimate user, the first processing unit can execute the corresponding operation based on the command type. These simple commands may include queries for basic device information or simple settings operations, such as "check battery level," "check network connection status," or "turn the volume up to 50%."
[0111] The second processing unit of the electronic device is capable of performing voice interaction processing based on the audio signal from the audio device, and may include:
[0112] Step S21: The second processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the second speech recognition module.
[0113] The second speech recognition module is somewhat similar to the first speech recognition module in terms of its functional implementation logic. Both of them have the two core functions of user authentication and wake word recognition.
[0114] In terms of user authentication, both methods can extract user voiceprints by analyzing biometric features in speech signals. Specifically, they delve into key biometric information such as spectral structure, formant patterns, and pronunciation habits in the speech signal to extract a voiceprint that represents the user's uniqueness. The extracted voiceprint is then meticulously compared with a pre-stored legitimate user voiceprint model in the system. If the matching degree reaches a preset accuracy threshold, the voice is determined to be from a legitimate user; if the matching degree is low, it may be determined to be from an unauthorized user or due to environmental noise interference. This effectively prevents unauthorized users from waking up the device or performing operations, ensuring device security and user privacy.
[0115] Regarding wake-up word recognition, both modules support the recognition of preset wake-up words. The first voice recognition module can recognize specific wake-up words or phrases such as "Xiao X Assistant" or "Wake up the device" to activate the electronic device from a non-working state to a working state. The second voice recognition module also supports the recognition of preset wake-up words (such as "Xiao X Assistant"), thereby triggering device state switching or function module activation.
[0116] However, differences also exist between the two. The first and second processing units may differ significantly in computing power, with the second unit typically possessing far greater computational capabilities. Based on this difference in computing power, the functions supported by the second speech recognition module are more complex and in-depth than those of the first. For example, during speech recognition, the second speech recognition module can handle more complex and varied speech scenarios, achieving more accurate recognition and analysis of speech signals with different accents, speaking speeds, intonations, and under complex environmental noise. In natural language processing, the second speech recognition module can employ more advanced algorithms to deeply understand the complex semantics behind user voice commands, thereby executing corresponding operations more accurately and providing users with a more comprehensive and higher-quality voice interaction experience.
[0117] In this embodiment, the first processing unit and the second processing unit are respectively equipped with a first speech recognition module and a second speech recognition module, with clearly defined functions. When the electronic device is in a non-working state, the first processing unit uses the first speech recognition module to perform preliminary voice interaction processing, such as recognizing voice wake-up commands and simple command interactions, quickly responding and executing corresponding operations, enabling the device to switch from a non-working state to a working state or complete simple tasks in a timely manner. When the device switches to a working state, the second processing unit, relying on the second speech recognition module and its more powerful computing capabilities, processes more complex and varied voice scenarios, deeply understands the complex semantics behind the user's voice commands, executes more precise operations, and provides the user with a more comprehensive and high-quality voice interaction experience. This division of labor and cooperation fully leverages the advantages of the two processing units, achieving complementary strengths and improving the overall voice interaction capabilities of the electronic device in different states.
[0118] Furthermore, when the first speech recognition module in the first processing unit processes the user's voice signal, all of the user's audio content is only inferred within the first processing unit (e.g., an AI chip), and is not uploaded to the operating system (OS) or the network. This means that the user's voice information will not be leaked to the external environment before the user utters the wake word, effectively avoiding the privacy risks that may arise from data uploading.
[0119] As another optional embodiment of this application, this is a processing method provided in Embodiment 3 of this application. This embodiment is mainly an implementation of the above-mentioned step S11. In this embodiment, as follows... Figure 2 As shown, the audio input device may include: a first audio input device and a second audio input device.
[0120] The first audio input device is connected to the first processing unit (e.g., an AI chip).
[0121] The second audio input device is connected to the second processing unit (e.g., SOC).
[0122] The first audio input device operates independently of the second audio input device, without interfering with each other, ensuring accurate acquisition of voice signals under different conditions.
[0123] Both the first and second audio input devices can be configured with microphone arrays. The microphone array can be flexibly set to include two or more microphones according to actual needs.
[0124] In addition, there can be multiple ways to connect the AI chip and the SOC, such as transmitting audio signals through the PDM (Pulse Density Modulation) interface or transmitting commands through the USB interface, so as to perform data interaction and collaborative work when necessary.
[0125] The first audio input device remains continuously active, capable of receiving voice signals, both when the electronic device is in a non-operating state (including sleep, power-off, or specific low-power modes) and when it is in operation. Its core responsibility is to capture user voice commands at all times, ensuring that user voice input is promptly perceived regardless of the device's state.
[0126] The direct connection between the first audio input device and the first processing unit allows the audio signal acquired by the first audio input device to be transmitted to the first processing unit at extremely high speed and stably, minimizing interference and delays that the signal may encounter during transmission. This feature provides the first processing unit with high-quality raw data for initial voice interaction processing, enabling the device to quickly respond to simple voice commands, such as waking up the device or querying basic information, even when not in operation.
[0127] The second audio input device also maintains its voice signal reception capability whether the electronic device is in operation or not. When the electronic device is in operation, users often issue more complex and diverse voice commands. With its high-precision acquisition capabilities, the second audio input device can accurately capture these commands, providing reliable raw data for subsequent processing.
[0128] The second audio input device is closely connected to the second processing unit. The second processing unit possesses powerful computing capabilities and rich functionality. Leveraging this advantage, the voice signals acquired by the second audio input device can achieve more advanced voice interaction processing. For example, it can perform complex natural language understanding, accurately grasp user intent, support multi-turn dialogue, and achieve a coherent and natural interaction process, bringing users an intelligent, convenient, and immersive voice interaction experience.
[0129] Step S11 may include, but is not limited to, the following steps:
[0130] Step S111: The first processing unit of the electronic device performs speech recognition based on the audio signal from the first audio input device using the first speech recognition module to obtain a first recognition result.
[0131] When the electronic device is in a non-operating state (such as sleep, power off, or a specific low-power mode), the first audio input device continuously collects voice signals from the surrounding environment.
[0132] The voice activity detection unit of the first processing unit monitors the audio signal collected by the first audio input device in real time to determine whether there is voice input. When a voice signal is detected, subsequent processing can be triggered.
[0133] A dedicated CPU receives signals from the voice activity detection unit. Once it determines that there is voice input, it can instruct the digital signal processing unit to preprocess the audio signal, such as noise reduction and filtering, to improve the quality of the voice signal.
[0134] A dedicated CPU can access commonly used speech models and parameters stored in the local storage unit to perform in-depth analysis and processing of the pre-processed speech signal. First, it can use voiceprint recognition to determine if the user is legitimate. If the user is legitimate, it can further perform keyword wake-up recognition on the speech signal.
[0135] If a valid wake-up command is determined to be issued by a legitimate user, the first processing unit can send a wake-up signal to the power management module of the electronic device. The power management module then gradually wakes up the various hardware components in the device according to a preset procedure, causing the device to switch to the working state.
[0136] If a simple command issued by a legitimate user is identified, such as checking battery level, network connection status, or adjusting volume, the first processing unit executes the corresponding operation based on the command type. After the device is woken up, it can transmit relevant information to the second processing unit as needed, so that the second processing unit can perform further processing or recording after the device is fully started.
[0137] Step S21 may include, but is not limited to, the following steps:
[0138] Step S211: The second processing unit of the electronic device performs speech recognition based on the second speech recognition module, at least according to the audio signal of the second audio input device.
[0139] Once the electronic device switches to operating mode, the second processing unit can perform speech recognition on the audio signal collected by the second audio input device based on the second speech recognition module. The second speech recognition module can use more advanced algorithms to accurately identify and analyze speech signals with different accents, speaking speeds, intonations, and under complex environmental noise, thereby gaining a deeper understanding of the complex semantics behind the user's voice commands.
[0140] After successfully completing speech recognition and accurately grasping the complex semantics of the user's voice commands, the second processing unit can interact with the operating system (OS) and other applications (APP) based on the recognized user commands to perform corresponding operations, such as launching a browser, opening a document, and performing complex queries, providing users with a comprehensive and high-quality voice interaction experience.
[0141] In this embodiment, the design of the dual audio input device enables the use of dedicated audio acquisition equipment in different states, reducing signal interference and transmission delay, and providing a clearer and more accurate audio signal for the speech recognition module.
[0142] Through this processing method, electronic devices can respond quickly and accurately to user voice commands in different states. Whether it is a simple wake-up and basic query, or a complex application operation and multi-turn dialogue, it can provide users with a smooth and convenient voice interaction experience, meeting users' high expectations for the interaction and audio experience of electronic devices.
[0143] As another optional embodiment of this application, refer to Figure 3 This is a flowchart illustrating a processing method provided in Embodiment 4 of this application. In this embodiment, the audio device may include: a first audio input device, a second audio input device, and an audio output device. Figure 3 As shown, the method may include, but is not limited to, the following steps:
[0144] Step S201: Obtain the first parameter; the first parameter is used to identify the system state of the electronic device.
[0145] For a detailed description of step S201, please refer to the relevant description of step S101 in Example 1, which will not be repeated here.
[0146] Step S202: If the first parameter represents the system state of the electronic device as being in a non-working state, the first processing unit of the electronic device performs speech recognition based on the audio signal from the first audio input device using the first speech recognition module to obtain a first recognition result.
[0147] For a detailed description of step S202, please refer to the relevant description of step S111 in the above embodiments, which will not be repeated here.
[0148] Step S203: After the system state of the electronic device is switched to the working state, the first processing unit determines that the electronic device has entered the user login state and identifies whether the audio output device is in the playback state.
[0149] When the audio output device is in playback mode, the played audio will be picked up again by the first audio input device through air propagation (such as sound waves spreading in space and being received by the microphone) or internal coupling of the device (such as circuit signal interference, mechanical vibration transmission, etc.), forming an echo. This echo, when mixed with the user's actual speech signal, will seriously interfere with speech recognition.
[0150] Considering power consumption and overall cost, the first processing unit is typically designed for relatively simple, basic speech processing tasks, and its hardware configuration and computing power are relatively weak. Echo cancellation is a complex process that requires real-time analysis, modeling, and filtering of the audio signal, involving numerous mathematical operations and signal processing algorithms, such as adaptive filtering algorithms. The computing power of the first processing unit is insufficient to support the efficient operation of these complex algorithms, and it cannot accurately separate and cancel echoes in a short time.
[0151] During playback, the first processing unit continuously performs speech recognition on the signal from the first audio input device. This requires the operation of components such as speech activity detection, digital signal processing, and a dedicated CPU, which consumes additional power. However, due to echo interference, the recognition effect is difficult to guarantee. Continuing to run the system not only fails to obtain accurate results but also wastes power.
[0152] Based on this, the first processing unit can determine whether an electronic device is playing audio content after entering the user login state by monitoring relevant status signals or parameters of the audio output device. For example, it can determine this by checking the power status of the audio output device, the signal status of the audio signal transmission line, etc.; or it can obtain volume information through the audio driver. The audio driver is the bridge between the operating system and the audio hardware, responsible for managing various operations and parameter settings of the audio device. The first processing unit can interact with the audio driver, call relevant interface functions, and obtain the current volume setting value of the audio output device. If the volume value is not zero, it means that the user may have set the playback volume, and the audio output device may be playing audio; further, by combining the volume change situation, such as whether the volume remains stable or shows dynamic changes over a period of time, it can more accurately determine whether it is actually playing audio. If the volume value is zero, it can be basically determined that the audio output device is not playing audio.
[0153] The first processing unit can determine whether to enter the user login state through, but is not limited to, any of the following methods:
[0154] Electronic devices incorporate dedicated hardware circuitry to transmit specific electrical signals to the first processing unit when a user's login status changes. For example, upon successful login, the system triggers a voltage level change signal, which is transmitted to the first processing unit via the hardware circuitry. The first processing unit detects the voltage level on this circuitry (e.g., a high level indicates login, a low level indicates no login) to determine whether the device has entered the user login state.
[0155] The system control module of the electronic device can be connected to the first processing unit via a GPIO interface. When the user's login status changes, the system control module controls the corresponding GPIO pin to output a specific logic level. The first processing unit periodically reads the level value of the GPIO pin and determines whether the device has entered the user login state according to preset logic rules (e.g., high level represents logged in, low level represents not logged in).
[0156] The first processing unit can obtain the current user login status by calling operating system interfaces. For example, in the Windows operating system, the Windows API can be used to obtain the currently active console session ID, and then other related functions can be used to further determine whether the user corresponding to that session has successfully logged in.
[0157] If so, proceed to step S204.
[0158] Step S204: The first processing unit stops performing speech recognition based on the audio signal from the first audio input device and sends the audio signal from the first audio input device to the second processing unit, so that the second processing unit of the electronic device performs speech recognition based on the audio signal from the first audio input device and the audio signal from the second audio input device using the second speech recognition module.
[0159] Different audio input devices may be located in different positions, and the audio signals they acquire will have different characteristics and angular information. Therefore, the first processing unit can send the audio signal from the first audio input device to the second processing unit. The audio signals acquired by the first and second audio input devices complement each other, providing richer and more comprehensive speech information. For example, one audio input device may be closer to the user's mouth, acquiring a clearer and more direct speech signal; another audio input device may be acquiring from a different direction, capturing some obscured or reflected speech components. The combination of the two can more completely reconstruct the user's speech content.
[0160] The second processing unit of the electronic device performs speech recognition based on the audio signals of the first audio input device and the audio signals of the second audio input device using the second speech recognition module, which is one embodiment of step S211 above.
[0161] The secondary processing unit is often equipped with a more powerful processor, such as a high-performance digital signal processor (DSP) or a dedicated artificial intelligence chip (such as an NPU). These processors have higher computing speeds and parallel processing capabilities, enabling them to quickly process large amounts of audio data and run complex echo cancellation algorithms. For example, the secondary processing unit in some high-end smart devices can process signals from multiple audio channels in real time, simultaneously performing various speech enhancement processes such as echo cancellation and noise suppression to ensure the accuracy of speech recognition.
[0162] In this embodiment, the first processing unit supports simultaneously providing audio signals from the first audio input device to both the first processing unit and the second processing unit. For example, such as... Figure 4 As shown, the DSP of the first processing unit can be connected to the first audio input device and the second processing unit (e.g., SOC). This connection method allows the first processing unit to support three modes.
[0163] The first mode can be used internally by the first processing unit (e.g., an AI chip). For example, in the non-working state, the DSP of the first processing unit sends the audio signal from the first audio input device to the CPU of the first processing unit for speech recognition to obtain a first recognition result.
[0164] The second mode involves feeding the audio signal from the first audio input device to the second processing unit, while the first processing unit ceases to be used. For example, if the audio output device is detected to be in playback mode, the audio signal from the first audio input device can be fed to the second processing unit, which then performs corresponding processing using the echo cancellation function.
[0165] The third mode allows the audio signal from the first audio input device to be used by the second processing unit, while the first processing unit performs processing without affecting the application of the second processing unit. For example, in operation, if a user inputs a voice query for the weather, the first processing unit can send the audio signal from the first audio input device to the second processing unit while simultaneously providing a simple response based on the audio signal, which is then processed by the second processing unit for full processing.
[0166] In this embodiment, when the electronic device is in a non-working state, the first processing unit performs preliminary voice interaction processing to reduce power consumption. After switching to the working state, when it is determined that the user has entered the login state and the audio output device is playing, the first processing unit stops recognizing and sends the audio signal to the second processing unit, which has powerful computing power and echo cancellation capabilities. The dual audio input devices provide comprehensive voice information, achieving complementary advantages, reducing power consumption, and ensuring stable, accurate and comprehensive voice interaction functions, thus meeting users' high expectations for the interaction and audio experience of electronic devices.
[0167] As another optional embodiment of this application, this embodiment provides a processing method for embodiment 5 of this application. This embodiment is mainly an implementation of step S11 in embodiment 2. In this embodiment, the non-working state may include the power-off state, and step S11 may include, but is not limited to, at least one of the following:
[0168] Step S112: Based on the first identification result, it is determined that the operating system needs to be woken up. The first processing unit of the electronic device sends a wake-up command to the control unit, so that the control unit responds to the wake-up command and triggers the electronic device to switch from the power-off state to the power-on state.
[0169] When an electronic device is powered off, there are several possible scenarios that might trigger the operation to wake up the operating system. Besides explicitly saying "power on" or "wake up the device" in voice commands, users might also issue voice commands to use specific functions of an application (app) on the device. In these cases, the user essentially has a latent need to wake up the operating system. For example, if a user wants to play a song using a music player app, even without directly saying anything related to powering on, issuing a command like "play music" implies a desire to restore the device from its powered-off state so that the desired function can be used.
[0170] The first audio input device (such as a microphone) is continuously in a state capable of receiving voice signals and will collect voice signals from the surrounding environment in real time. The first processing unit performs voice recognition processing on the audio signals collected by the first audio input device based on the first voice recognition module to obtain a first recognition result.
[0171] The first voice recognition module performs two key recognition operations. First, voiceprint recognition analyzes the voiceprint features of the collected voice signal and compares them with pre-stored voiceprint information of authorized users to determine if the voice command was issued by an authorized user. This ensures the security of device operation and prevents unauthorized misoperation or malicious operation. Second, keyword wake-up recognition matches the voice signal with a preset wake-up word library to determine if the voice contains a preset wake-up word. These wake-up words include not only direct instructions like "power on" or "wake up the device," but also a series of potential command words related to using application functions that may trigger power-on.
[0172] If the user is identified as legitimate through voiceprint recognition, and the keyword wake-up recognition determines that the first recognition result is an instruction to wake up the operating system (whether it is a direct power-on instruction or a wake-up request instruction indirectly generated by using application functions), then the first processing unit will generate a corresponding wake-up instruction based on the recognition result and send the wake-up instruction to the control unit (EC) through a specific communication interface (such as a low pin count (LPC) interface).
[0173] After receiving a wake-up command, the EC can trigger the electronic device to switch from a powered-off state to a powered-on state. Specifically, the EC will control the power management module to gradually supply power to core components such as the main processing unit (SOC), memory, and hard drive in a certain order and rhythm, restoring them from a power-off state to a normal working state, and then starting the operating system, allowing the electronic device to enter an operable working mode.
[0174] For example, if a user suddenly has a specific need at a certain moment, such as wanting to listen to their favorite music to relax or start the day off right, when the user issues a natural voice command like "play my favorite music", the first audio input device (such as a microphone) that has been in standby mode will quickly collect the audio signal corresponding to the command and immediately transmit it to the first processing unit.
[0175] Upon receiving the audio signal, the first processing unit quickly begins operation using the first speech recognition module. First, it performs voiceprint recognition, using sophisticated algorithms to analyze the voiceprint features in the audio signal and meticulously compares them with pre-stored legitimate user voiceprint information to confirm that the voice command was issued by a device-approved, legitimate user, ensuring the security and legitimacy of the operation. After voiceprint recognition, it proceeds with keyword wake-up recognition. This step goes beyond simply recognizing straightforward commands like "power on" or "wake up," possessing powerful semantic understanding capabilities to deeply analyze the underlying intent of the voice command. For example, in this case, the user's command to "play my favorite music" accurately indicates that the command implicitly requires waking up the operating system to use the music playback function to satisfy the user's music listening needs.
[0176] Once it is determined that the operating system needs to be woken up, the first processing unit can generate a corresponding wake-up command and send the wake-up command to the control unit (EC) quickly and accurately through a specific communication interface (such as a low pin count (LPC) interface).
[0177] Upon receiving a wake-up command, the EC triggers the electronic device to switch from its current state to a power-on state. As the core components are gradually activated, the operating system also starts up, and the electronic device gradually enters an operable working mode.
[0178] Subsequently, the music playback app can automatically launch based on the user's command, quickly locate the user's favorite music list, and begin smoothly playing the user's favorite music. Throughout the entire process, the user does not need to manually operate the power button, or even consciously pay attention to whether the device is off. Simply issuing a simple voice command allows for easy activation and enjoyment of music, truly achieving a seamless wake-up process and enabling electronic devices to better serve daily life.
[0179] Step S113: Based on the first identification result, it is determined that the Basic Input / Output System (BIOS) configuration needs to be performed. The first processing unit of the electronic device obtains the BIOS configuration instruction from the first identification result and sends the configuration instruction to the control unit, so that the control unit stores the configuration instruction and transmits the configuration instruction to the BIOS, so that the BIOS performs initialization settings based on the configuration instruction during power-on.
[0180] When the electronic device is in the S5 (power off) state, the user emits a voice signal containing a specific intention, such as expressing a request to "set the graphics card to dedicated graphics mode" or "set the USB drive as the first boot device." The first audio input device (such as a microphone) picks up the signal and transmits it to the first processing unit (such as an AI chip).
[0181] The AI chip can infer the corresponding BIOS configuration commands (such as "discrete graphics mode," "USB boot," etc.) and the system wake-up command, and quickly transmit the BIOS configuration commands and wake-up commands to the EC (Engineer Control Center). The EC temporarily stores the received BIOS configuration commands. At this stage, the computer is not yet powered on, but the EC has successfully saved the user's new BIOS configuration requirements, laying the foundation for subsequent rapid configuration.
[0182] Upon receiving the wake-up command, the EC immediately triggers the computer's power-on process. Simultaneously, the BIOS reads the BIOS configuration instructions from the EC. It's important to note that the EC does not transmit the entire BIOS configuration, but rather the portion that the user needs to modify based on specific requirements. The computer's basic configuration is still obtained using the existing loading mechanism. During this power-on process, the BIOS can directly read the new configuration information transmitted by the EC and quickly complete the loading.
[0183] Compared to manually configuring the BIOS after powering on the computer and entering the BIOS interface (when the computer powers on via the power button, it performs an initial power-on self-test, during which the basic BIOS configuration is already loaded. New configurations modified by the user at this point, such as graphics card mode and boot device parameters, cannot directly overwrite the already loaded configuration. To make the new configuration take effect, the user must save the settings (e.g., press F10), then restart the computer to power it on again and reread the newly saved BIOS configuration, thus replacing the original configuration), this embodiment eliminates the need to restart the device, greatly improving configuration efficiency.
[0184] Next, taking the user's input of "set the graphics card mode to dedicated graphics mode" as an example, we will explain in detail the complete process of electronic devices from voice command input to completing BIOS configuration and entering the operating system.
[0185] For example, such as Figure 5 As shown, when the electronic device is powered off, the user utters the voice command "Set the graphics card mode to dedicated graphics mode." The Voice Activity Detection (VAD) unit in the first processing unit monitors the audio signal collected by the first audio input device in real time to determine whether there is voice input. When a voice signal is detected, subsequent processing is quickly triggered to avoid unnecessary computation and power consumption waste.
[0186] A dedicated CPU receives signals from the voice activity detection unit. After determining that there is voice input, it directs the digital signal processing unit (DSP) to preprocess the audio signal, such as noise reduction and filtering, to improve the quality of the voice signal and provide clearer and more accurate input for subsequent speech recognition.
[0187] A dedicated CPU accesses commonly used speech models and parameters stored in the local storage unit to perform in-depth analysis of the pre-processed speech signal. First, it uses voiceprint recognition to determine if the user is legitimate. If legitimate, it further performs keyword wake-up recognition on the speech signal. For example, it identifies keywords related to BIOS configuration, such as "graphics card mode" and "dedicated graphics card mode."
[0188] Based on the identified keywords, it is determined that relevant system functions need to be woken up to perform BIOS configuration operations. Therefore, the dedicated CPU sends the inferred BIOS configuration instructions (i.e., "set the graphics card mode to discrete graphics mode") and the wake-up instructions to the control unit (EC).
[0189] The EC (Electronic Bootloader) can first save the BIOS configuration instructions in its own storage area, waiting to be passed to the BIOS later. Then, in response to the wake-up command, the EC quickly triggers the electronic device to switch from a power-off state to a power-on state, officially initiating the device's boot process. At this stage, although the computer has not yet fully powered on, the EC has successfully saved the user's new BIOS configuration requirements, fully preparing for the subsequent rapid configuration to take effect.
[0190] When an electronic device powers on, the BIOS starts up and performs initialization settings. During this process, the BIOS actively establishes a communication connection with the Electronic Control Panel (EC), retrieving previously stored BIOS configuration commands from the EC via a specific communication protocol. The BIOS precisely configures the graphics card mode during initialization, switching it to dedicated graphics mode. This process ensures that the BIOS can complete the corresponding hardware configuration based on the user's voice commands, laying the foundation for the normal operation of the operating system.
[0191] After the BIOS completes its initialization settings, it continues to boot the electronic device to load the operating system (OS). Since the BIOS has already completed key configurations such as the graphics card mode based on the user's voice commands, the electronic device can directly enter an operating system environment perfectly matched to the new configuration. The entire process does not require restarting the device, avoiding cumbersome restart steps, greatly simplifying the operation process, saving users a significant amount of time, and providing a more convenient and efficient user experience.
[0192] As another optional embodiment of this application, a processing method is provided for Embodiment 6 of this application. In this embodiment, the audio device may include: an audio input device and an audio output device. The method may include, but is not limited to, the following steps:
[0193] Step S301: Obtain the first parameter; the first parameter is used to identify the system state of the electronic device.
[0194] For a detailed description of step S301, please refer to the relevant description of step S101 in Example 1, which will not be repeated here.
[0195] Step S302: If the first parameter indicates that the system state of the electronic device is in a non-working state, the first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module to obtain a first recognition result, and performs a first operation based on the first recognition result, so that after the system state of the electronic device is switched to a working state, the second processing unit of the electronic device can perform voice interaction processing based on the audio signal from the audio device.
[0196] For a detailed description of step S302, please refer to the relevant description of step S11 in the above embodiments, which will not be repeated here.
[0197] Step S303: After the system state of the electronic device is switched to the working state, it is determined that the electronic device has not entered the user login state. The first processing unit continues to perform voice recognition based on the audio signal of the audio input device by the first voice recognition module to obtain the second recognition result.
[0198] In this embodiment, the method for determining whether to enter the user login state can be referred to the relevant description in step S203 of embodiment 4, and will not be repeated here.
[0199] When a user starts their electronic device but hasn't logged in yet for various reasons, they might habitually call out the name of a voice assistant for an application on the device, such as "Xiao X," hoping to interact with that voice assistant and use the application's functions. The audio input device accurately captures the user's voice signal, such as "Xiao X," and quickly transmits it to the first processing unit. Upon receiving the signal, the first processing unit immediately performs speech recognition, analyzes and processes it to obtain a second recognition result, identifying that the user is attempting to call a voice assistant for interaction.
[0200] Step S304: Based on the second recognition result, it is determined that the user has a voice interaction requirement. The first processing unit generates voice feedback content based on the second recognition result and plays it through the audio output device.
[0201] The first processing unit generates corresponding voice feedback based on the specific content of the second recognition result. For example, in the scenario where the user calls "Xiao X," since the device is not logged in, the voice assistant cannot start normally to provide full service. The first processing unit can generate a simple reply such as "Yes." This reply is concise and clear, letting the user know that the device has received the call signal, giving the user the most basic response, and avoiding the user feeling that the device is unresponsive due to a long period of no feedback. For example, if the user simply habitually calls out "Xiao X," and the device quickly replies "Yes," the user can immediately perceive that the device is in an interactive state.
[0202] If the first processing unit determines that the user may have further operational intentions but is restricted due to not being logged in, it can generate feedback such as, "Hello, Xiao X is here! But the device isn't logged in yet. Once logged in, I can help you better. Would you like to log in now?" This feedback is both a friendly response to the user and proactively guides the user to log in to unlock more functions, improving the completeness and smoothness of the user's device usage. For example, if a user wants to check their schedule while not logged in, calling "Xiao X" and receiving such a response will clearly show them that logging in will provide access to more comprehensive services.
[0203] When the first processing unit identifies that a user's call may be related to a specific function, and that function is partially available when the user is not logged in, it generates feedback such as, "Hi, Little X, you heard me! The device isn't logged in. However, basic functions like local music playback are available now. If you want to use the online music library or personalized recommendations, you'll need to log in first." This feedback details the differences between logged-in and logged-out states for a specific function, allowing users to decide whether to log in based on their needs. This provides users with more accurate information and enhances their understanding and sense of control over the device's functions.
[0204] If a user calls out, "Hey, what's the weather like today?", the voice activity detection unit of the first processing unit (e.g., the AI chip) will detect the audio signal collected by the first audio input device in real time to determine if there is voice input. When a voice signal is detected, the dedicated CPU of the first processing unit can immediately perform keyword wake-up recognition on the voice signal.
[0205] If the voice contains a preset wake-up keyword (such as "Xiao x"), the dedicated CPU will further use the voiceprint recognition function to compare and analyze the collected voiceprint features with the pre-stored legitimate user voiceprint templates to determine whether the caller is a legitimate user.
[0206] If the user is identified as legitimate, since the device is currently not logged in, it cannot directly call service interfaces that require account login to obtain detailed and personalized weather data. However, the first processing unit can extract basic weather conditions from the device's locally stored general weather information (such as pre-installed weather data at the factory or recently cached public weather data). Figure 6 As shown, the system can generate voice feedback such as, "Hi, I can only provide you with some basic weather information right now. It's currently sunny locally, and the temperature is around 20 degrees Celsius. However, if you log in, I can get a more accurate and detailed weather forecast, including weather changes and air quality information for the next few days. Would you like to log in and try it?" This feedback is then played to the user through an audio power amplifier (AMP) and an audio playback unit (i.e., one implementation of an audio output device). This satisfies the user's need to check the weather as much as possible even when they are not logged in, while also guiding them to log in for better service.
[0207] It should be noted that when the user is not logged in, although the second processing unit may also have voice signal input and is connected to the audio output device, it cannot respond to the user's voice commands due to the lack of necessary permissions and data support. Its function is limited. Only after the user completes the login operation can the second processing unit function normally according to the system settings and user needs.
[0208] In this embodiment, when the user is not logged in, the first processing unit generates and plays targeted feedback on the user's voice commands in a timely manner based on locally stored general information. This satisfies the user's basic voice interaction needs and guides the user to log in, ensuring a basic voice interaction experience when not logged in and effectively improving the user's satisfaction with the interaction and audio experience of using electronic devices.
[0209] As another optional embodiment of this application, a processing method provided in Embodiment 7 of this application, in this embodiment, the audio device may include: an audio input device. This embodiment is mainly an implementation of the user's voice interaction processing performed by the first processing unit of the electronic device according to the audio signal of the audio device in Embodiment 1. Specifically, it may include, but is not limited to, at least one of the following:
[0210] Step S31: The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module, and determines the user's direction based on the first recognition result obtained from the speech recognition.
[0211] During speech recognition, the first processing unit can determine the user's approximate direction relative to the electronic device by analyzing information such as the time difference and intensity differences of the speech signals, combined with the layout of the microphone array. For example, by comparing the intensity of the speech signals received by different microphones, it can determine whether the user is located to the left, right, or directly in front of the device.
[0212] Step S32: Based on the user's orientation, drive the display screen to rotate so that the rotated display screen faces the user.
[0213] Based on the user orientation determined in step S31, the first processing unit generates a corresponding screen rotation command. For example, if the user is on the left side of the device, a rotation command to turn left is generated; if the user is on the right side of the device, a rotation command to turn right is generated.
[0214] The first processing unit controls the motor to drive the display screen to rotate according to the generated rotation command. The motor can be a stepper motor or a servo motor, etc., which can precisely control the rotation angle and speed.
[0215] During rotation, the first processing unit can monitor the rotation angle and speed in real time to ensure a smooth and accurate rotation process. At the same time, it can adjust the motor's output torque to avoid excessive vibration or noise during rotation.
[0216] Step S33: During the rotation of the display screen, drive the indicator light on the display screen to flash to indicate the rotation direction of the display screen.
[0217] Indicator lights are installed on the display screen or the device casing to indicate the screen's rotation direction. These indicator lights can be LEDs, which offer advantages such as high brightness and long lifespan.
[0218] During screen rotation, the first processing unit drives the indicator lights to flash according to preset logic. For example, when the screen rotates to the left, the indicator lights can flash to the left; when the screen rotates to the right, the indicator lights can flash to the right. Through the flashing of the indicator lights, users can intuitively understand the screen rotation direction, improving the user experience.
[0219] In this embodiment, regardless of whether the electronic device is in a non-working state (such as sleep, power off, or a specific low-power mode) or a working state, the first processing unit can perform voice recognition and processing, as well as generate screen rotation commands and control the motor to drive the screen to rotate. Simultaneously, an indicator light can flash to indicate the rotation direction.
[0220] For example, in a non-working state, such as Figure 7 As shown, the user stands to the left of the device and issues a voice command to the device to "turn left". An audio input device (such as a microphone) can pick up the voice signal and transmit it as an electrical signal to the first processing unit (such as an AI chip).
[0221] Upon receiving the audio signal, the first processing unit immediately activates the voice recognition function to accurately analyze the command. While recognizing the command "turn left," it uses a built-in positioning algorithm or combines it with the directional information from the audio input to determine that the user's current position is on the left side of the device.
[0222] Based on the above identification and location results, referring to Figure 8 In this connection method, the first processing unit (AI chip) can generate a corresponding rotation command to turn left and send the command to the motor control module. Based on the received command, the motor control module precisely controls the motor (which can be represented as a Motor) to drive the screen (which can be represented as a Panel) to rotate to the left as required.
[0223] Meanwhile, to allow users to intuitively understand the device's rotation direction, the first processing unit generates a control command and sends it to the indicator light control module. Based on the command, the indicator light control module controls the indicator light (which can be represented as an LED) to flash to the left to indicate the screen's rotation direction. By observing the screen's rotation and the indicator light's flashing, users can clearly understand the device's response to voice commands.
[0224] Or, when the electronic device is not in operation, such as Figure 7As shown, the user issues a voice command to the device saying "Hi, Xiaox". The audio input device on the device (such as a microphone) captures the voice signal and transmits it to the first processing unit.
[0225] The first processing unit performs speech recognition processing on the audio signal, recognizing the user's request to inquire about the weather, and determines that this instruction requires waking up the system from its sleep state. Based on this determination, the first processing unit generates a wake-up instruction and sends it to the embedded controller (EC).
[0226] Upon receiving the wake-up command, the EC immediately performs a wake-up operation on the system-on-a-chip (SOC), restoring it from sleep mode to normal operation. Simultaneously, the EC sends power-on commands to the power control modules of the audio power amplifier unit (AMP) and the audio playback unit (Speaker), respectively, turning on the power to the AMP and Speaker to prepare for subsequent voice feedback.
[0227] During the wake-up process, in order to promptly inform the user that the system has responded to their command, the first processing unit generates simple voice feedback, such as "Yes." This voice feedback is sent to the AMP for audio signal amplification and then played back through the speaker, allowing the user to hear the device's response and confirm that the system has begun processing their inquiry.
[0228] like Figure 9 As shown, when electronic devices are in normal working condition, they can exhibit rich and intelligent interactive functions.
[0229] If a user issues a "turn left" command, the device's sensitive audio input (such as a microphone) will quickly capture this voice signal and transmit it instantly to the first processing unit. Leveraging its powerful voice recognition capabilities, the first processing unit quickly parses the "turn left" command. It then accurately generates a control command and sends it to the motor drive module. Based on the received command, the motor drive module precisely controls the motor's operation, thereby driving the display screen to rotate to the left as required, presenting the user with the desired visual effect.
[0230] When a user says, "Hey X, what's the weather like today?", the device's interaction process becomes more complex and intelligent. Assuming the user has successfully logged into the system, the audio input device immediately captures the voice signal and transmits it to the first processing unit. The first processing unit first performs wake-up word recognition, accurately determining that the user is waking up the device to interact. Then, it initiates a strict user verification mechanism, comparing pre-stored user information (such as voiceprint features, account password, etc.) to confirm the legitimacy of the user currently performing the interaction.
[0231] After completing wake word recognition and user verification, the first processing unit sends a notification command to the second processing unit. Upon receiving the command, the second processing unit quickly initiates a series of operations based on the specific app installed on the device. It activates the large-scale model inference function, calling upon a powerful artificial intelligence model to provide intelligent support for subsequent weather queries and responses. Simultaneously, the second processing unit also activates the audio input device's receiving function to ensure that it can receive any subsequent supplementary commands or interactive information from the user in real time and accurately.
[0232] At this point, powered by a large model, the app begins weather query inference. It connects to online weather data interfaces to obtain real-time, accurate weather information and uses natural language processing technology to translate this information into natural, fluent language. For example, it might say, "Hello, according to the latest meteorological data, today our area is sunny with temperatures between 20°C and 28°C and excellent air quality, perfect for outdoor activities. However, the wind will be slightly stronger in the afternoon, reaching level 3-4, so remember to take precautions when going out. Also, the probability of precipitation today is low, only 10%, so don't worry about sudden rain." Finally, the app clearly broadcasts the weather conditions to the user through the device's speaker, allowing them to stay informed about the current weather and providing a reference for travel and activity planning.
[0233] In this embodiment, whether in a non-working state or a working state, the first processing unit can quickly recognize the user's voice commands and control the screen rotation or wake up the system. At the same time, the rotation direction is intuitively indicated by flashing indicator lights, providing the user with a convenient, intuitive and intelligent interactive experience.
[0234] As another optional embodiment of this application, this embodiment provides a processing method for embodiment 8 of this application. This embodiment is mainly an implementation of the above-mentioned user voice interaction processing performed by the first processing unit of the electronic device according to the audio signal of the audio device. In this embodiment, the audio input device may include: an audio input device and an audio output device.
[0235] The audio input device may include the first audio input device and the second audio input device mentioned above, and the specific connection method will not be described in detail here.
[0236] like Figure 10 As shown, the audio output device may include: an audio power amplifier unit (which may be denoted as AMP) connected to both the first processing unit and the second processing unit, and an audio playback unit (which may be denoted as Speaker) connected to the audio power amplifier unit. This design allows both processing units to output audio signals to the audio power amplifier unit (AMP) without interfering with each other.
[0237] Embedded controllers (ECs) play a crucial role in the power management of audio power amplifier units (AMPs). ECs can precisely control the power-on and power-off operations of the AMPs. For example, when the electronic device is not in operation, the EC can cut off the power to the AMP, putting it into a low-power mode, thereby reducing the overall power consumption of the device; conversely, when the electronic device needs to use voice interaction functions, the EC can promptly power on the AMP, restoring it to normal operation. This flexible power control not only helps save energy but also extends the lifespan of the device.
[0238] The AMP can be connected, but is not limited to, via I2S and an AI chip (i.e., an implementation of a first processing unit), and the AMP can be connected, but is not limited to, via Swire and a SOC (i.e., an implementation of a second processing unit).
[0239] In this embodiment, Figure 10 The connection methods of the audio power amplifier unit (AMP) and each processing unit can be found in the relevant technical information and standards in this field. The details of these interfaces will not be described in detail here.
[0240] The first processing unit of the electronic device performs user voice interaction processing based on the audio signal from the audio device, which may include, but is not limited to, the following steps:
[0241] Step S41: Control the audio power amplification unit to enter the first mode; in the first mode, the audio power amplification unit is only allowed to process the audio feedback signal output by the first processing unit.
[0242] For example, in non-operating mode, when a user issues a voice command, the first audio input device acquires the signal and immediately transmits it to the first processing unit. After performing voice recognition and processing, if the first processing unit needs to provide feedback to the user, such as confirming receipt of the command or giving a simple reply, it generates a corresponding audio feedback signal. At this time, since the audio power amplification unit is in the first mode, it focuses only on processing this feedback signal output by the first processing unit and then plays it to the user through the audio playback unit. This mode setting has significant advantages: in non-operating mode, it effectively avoids interference signals that may be generated by the second processing unit, while greatly reducing power consumption, because the audio power amplification unit does not need to process extra signals and only focuses on the feedback output of the first processing unit, thereby improving the energy efficiency of the device.
[0243] Step S42: After the electronic device switches from the non-working state to the working state, the audio power amplification unit is controlled to switch from the first mode to the second mode; in the second mode, the audio power amplification unit is allowed to process the audio feedback signals output by the first processing unit and the second processing unit.
[0244] In operation, there may be situations where only the first processing unit generates an audio feedback signal, only the second processing unit generates an audio feedback signal, or both the first and second processing units generate audio feedback signals.
[0245] In the second mode, when only the first processing unit generates an audio feedback signal, the audio power amplification unit can receive and process only the audio feedback signal generated by the first processing unit.
[0246] For example, when a user is not logged in, some applications (APPs) on the electronic device that require user login information to function properly have limited functionality. In this case, if the user issues basic commands, such as asking for the current time or checking the local weather—requirements that don't require login—the first processing unit will process these commands directly. For instance, if a user asks "What time is it?", the first processing unit obtains the system time and generates an audio feedback signal saying "It is 10:00 AM." Because the user is not logged in, the relevant APP cannot participate in the interaction. In the second mode, the audio power amplification unit only receives this audio feedback signal output by the first processing unit, processes it, and plays it to the user through the audio playback unit. This processing method ensures that users can quickly obtain basic information even when not logged in, guaranteeing the availability of the device's basic functions.
[0247] In the second mode, when only the second processing unit generates an audio feedback signal, the audio power amplification unit can receive and process only the audio feedback signal generated by the second processing unit.
[0248] For example, after a user logs in and enters normal working mode, they interact with the electronic device through a specific application (APP). At this point, the second processing unit plays a crucial role. For instance, if a user uses a music app and commands, "Play rock songs from my favorites list," the first processing unit initially identifies the command type and then transfers the task to the second processing unit. The second processing unit connects to the music app's database, filters rock songs from the user's favorites list, and generates an audio feedback signal saying, "Playing rock songs from your favorites list for you." In this second mode, the audio power amplification unit only receives and processes this signal from the second processing unit before playing it to the user. This approach fully utilizes the second processing unit's ability to deeply interact with the application, providing users with personalized and precise services.
[0249] In the second mode, when both the first and second processing units generate audio feedback signals, the audio power amplification unit can receive the audio feedback signals generated by the first and second processing units. The audio feedback signals can be processed sequentially according to preset rules.
[0250] For example, in some complex scenarios involving multi-task collaborative processing, the first processing unit and the second processing unit will simultaneously generate audio feedback signals.
[0251] For example, when a user exercises using a fitness app, the first processing unit monitors the user's exercise data in real time, such as heart rate and exercise duration. When the user's heart rate exceeds the safe range, the first processing unit generates an audio feedback signal saying, "Your current heart rate is too high; please adjust your exercise intensity." Simultaneously, the second processing unit, based on the user's set fitness goals and current exercise progress, connects to a cloud database to obtain more fitness suggestions and generates an audio feedback signal saying, "Based on your goals, we suggest you perform 10 minutes of stretching exercises next." At this time, the audio power amplification unit receives both audio feedback signals from the first and second processing units simultaneously in the second mode. To ensure the user receives information clearly and systematically, the audio power amplification unit can process the audio feedback signals according to preset rules, such as first playing health and safety-related prompts (feedback from the first processing unit), then playing information about subsequent exercise suggestions (feedback from the second processing unit), and finally playing it to the user through the audio playback unit. This processing method allows the device to comprehensively utilize the advantages of both processing units, providing users with a more comprehensive and detailed service experience.
[0252] In this embodiment, the audio power amplifier unit (AMP) is connected to both the first processing unit and the second processing unit. When not in operation, the AMP is in a first mode that only processes the feedback signal from the first processing unit, thereby reducing power consumption and avoiding interference.
[0253] After switching to working mode, it enters a second mode that can process feedback signals from both processing units. It can flexibly respond to the situation where signals are generated by different processing units, process multiple signals according to preset rules, make full use of the advantages of the two processing units, ensure the stability, accuracy and comprehensiveness of the voice interaction function in different states, and improve the overall performance and user experience of voice interaction of electronic devices.
[0254] As another optional embodiment of this application, a processing method provided in embodiment 9 of this application, in which the audio device may include: an audio input device, and the method may include, but is not limited to, the following steps:
[0255] Step S401: Obtain the first parameter; the first parameter is used to identify the system status of the electronic device.
[0256] For a detailed description of step S401, please refer to the relevant description of step S101 in Example 1, which will not be repeated here.
[0257] Step S402: The ambient sound collected by the audio input device is detected by the voice activity detection unit (VAD) within the first processing unit.
[0258] The voice activity detection unit can distinguish voice signals from environmental noise (such as fan noise and traffic noise) through spectrum analysis.
[0259] Alternatively, the short-time energy of the audio signal can be calculated, and if it exceeds a threshold, it is considered valid speech.
[0260] For example, if the user does not speak, the VAD detects ambient noise (such as air conditioning noise) and does not wake up the subsequent processing unit.
[0261] When the user says a wake word (such as "Hi, Xiaox"), the VAD detects a valid voice signal and triggers the subsequent process.
[0262] Step S403: If the voice activity detection unit does not detect a valid sound signal, then control the other functional units in the first processing unit, excluding the voice activity detection unit, to be in a sleep state.
[0263] Controlling other functional units in the first processing unit, excluding the voice activity detection unit, to a sleep state may include, but is not limited to:
[0264] The clock signals of all functional units in the first processing unit, except for the voice activity detection unit, are turned off, and the register data is retained.
[0265] Alternatively, the power supply voltage of the other functional units in the first processing unit, excluding the voice activity detection unit, can be reduced to below 1.8V to reduce leakage current.
[0266] For example, such as Figure 11 As shown, the VAD (Voice Activity Detection Unit) can turn off the clock signals of functional units such as the DSP and CPU in the AI chip or reduce their power supply voltage.
[0267] Step S404: If the voice activity detection unit detects a valid sound signal, the other functional units are switched from the sleep state to the working state.
[0268] The corresponding method of turning off the clock signal can be used to re-enable the clock signal, restore the register state, and thus complete the switch from sleep state to working state.
[0269] The corresponding method of reducing the voltage can raise the voltage to the normal operating voltage (e.g., 1.8V → 3.3V).
[0270] Step S405: If the first parameter represents the system state of the electronic device as being in a non-working state, the first processing unit of the electronic device performs voice interaction processing on the user based on the audio signal of the audio device, so that after the system state of the electronic device is switched to a working state, the second processing unit of the electronic device can perform voice interaction processing on the audio signal of the audio device; wherein, the first processing unit can identify the user based on the audio signal.
[0271] For a detailed description of step S405, please refer to the relevant description of step S102 in Example 1, which will not be repeated here.
[0272] In this embodiment, the voice activity detection unit detects ambient sounds collected by the audio input device, accurately distinguishing between voice signals and ambient noise, thus avoiding invalid wake-ups. When no valid voice signal is detected, the other functional units in the first processing unit, except for the voice activity detection unit, are put into a sleep state, reducing power consumption. When a valid voice signal is detected, the other functional units are switched from sleep to working state, ensuring normal voice interaction processing. Overall, this effectively reduces the power consumption of the audio device when there is no valid voice, while ensuring timely response and interaction processing when there is valid voice.
[0273] In this embodiment, combined with Figure 12 This section explains the relevant functions in the voice interaction processing of electronic devices. For example,
[0274] The AI chip has a dynamic MIC channel switching function:
[0275] The first parameter, i.e. the system status, can be obtained from EC to determine whether the LA4 system is in S0 (i.e., working state).
[0276] If not, LA4 enters monitoring mode, where the AI chip receives the MIC signal for sound recognition but does not transmit sound to the PCH (i.e., SOC). Additionally, the AI chip can notify the AMP to enter I2S mode (i.e., mode one), in which the AMP only processes the audio feedback signal output by the AI chip.
[0277] If so, LA4 enters monitoring mode + transmission mode. The AI chip receives the MIC signal and sends it to the AI chip's internal processor (e.g., a dedicated CPU) for speech recognition, while simultaneously transmitting the MIC sound to the PCH (i.e., the SOC). Additionally, the AI chip can notify the AMP to enter I2S+Swire mode (i.e., the second mode), in which the AMP is allowed to process the audio feedback signals output by the AI chip and the PCH (i.e., the SOC).
[0278] When the PCH (i.e., SOC) is in S0 state, it can start receiving MIC sound signals. When the user logs in (i.e. the electronic device is in the user login state), the APP can turn on sound detection and perform corresponding processing.
[0279] When the app generates a voice feedback signal, the app can wake up the AMP and play it through the Swire interface.
[0280] Whether the LA system is S0 or not, the AI chip's dynamic switching of Voice ID / Wake up function can be triggered.
[0281] AI chip's dynamic switching of Voice ID / Wake up function:
[0282] The AI chip determines whether to log in.
[0283] If it is "Log in", the system detects whether playback is currently being sent through the speaker by receiving the volume from the audio driver.
[0284] If playback is performed through the Speaker, the PCH (i.e., SOC) is notified to call the APP for processing, based on the Voice ID / Wake up function of the CPU in the PCH (i.e., SOC).
[0285] If the audio is not played through the speaker, the Voice ID / Wake up function inside the AI chip can be invoked.
[0286] Voice ID / Wake up functionality within the AI chip:
[0287] Detect if there is sound.
[0288] If a sound is detected, determine if it is a wake word; if no sound is detected, continue detecting if a sound is detected.
[0289] If it is a wake word, determine if it is the registrant's voice; if it is not a wake word, continue to detect if there is any sound.
[0290] If the voice is not that of the registrant, continue to detect if there is any sound; if it is the voice of the registrant, the AI chip can be triggered to perform subsequent processing.
[0291] When the AI chip hears the registered user's voice, it can determine whether the current state is S0.
[0292] If not in S0 state, notify EC to wake up the system to enter S0 state, and wake up AMP via I2C or GPIO and enter I2S mode (i.e., first mode).
[0293] In non-S0 state, after the AMP is woken up, if it receives an audio signal from the AI chip via I2S Play audio, it can play the audio.
[0294] If the state is S0, determine whether the user should log in.
[0295] If Log in, notify the app to enable sound detection and perform appropriate processing.
[0296] AMP channel dynamic switching function:
[0297] In non-S0 state, the AMP enters I2S mode (i.e., first mode). If the current AMP is in a low-power state (e.g., power off), the AMP can be woken up.
[0298] In S0 state, the AMP enters I2S+Swire mode (i.e., the second mode). If the current AMP is in a low-power state (e.g., power off), the AMP can be woken up.
[0299] In another embodiment of this application, an electronic device is provided, comprising:
[0300] A first audio input device is connected to a first processing unit;
[0301] A second audio input device is connected to a second processing unit; the first audio input device operates independently of the second audio input device; the power consumption of the first processing unit is lower than that of the second processing unit.
[0302] The first processing unit is configured to:
[0303] If the first parameter represents the system state of the electronic device as being in a non-working state, the user's voice interaction processing is performed based on the audio signal from the audio device.
[0304] The second processing unit is used to perform voice interaction processing based on the audio signal from the audio device after the system state of the electronic device is switched to the working state; wherein, the first processing unit is able to identify the user based on the audio signal.
[0305] It should also be noted that the unit embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the unit embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0306] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0307] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0308] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable unit. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A processing method, comprising: Get the first parameter; The first parameter is used to identify the system status of the electronic device; If the first parameter indicates that the system state of the electronic device is in a non-working state, the first processing unit of the electronic device performs voice interaction processing on the user based on the audio signal of the audio device, so that after the system state of the electronic device is switched to a working state, the second processing unit of the electronic device can perform voice interaction processing on the audio signal of the audio device; wherein, the first processing unit can identify the user based on the audio signal.
2. The processing method according to claim 1, wherein the audio device comprises: Audio input device; The first processing unit of the electronic device performs user voice interaction processing based on the audio signal from the audio device, including: The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module to obtain a first recognition result, and performs a first operation based on the first recognition result; The second processing unit of the electronic device is capable of performing voice interaction processing based on the audio signal from the audio device, including: The second processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the second speech recognition module.
3. The processing method according to claim 2, wherein the audio input device comprises: First audio input device and second audio input device; The first audio input device is connected to the first processing unit; The second audio input device is connected to the second processing unit; the first audio input device operates independently of the second audio input device. The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module to obtain a first recognition result, including: The first processing unit of the electronic device performs speech recognition based on the audio signal from the first audio input device using the first speech recognition module to obtain a first recognition result. The second processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the second speech recognition module, including: The second processing unit of the electronic device performs speech recognition based on the second speech recognition module, at least according to the audio signal from the second audio input device.
4. The processing method according to claim 3, wherein the audio device further comprises: Audio output device; The processing method further includes: After the electronic device switches its system state to the working state, the first processing unit determines that the electronic device has entered the user login state and identifies whether the audio output device is in the playback state. If so, the first processing unit stops performing speech recognition based on the audio signal from the first audio input device and sends the audio signal from the first audio input device to the second processing unit; The second processing unit of the electronic device performs speech recognition based on the second speech recognition module, at least according to the audio signal from the second audio input device, including: The second processing unit of the electronic device performs speech recognition based on the audio signals from the first audio input device and the second audio input device using the second speech recognition module.
5. The processing method according to claim 2, wherein the non-working state includes a power-off state; the step of the first processing unit of the electronic device performing speech recognition based on the audio signal from the audio input device by the first speech recognition module to obtain a first recognition result, and performing a first operation based on the first recognition result, includes at least one of the following: Based on the first identification result, it is determined that the operating system needs to be woken up. The first processing unit of the electronic device sends a wake-up command to the control unit, so that the control unit responds to the wake-up command and triggers the electronic device to switch from the power-off state to the power-on state. Based on the first identification result, it is determined that basic input / output system configuration needs to be performed. The first processing unit of the electronic device obtains the basic input / output system configuration instruction from the first identification result and sends the configuration instruction to the control unit, so that the control unit stores the configuration instruction and transmits the configuration instruction to the basic input / output system, so that the basic input / output system performs initialization settings based on the configuration instruction during power-on.
6. The processing method according to claim 2, wherein the audio device further comprises: Audio output device; The processing method further includes: After the electronic device switches its system state to working state, it is determined that the electronic device has not entered the user login state. The first processing unit continues to perform voice recognition based on the audio signal of the audio input device by the first voice recognition module to obtain the second recognition result. Based on the second recognition result, it is determined that the user has a need for voice interaction. The first processing unit generates voice feedback content based on the second recognition result and plays it through the audio output device.
7. The processing method according to claim 1, wherein the audio device comprises: Audio input device; The user's voice interaction processing performed by the first processing unit of the electronic device based on the audio signal from the audio device includes at least one of the following: The first processing unit of the electronic device performs speech recognition based on the audio signal from the audio input device using the first speech recognition module, and determines the user's direction based on the first recognition result obtained from the speech recognition. Based on the user's orientation, drive the display screen to rotate so that the rotated display screen faces the user; During the rotation of the display screen, the indicator lights on the display screen are driven to flash to indicate the direction of rotation of the display screen.
8. The processing method according to claim 1, wherein the audio device comprises: Audio output device; The audio output device includes: an audio power amplification unit connected to both the first processing unit and the second processing unit, and an audio playback unit connected to the audio power amplification unit; The first processing unit of the electronic device performs user voice interaction processing based on the audio signal from the audio device, including: The audio power amplification unit is controlled to enter a first mode; in the first mode, the audio power amplification unit is only allowed to process the audio feedback signal output by the first processing unit; After the electronic device switches from the non-working state to the working state, the audio power amplification unit is controlled to switch from the first mode to the second mode; in the second mode, the audio power amplification unit is allowed to process the audio feedback signals output by the first processing unit and the second processing unit.
9. The processing method according to claim 1, wherein the audio device comprises: Audio input device; The processing method further includes: The ambient sound collected by the audio input device is detected by the voice activity detection unit inside the first processing unit; If the voice activity detection unit does not detect a valid sound signal, then the other functional units in the first processing unit, excluding the voice activity detection unit, are put into a sleep state. If the voice activity detection unit detects a valid sound signal, it switches the other functional units from the sleep state to the working state.
10. An electronic device, comprising: A first audio input device is connected to a first processing unit; A second audio input device is connected to the second processing unit; The first audio input device operates independently of the second audio input device; the power consumption of the first processing unit is lower than that of the second processing unit; The first processing unit is configured to: If the first parameter represents the system state of the electronic device as being in a non-working state, the user's voice interaction processing is performed based on the audio signal from the audio device. The second processing unit is used to perform voice interaction processing based on the audio signal from the audio device after the system state of the electronic device is switched to the working state; wherein, the first processing unit is able to identify the user based on the audio signal.