Method and electronic device for voice interaction
By reducing power consumption in time after the electronic device is lifted, the existing breath wake-up solution has solved the problem of high power consumption when the low-power memory space is insufficient, and the power consumption is effectively reduced when the universality remains unchanged.
Patent Information
- Application Number
- CN202411955423.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-28
AI Technical Summary
The existing breath wake-up solution consumes a high power consumption while ensuring universality, especially when the low-power memory space is insufficient, resulting in poor universality of the solution.
In the breath wake-up scenario, when the electronic device is lifted and then lifted, it will reduce the power consumption in time, eliminate unnecessary power consumption caused by lifting the hand without breath wake-up, and achieve a reduction in power consumption.
On the premise of ensuring the universality of the breath wake-up solution, the power consumption of breath wake-up is effectively reduced and unnecessary power consumption caused by wrong wake-up is avoided.
Smart Images

Figure CN119376683B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice interaction technology, and in particular to a voice interaction method and electronic device. Background Art
[0002] When users perform voice interaction on electronic devices such as mobile phones, tablets, smart wearable devices, car computers, smart speakers, etc., the traditional method is to say a specific wake-up word recorded in advance to wake them up. Users need to remember the different wake-up words for different electronic devices, and they cannot wake them up if they say the wrong ones. In order to solve this problem, the breath wake-up method has emerged. In the breath wake-up method, users only need to get close to the microphone of the electronic device with the screen off or on and speak the interactive voice. They no longer need to wake up with the wake-up word and then speak the interactive voice. However, since breath wake-up requires the electronic device to continue working even when the screen is off and the application processor (AP) is in sleep mode, this leads to higher power consumption requirements.
[0003] In order to reduce power consumption, a solution has emerged to deploy the breath wake-up module in the low-power space of the audio digital signal processor (ADSP), so that the ADSP can continuously detect breath wake-up in low-power mode and wake up the AP for voice interaction when the breath wake-up function is triggered. However, since only the low-power space of high-end (high-configuration) ADSPs has enough low-power space to achieve the above deployment, and the low-power memory space of more ADSPs is not enough to deploy the complete breath wake-up module, the above solution has poor universality. Under the premise of giving priority to universality, a solution has emerged to divide the breath wake-up module into functional parts and deploy them in the low-power space and non-low-power space of ADSP respectively, so that the ADSP can work in low-power mode most of the time when the screen is turned off, and work in normal working mode under certain trigger conditions to complete breath wake-up detection and wake up the AP. However, compared with the solution that is fully deployed in the low-power space of ADSP, this increases power consumption.
[0004] Therefore, how to reduce the power consumption of breath awakening as much as possible while ensuring the universality of the breath awakening solution is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The present application provides a voice interaction method and electronic device, which can effectively reduce the power consumption of breath awakening while ensuring the universality of the breath awakening solution.
[0006] In a first aspect, a method for voice interaction is provided, which is applied to an electronic device. The method includes: at a first moment, the electronic device is lifted, and the screen of the electronic device is in an off state. The power consumption of the electronic device changes from a first value to a second value. Before the first moment, the screen of the electronic device is in an off state, and the electronic device is not lifted. The second value is greater than the first value; at a second moment after the first moment, the electronic device is not lifted, and the screen of the electronic device is in a screen-off state. The power consumption of the electronic device changes from the second value to the first value.
[0007] In the technical solution of the present application, mainly in the voice interaction process in the breath wake-up scenario, when the electronic device is lifted and then the lifting ends, the power consumption is timely reduced back to a low level, so as to eliminate the unnecessary power consumption caused by the non-breath wake-up hand-lifting, and reduce the overall power consumption.
[0008] Combined with the first aspect, in some implementation manners of the first aspect, at a second moment after the first moment, when the electronic device is not lifted and the screen of the electronic device is in a screen-off state, and the power consumption of the electronic device changes from the second value to the first value, it may include: at the second moment, based on the electronic device changing from the lifted state to the non-lifted state, the power consumption of the electronic device changes from the second value to the first value. In this implementation manner, it explains why the power consumption value changes at the second moment, because the electronic device changes based on the change of the lifted state in the screen-off state. It can also be understood that the power consumption change at the second moment occurs under the trigger of the change of the lifted state from lifted to non-lifted. Combining with the above-mentioned first moment when the lifted state of the electronic device changes from non-lifted to lifted and increases, it constitutes a complete solution for changing the power consumption based on the change of the lifted state of the electronic device. In the traditional solution, there will be no such second moment as in the present application (or it can be understood that the second moment in the traditional solution will not be like in the present application where the electronic device will end the process based on the change of the lifted state). Once the electronic device is lifted in the traditional solution and the power consumption changes from the first value to the second value, it needs to wait at least for the following third moment (no appropriate voice data is collected within the preset duration) or the sixth moment (the entire voice interaction process is completed) to return to the first value again. Therefore, compared with this, the present application plays the effect of "timely stopping losses", and can end the process in time after the lifted state changes, without wasting power by executing the subsequent process in vain.
[0009] In combination with the first aspect, in some implementations of the first aspect, the above method further includes: before the first moment, the first processor of the electronic device continuously detects a lift event in the low-power mode, where the lift event is used to indicate whether the electronic device is lifted. When the lift event is detected, it means the electronic device is lifted, and when a non-lift event is detected, it means the electronic device is not lifted; at the first moment, in response to detecting the lift event, the first processor switches from the low-power mode to the normal operating mode, and starts to determine whether to trigger a breath wake-up event in the normal operating mode, and continues to obtain the detection result of the lift event in the normal operating mode; at the second moment, in response to detecting a non-lift event, the first processor switches from the normal operating mode to the low-power mode, and continuously detects the lift event in the low-power mode. In this implementation, the first processor continuously detects the lift event in the low-power mode. Once the lift event is detected, it means the electronic device is lifted, that is, the first moment arrives. The electronic device switches from the low-power mode to the normal operating mode in response to detecting the lift event, so that the power consumption increases from the previous first value to the subsequent second value starting from the first moment. From the first moment until the second moment, the first processor is in the normal operating mode. Until the second moment, the electronic device detects a non-lift event again, and then it knows that the lifted state has ended, so it directly returns the first processor to the low-power mode based on this, instead of waiting for a sufficient duration and executing subsequent processes, thus effectively reducing the power consumption.
[0010] In combination with the first aspect, in some implementations of the first aspect, the above method further includes: between the first moment and the third moment, the electronic device remains lifted, and the screen of the electronic device is in the off-screen state. The time interval between the first moment and the second moment is less than the time interval between the first moment and the third moment, and the time interval between the first moment and the third moment is a preset time interval; at the third moment, the power consumption of the electronic device changes from the second value to the first value. In this implementation, the step of waiting for a fixed duration in the traditional solution is continued, so that on the basis of superimposing the solution of this application, the timely end of the process can be ensured. That is, in the case where the second moment does not occur, that is, the lifted state of the electronic device does not change, the process ends after counting the preset time interval from the first moment, and the power consumption is reduced back to the first value.
[0011] In one example, at the third moment, the power consumption of the electronic device changes from the second value to the first value, including: at the third moment, based on the fact that the electronic device does not collect voice data that meets the first preset condition, the power consumption of the electronic device changes from the second value to the first value. The first preset condition includes: the signal strength of the voice data collected by the first microphone of the electronic device is greater than or equal to the first preset signal strength threshold, the signal strength of the voice data collected by the second microphone of the electronic device is greater than or equal to the second preset signal strength threshold, and the difference between the signal strength of the voice data collected by the first microphone and the signal strength of the voice data collected by the second microphone is greater than or equal to the preset signal strength difference threshold. In this example, the triggering factor for adjusting the power consumption at the third moment is given, that is, based on the fact that the electronic device does not collect voice data that meets the first preset condition. Combining the above, there are two ways for this application to trigger "return to low power consumption", the lifted state is changed (or understood as the electronic device ends the lifted state) and no voice data that meets the first preset condition is collected within the preset time interval. The two triggering methods can be executed independently or superimposed. Only executing the first method is the solution of this application, only executing the second method is the traditional solution, and executing both methods has the best power consumption reduction effect.
[0012] In one example, at the third moment, based on the fact that the electronic device does not collect voice data that meets the first preset condition, the power consumption of the electronic device changes from the second value to the first value, including: before the first moment, the first processor of the electronic device operates in the low-power mode; between the first moment and the third moment, the first processor operates in the normal working mode; the first processor continuously reads the voice data collected by the first microphone and the second microphone starting from the first moment; at the third moment, based on the fact that the electronic device is lifted and the electronic device does not collect voice data that meets the first preset condition, the first processor determines not to trigger the breath wake-up event, controls the first processor to switch from the normal working mode to the low-power mode, and continuously detects the lift event in the low-power mode. In this example, it mainly shows that the change in the power consumption value is brought about by the change in the working mode of the first processor. Before the first moment, the first processor is in the low-power mode, so the power consumption is the lowest first value. Between the first moment and the third moment, it is in the normal working mode, so the power consumption is the higher second value. At the third moment, since no voice data that meets the first preset condition is collected and the breath wake-up event is not triggered, the first processor works in the low-power mode again, so the power consumption changes from the first value to the second value and then back to the third value during the whole process.
[0013] In combination with the first aspect, in some implementations of the first aspect, the above method further includes: at a fifth moment after the first moment, the electronic device turns on the screen, and the power consumption of the electronic device is a third value, and the third value is greater than the second value. In this implementation, the power consumption change brought about by the successful breath wake-up is mainly given.
[0014] In one example, at a fifth moment after the first moment, the electronic device turns on the screen, and the power consumption of the electronic device is a third value, and the third value is greater than the second value, including: at the first moment, the first processor of the electronic device switches from the low-power mode to the normal operating mode, so that the power consumption of the electronic device increases from the first value to the second value; at a fourth moment after the first moment, based on the electronic device starting to collect voice data that meets the first preset condition, the first processor starts to continuously read the collected voice data; the first preset condition includes: the signal strength of the voice data collected by the first microphone of the electronic device is greater than or equal to the first preset signal strength threshold, the signal strength of the voice data collected by the second microphone of the electronic device is greater than or equal to the second preset signal strength threshold, and the difference between the signal strength of the voice data collected by the first microphone and the signal strength of the voice data collected by the second microphone is greater than or equal to the preset signal strength difference threshold; at a fifth moment after the fourth moment, the first processor finishes reading the first voice data, and the first processor wakes up the second processor based on the first voice data meeting the second preset condition, so that the second processor switches from the low-power mode to the normal operating mode and performs voice recognition and response on the first voice data in the normal operating mode, so that the power consumption of the electronic device increases from the second value to the third value, and meeting the second preset condition is used to indicate that the first voice data can be used to trigger the breath wake-up event. In this example, how the power consumption changes from the first moment to the fifth moment as the working mode of the processor changes is given. Before the first moment, both the first processor and the second processor are in the low-power mode. Between the first moment and the fourth moment, the first processor is in the normal operating mode. Between the fourth moment and the fifth moment, the first processor is in the normal operating mode. After the fifth moment, both the second processor and the first processor are in the normal operating mode, so that the power consumption continues to increase in the case of successful wake-up.
[0015] In another example, the above method further includes: at the fifth moment, based on the first voice data not meeting the second preset condition, the first processor switches from the normal operating mode to the low-power mode, so that the power consumption of the electronic device changes from the second value to the first value. In this example, when the first processor determines not to trigger the breath wake-up event, it switches back to the low-power mode.
[0016] In another example, the above method further includes: at a sixth moment after the fifth moment, the screen of the electronic device changes from the lit state to the off state, and the power consumption of the electronic device changes from a third value to a first value. In this example, when the screen of the electronic device goes off again, the power consumption returns to the first value.
[0017] In another example, at a sixth moment after the fifth moment, the screen of the electronic device changes from the lit state to the off state, and the power consumption of the electronic device changes from a third value to a first value, including: at the sixth moment, based on the fact that the second processor has finished responding to the first voice data, both the first processor and the second processor switch to the low-power mode, so that the power consumption of the electronic device changes from the third value to the first value. In this example, after the second processor finishes responding, both processors are switched to the low-power mode, thereby reducing the power consumption to the first value.
[0018] In another example, the above method further includes: at a fourth moment, based on the fact that the electronic device has not collected voice data that meets the first preset condition, the first processor has not read voice data that meets the first preset condition, and the electronic device continues to collect voice data until a third moment, and the time interval between the third moment and the first moment is a preset time interval. In this example, if no voice data has been collected at the fourth moment, the collection will continue until the third moment arrives.
[0019] In a second aspect, a voice interaction device is provided, and the device includes units composed of software and / or hardware for executing any one of the methods in the first aspect.
[0020] In a third aspect, an electronic device is provided. The electronic device includes a memory, multiple processors, sensors, an audio collection device, and a display screen. The sensors are used to collect the motion data of the electronic device, the audio collection device is used to collect the voice of the electronic device, the display screen is used to display an interface, the memory is used to store instructions, and the multiple processors are used to run the instructions stored in the memory, so that the electronic device executes any one of the methods in the first aspect.
[0021] Fourthly, a system is provided, which is applied to an electronic device. The chip system includes a first processor and a second processor. The first processor is used for performing breath wake-up detection, and the second processor is used for responding to voice interaction for the voice data collected by the second processor under the wake-up of the first processor; at a first moment, the electronic device is lifted and the screen of the electronic device is in an off state. The first processor triggers the power consumption of the electronic device to change from a first value to a second value. Before the first moment, the screen of the electronic device is in an off state and the electronic device is not lifted, and the second value is greater than the first value; at a second moment after the first moment, the electronic device is not lifted and the screen of the electronic device is in a screen-off state, and the power consumption of the electronic device changes from the second value to the first value.
[0022] Optionally, the chip further includes a memory, and the memory is electrically connected to the processor.
[0023] Optionally, the chip may further include a communication interface.
[0024] Fifthly, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by an electronic device, any method in the first aspect can be implemented.
[0025] Sixthly, a computer program product is provided. The computer program product includes a computer program, and when the computer program is executed by an electronic device, any method in the first aspect can be implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figures 1 - 3 is a schematic interaction diagram of a voice interaction method according to an embodiment of the present application.
[0027] Figure 4 is a schematic software architecture diagram of an electronic device according to an embodiment of the present application.
[0028] Figure 5 is a schematic diagram of the execution flow of a breath wake-up module during a voice interaction according to an embodiment of the present application.
[0029] Figure 6 is a schematic flowchart of a voice interaction method according to an embodiment of the present application.
[0030] Figure 7 is a schematic diagram of the execution flow of a breath wake-up detection according to an embodiment of the present application.
[0031] Figure 8 is a schematic flowchart of a voice recognition stage according to an embodiment of the present application.
[0032] Figure 9 is a schematic flowchart of a dialogue stage according to an embodiment of the present application.
[0033] Figure 10 It is a power consumption comparison diagram between the solution of the embodiment of the present application and the traditional solution in the same interaction scenario.
[0034] Figure 11 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0035] The solution of the embodiment of the present application will be introduced below with reference to the accompanying drawings.
[0036] Voice interaction is a human-computer interaction method. An electronic device can perform corresponding actions based on voice commands input by a user, such as performing information queries, playing music, starting a certain function, etc. If a user wants to perform voice interaction with the electronic device, the voice interaction module of the electronic device needs to be awakened first. The ways to awaken the voice interaction module can include: wake-up by a shortcut key, wake-up by a wake-up word, and wake-up by breath. The wake-up by a wake-up word requires the user to input a specific wake-up word. The wake-up by breath allows the user to wake up the device or application program by emitting a certain voice (such as a voice command, a breathing sound, a coughing sound, or other non-specific vocabulary sounds) without inputting a specific wake-up word. Among them, the wake-up by breath can also be called wake-up without a wake-up word, and the present application does not specifically limit its name.
[0037] The wake-up by breath needs to continuously detect or periodically check a signal (such as a voice signal) to determine whether to wake up the voice interaction module. When the wake-up by breath requires the participation of a main processor such as an application processor (AP), even a low-frequency periodic check may frequently wake up the main processor, which conflicts with the sleep state of the main processor and causes an increase in the power consumption of the electronic device. In order to realize waking up the voice interaction module by the wake-up by breath when the electronic device is in the screen-off state (the screen of the electronic device is in the off state), that is, the main processor (such as AP) enters the sleep state, and to ensure that the power consumption of the electronic device is relatively low, the breath wake-up module for realizing the wake-up by breath in the embodiment of the present application is deployed in the low-power space of other processors (non-main processors), such as the low-power memory space. The design of the low-power space allows specific tasks, such as processing audio data or sensor data, to be continued when other parts of the system (such as AP) are in the sleep state, thereby reducing the power consumption of the electronic device.
[0038] Exemplarily, other processors can be a coprocessor, also known as an auxiliary processor. Whether the main processor (AP) is in a sleep state or not, the coprocessor can keep working. For example, the coprocessor can be an audio digital signal processor (ADSP), or a system control processor (SCP), or a sensor hub chip. When the user wakes up by voice, other processors can wake up the AP to perform subsequent operations.
[0039] When the low-power memory space of the ADSP in the chip platform is large enough, the breath wake-up module can be fully deployed in the low-power memory space of the ADSP. However, if the low-power memory space of the ADSP in the chip platform is insufficient, it is impossible to fully deploy the breath wake-up module in the low-power memory space of the ADSP. That is to say, limited by the finiteness of the low-power memory space of the ADSP, the breath wake-up method is difficult to be universalized, that is, it is difficult to be extended to electronic devices with too small low-power memory space. Taking a mobile phone as an example, mobile phones with high chip configurations generally have relatively high ADSP configurations and sufficient low-power space, and can deploy a complete breath wake-up module without pressure. However, other mobile phones with ordinary chip configurations cannot deploy it. This undoubtedly results in that the breath wake-up solution can only be applied to high-configured mobile phones and cannot be popularized to all mobile phones. To solve this problem, a solution has emerged in which after the breath wake-up module is functionally divided, a part of it is deployed in the low-power memory space, and the remaining part is deployed in the non-low-power space of the ADSP. In this solution, the gesture detection module in the low-power memory space of the ADSP is mainly used to continuously perform gesture detection in the low-power working mode in the screen-off state, and only after detecting a raising hand gesture will it wake up the relevant modules in the non-low-power memory space of the ADSP to perform the relevant steps of breath wake-up detection, including collecting voice signals and judging whether to wake up the main processor AP. Otherwise, it will return to the low-power working mode and continue to only perform gesture detection in the low-power mode. The above solution of gradually waking up the ADSP and the AP step by step enables the breath wake-up solution to be deployed on electronic devices with mid-range and low-end chips (with relatively small low-power memory space), improving the universality of breath wake-up.
[0040] However, the breath wake-up solution that emerged to improve universality still increases extra power consumption to a certain extent. That is, the ADSP needs to perform operations such as detection and discrimination of breath wake-up in the normal working mode (non-low-power working mode). Especially to ensure the accuracy of detection, after waking up the ADSP, it needs to be kept for a long enough time to ensure accurate collection of audio signals and discrimination and waking up the AP. This results in that the power consumption during this waiting time is also inevitable.
[0041] In order to further reduce power consumption on the premise of ensuring the universality of the breath wake-up solution, the solution of this application came into being. Through analysis, it is found that some incorrect wake-ups cause unnecessary power consumption of the ADSP. For example, when the user may flip the electronic device, pick up and immediately put down the electronic device, or pick up the electronic device but not bring it close to the mouth, etc., a lift gesture will be detected, thus waking up the ADSP, and it is also necessary to wait for a sufficient duration to collect audio signals and perform subsequent discrimination. In fact, in the above scenarios, the user does not intend to perform breath wake-up. By analyzing the commonalities of these incorrect wake-up situations, a solution is found, that is, to "correct" the detection result of the lift gesture, so as to eliminate and terminate the subsequent steps of the above incorrect wake-up in a timely manner, effectively reducing power consumption, and thus the solution of this application came into being.
[0042] To facilitate the understanding of the solution of this application, the above low-power mode, sleep state, etc. will be illustrated by examples below. Taking the main processor as the AP for example, the AP entering the sleep state can be understood as the AP entering the low-power mode, which can reduce energy consumption and extend the usage time of the electronic device. When the AP enters the sleep state, it includes but is not limited to the following situations: 1. The clock frequency of the AP decreases or stops. The clock is the basis for the processor to execute instructions. Reducing or stopping the clock frequency can reduce the power consumption of the processor. 2. In order to match the reduced clock frequency, the supply voltage of the AP decreases accordingly. 3. Some circuits are turned off or paused, such as the cache, graphics processor, etc.
[0043] For an electronic device with an indicator light, when the AP enters the sleep state, the indicator light of the electronic device can flash in a specific manner or remain constantly on / off. For an electronic device with a display screen, when the AP enters the sleep state, the screen of the electronic device may be turned off or display a specific sleep screen saver.
[0044] Taking the main processor as the AP for example, all operations executed by the AP can be implemented by the main processor. The situations where the main processor enters the sleep state or the low-power mode can include the situations where the above AP enters the sleep state or the low-power mode. This application embodiment only takes the AP as an example for illustration, and does not mean that this application is limited thereto.
[0045] The electronic device provided by this application embodiment is equipped with a voice interaction module. For example, the electronic device installs a voice assistant (such as YOYO, etc.) application program that can provide voice interaction functions. The user wakes up the voice assistant to achieve voice interaction between the user and the electronic device. Waking up the voice assistant installed on the electronic device also means waking up the voice interaction module of the electronic device.
[0046] The "voice assistant" involved in the embodiments of this application can also be referred to as the "intelligent assistant", and can also be referred to as the "digital assistant", "virtual assistant", "intelligent automation assistant" or "automatic digital assistant", etc. The "voice assistant" can be understood as an information processing system that can recognize natural language inputs in voice form and / or text form to infer the user's intention and perform corresponding actions based on the inferred user intention. The system can output a response to the user's input in an audible (e.g., voice) and / or visual form.
[0047] The electronic device provided in the embodiments of this application is equipped with a breath wake-up module for waking up the voice interaction module. For example, the voice assistant includes a voice wake-up module and a breath wake-up module. When the voice wake-up module is turned on, the user can wake up the intelligent assistant by using the wake-up word wake-up method. The electronic device wakes up the intelligent assistant in response to receiving a specific wake-up word. When the breath wake-up module is turned on, the user can wake up the intelligent assistant without using a wake-up word. The user can wake up the intelligent assistant by the breath wake-up method. The electronic device receives the audio data sent by the user in response to the electronic device being in a lifted state or in response to a hand-raising event occurring, and wakes up the intelligent assistant based on the audio data.
[0048] Among them, the breath wake-up module can also be called the "wake-up word-free wake-up module" or the "direct dialogue wake-up module", etc. This application does not make specific limitations on this.
[0049] Figures 1 - 3 It is a schematic interaction diagram of a voice interaction method according to an embodiment of this application. As Figure 1 shown, a desktop interface 101 is displayed on the electronic device. The desktop interface includes icons of various applications (apps), such as icons of the phone app, contacts app, settings app, and messages app. However, it should be understood that this application does not limit the number and specific icons included in the interface 101, nor the positions of the icons in the interface.
[0050] The electronic device can be a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a home device, a smart wearable device, etc. This application does not impose any restrictions on the specific type of the electronic device, but the user needs to be able to perform a lift gesture operation on the electronic device, and the electronic device can detect the lift gesture, and the electronic device also needs to support intelligent voice interaction. The electronic device can be a device running the Android system, the IOS system, the Windows system or other operating systems. This application does not limit the type of the electronic device for executing the voice interaction method provided by this application and the operating system run by the electronic device, as long as it can be used to execute the voice interaction method of this application.
[0051] Suppose the user clicks on the icon of the Settings app in interface 101. In response to this click operation, the electronic device displays interface 102, which is the main interface of the Settings app. Interface 102 contains option bars for various services, such as WLAN, Bluetooth, mobile network, notification and status bar, desktop and wallpaper, and intelligent assistant, etc.
[0052] Suppose the user clicks on the intelligent assistant option bar (Option B) in interface 102. In response to this click operation, the electronic device displays interface 103, which is the main interface of the intelligent assistant. Interface 103 includes the wake-up word "Hello, YOYO", and also includes controls (option bars) for various intelligent services and various intelligent interactions. The intelligent services can include, for example, YOYO suggestions, the negative first screen, intelligent vision, etc., mainly some services that can use various models to provide intelligent services for users. The intelligent interaction can include intelligent voice, and the intelligent interaction can mainly give intelligent feedback during the human-computer interaction process. The intelligent voice control here is the management option corresponding to the voice interaction of this application.
[0053] Suppose the user clicks on the intelligent voice option bar (Control C) in interface 103. In response to this click operation, the electronic device displays interface 104. Interface 104 is the management interface of the intelligent voice. As shown in interface 104, this interface includes a schematic diagram of the human-computer interaction of the intelligent voice interaction, and also includes a functional explanation of the intelligent voice "Through natural voice conversations, help you use the device efficiently and answer your questions", and also includes various controls, and these controls respectively correspond to a specific type of intelligent voice. Here, taking interface 104 including a voice wake-up control, a breath wake-up control, a button wake-up control, a desktop shortcut control, and a voice broadcast tone control, etc. as an example, in practice, it can also include other controls, or reduce some of the above controls, without limitation.
[0054] Assume that the user clicks the breath wake-up control (Control D) in interface 104, and the electronic device responds to this click operation and displays Figure 2 the interface 105 as shown.
[0055] Interface 105 is the management interface for breath wake-up. This interface includes explanatory content, operation prompts, a switch control (i.e., Control E), etc. Among them, the operation prompt can show the user how to wake up the intelligent assistant through the breath wake-up function by means of animation. The operation prompt can also include text prompts, such as "Natural conversation, answer immediately without saying the wake-up word 'Hello YOYO'". The explanatory content of breath wake-up is used to introduce the operation method of the breath wake-up method. This explanatory content can, for example, include a text prompt: Lift the mobile phone, bring the bottom of the mobile phone close to the mouth (within 7 cm), and directly speak the voice command, such as "What's the weather like today" directly facing the bottom microphone. Thus, when the breath wake-up module is running, the user can wake up the intelligent assistant by directly outputting a voice command without having to output the wake-up word first. It should be understood that the above is only an example of interface 105 and there is no specific limitation. For example, the distance close to the mouth can be 5 centimeters (cm) or 6 centimeters or other distance thresholds. As long as the intensity difference of the audio signals received by the two microphones at the top and bottom of the electronic device is large enough when the user is speaking to the electronic device, and the signal intensity of the audio signal received by the microphone close to the mouth is large enough, other situations will not be listed one by one.
[0056] As shown in interface 105, the initial state of the breath wake-up switch control (Control E) is off. Assume that the user clicks Control E in interface 105, and the electronic device responds to this click operation and displays interface 106. It can be seen that other contents in interface 106 are the same as those in interface 105, except that the state of Control E is the on state, thus enabling the breath wake-up function.
[0057] It should be understood that Figures 1 - 2 this is mainly an example of the interaction process of how to enable the breath wake-up function in the settings app, but in practice, the specific interaction process can also be Figures 1 - 2 not exactly the same. For example, in the case where the intelligent assistant service is created as a desktop shortcut icon or a desktop card, the user can directly enter the intelligent assistant interface 103 by clicking the above desktop shortcut icon or desktop card. Other situations will not be listed one by one.
[0058] After the electronic device enables the breath wake-up function, the user can directly wake up the intelligent voice assistant in the screen-off state by means of breath wake-up, and thus perform voice interaction. The following takes Figure 3 as an example, but it should be understood that Figures 1 - 2The steps of setting the on / off state of the breath wake-up function do not have to be executed before each breath wake-up. Instead, each breath wake-up can be executed on the premise that the breath wake-up function is turned on.
[0059] Figure 3 It is a schematic interaction diagram of a voice interaction method according to an embodiment of the present application. Figure 3 The process shown needs to be carried out on the premise that the electronic device turns on the breath wake-up function.
[0060] As Figure 3 shown in (a) below, the electronic device is in the screen-off state, the AP processor is in the sleep state, and a microphone is provided at the top of the electronic device ( Figure 3 denoted by microphone M in the figure), and another microphone is provided at the bottom ( Figure 3 denoted by microphone N in the figure). However, it should be understood that it is also possible that the microphone at the bottom is denoted by microphone M and the microphone at the top is denoted by microphone N, without limitation.
[0061] Assume that the user raises the electronic device so that the microphone N at the bottom is close to the user's mouth. The electronic device responds to this raising gesture and starts to collect the voice signal input by the user. The raising gesture can be collected by an acceleration sensor, a gyroscope, and / or an Inertial Measurement Unit (IMU) in the electronic device, etc.
[0062] The distance between the electronic device and the user's mouth can be kept within 0 - 7 cm, which is convenient for the microphones M and N of the mobile phone to accurately collect the user's audio data.
[0063] From Figure 3 the gesture state shown in (a) below to Figure 3 the gesture state shown in (b) below is an example of the raising operation. That is, the user can raise the electronic device through the raising gesture of the wrist, making the electronic device closer to the user's mouth. The motion data of the electronic device that can be collected by the sensors in the electronic device can include data such as the acceleration, tilt, impact, vibration, and rotation of the electronic device. As Figure 3 shown in (a) and (b) below, the electronic device is raised, and the electronic device is currently in the hand-held raised state (i.e., the raised state) or the electronic device has a raising event. Among them, the electronic device being in the hand-held raised state can be understood as the user holding the electronic device in the hand. It should be understood that Figure 3Figures (a) and (b) are just examples of the lifting gesture, and there is no limitation on the actual scenario of the lifting gesture. As long as the lifting gesture brings the electronic device closer to the user's mouth, it is acceptable. Since voice activation by breath necessarily requires getting close to the microphone to speak, regardless of the initial state of the electronic device, there will inevitably be a lifting movement process that brings the microphone of the electronic device closer to the user's mouth. Therefore, the electronic device only needs to detect this posture change to determine whether the lifting gesture has occurred.
[0064] As Figure 3 As shown in Figure (b), assume that after the user brings the microphone N closer to the mouth by the lifting gesture and says the voice "How long have I exercised today", the electronic device lights up the screen and displays interface 107 in response to this voice input. It should be understood that voice activation by breath is performed in the screen-off state. Therefore, after the electronic device detects the lifting gesture, it needs to first collect voice through two microphones, and then determine whether voice activation by breath is required based on this voice. When it is determined that voice activation by breath is to be performed, the AP processor is awakened, the screen is lit up, and subsequent responses are made.
[0065] Since the positions of microphone M and microphone N on the electronic device are different, the distances between microphone M and the sound source (user's mouth) and between microphone N and the sound source (user's mouth) are different, the signal intensities of the voices received by microphone M and microphone N may be different, and there may be an energy difference between the voice signals received by microphone M and microphone N. When the user speaks to the top of the electronic device, the voice energy of the voice collected by the top microphone M will be higher than that of the voice collected by the bottom microphone N. Similarly, when the user speaks to the bottom of the electronic device, the voice energy of the voice collected by the bottom microphone N will be higher than that of the voice collected by the top microphone M.
[0066] It should also be understood that when making a call, one may also get close to a certain microphone to speak, but this situation occurs during a call, and during this period, the AP is in the working state (non-sleep state) after being awakened, and the electronic device will not perform detection and response for voice activation by breath. Therefore, this situation does not belong to the scenario involved in this application.
[0067] The ADSP determines whether the voice activation by breath event is triggered. When it detects that the voice activation by breath event is triggered, the ADSP wakes up the AP, and the AP starts or resumes the intelligent assistant. When the voice activation by breath event is not triggered, the ADSP does not wake up the AP, nor does it start or resume the intelligent assistant.
[0068] For example, when the electronic device is lifted, after the electronic device collects the corresponding voice when the user speaks through microphone M and microphone N, and the motion data collected by the sensor, it sends the collected voice and motion data to the ADSP. The ADSP detects the input voice data and motion data through the voice activation by breath module to determine whether the voice activation by breath event is triggered.
[0069] As Figure 3 shown in interface 107, when a trigger breath wake-up event is detected according to the input voice data and / or motion data, the ADSP wakes up the AP, and the electronic device turns on the screen and displays the wake-up interface 107 of the intelligent assistant. The wake-up interface 107 includes the text content converted from the input voice data and the control F that prompts the user that the intelligent assistant has been started. As Figure 3 shown in (b) therein, the audio data input by the user is "How long did I exercise today", and the corresponding text content in interface 107 is "How long did I exercise today". The user can click on the control F to trigger the intelligent assistant to end the speech recognition.
[0070] In some embodiments, the wake-up interface 107 may further include the interface content before the intelligent assistant is woken up. For example, the interface 107 further includes a part of the content of interface 106.
[0071] After recognizing the text content of the voice data input by the user, the electronic device can respond. Here, since "How long did I exercise today?" is an intelligent question and answer for the sports health app, the electronic device can, in response to this text content, dispatch the intelligent agent of the sports health app and its corresponding question and answer large model to infer the corresponding answer and display it, as shown in interface 108. In interface 108, the user's question "How long did I exercise today?" and the corresponding answer "You have walked for 15 minutes and run for 30 minutes today, for a total of 45 minutes of exercise. From the sports health app" are displayed in the form of a conversation. It should be understood that interface 108 is only an example of how to respond after voice recognition. In practice, it can also be presented in other forms or make other corresponding responses, without limitation. For example, interface 108 may not include the character image of the breath wake-up demonstration. Another example is that it may not be in the form of a dialogue bubble, but in a list form. Another example is that if the question is not for the sports health app, such as a question corresponding to other apps like weather, other content can be displayed in other forms. Another example is that if the user inputs "Help me open the gallery", interface 108 can display "Help me open the gallery", as well as the feedback "Okay, about to help you open the gallery app", and automatically jump to the display interface of the gallery app after the display is complete. Or, under this premise, instead of displaying interface 108, directly display the running interface of the gallery app and display "Okay, have helped you open the gallery" in the form of voice broadcast or light prompt. Another example is that if the user inputs "Call Zhang San", interface 108 can display "Call Zhang San", as well as the feedback "Okay, have found Zhang San's phone number 123456 for you from the contacts, about to open the phone app", and automatically jump to the display interface of the phone app and display Zhang San's phone number. Or, under this premise, instead of displaying interface 108, directly display the running interface of the phone app and display Zhang San's phone number, and voice broadcast "Okay, have helped you input Zhang San's phone number. Please confirm whether to call this number, thank you". Other situations will not be listed one by one.
[0072] When the ADSP confirms the need to wake up the intelligent assistant, the ADSP generates a wake-up signal and transmits the wake-up signal to the AP. For example, when the ADSP detects that the breath wake-up event is triggered, the ADSP generates a wake-up signal. Another example is that when the ADSP detects the wake-up word, the ADSP generates a wake-up signal. The ADSP can send the wake-up signal to the AP through the internal communication interface of the system, such as I2C, SPI or a dedicated wake-up line.
[0073] If the AP was previously in a low-power mode such as sleep state, it will resume to the normal operating mode during the wake-up process. Exemplarily, after receiving the wake-up signal, the AP will wake up or configure the necessary system resources, including but not limited to: adjusting the power management policy, increasing the clock frequency, enabling the necessary hardware interfaces, etc. For example, after the power management unit (PMU) of the AP recognizes the wake-up signal, the PMU changes the power state of the AP and provides the required voltage and current to the AP. For example, it activates the power converter, adjusts the voltage level or increases the supply current. Another example is that the clock management unit (CMU) of the AP adjusts the clock frequency.
[0074] After the AP receives the wake-up signal, the AP checks the status of the intelligent assistant and performs corresponding processing according to the status of the intelligent assistant.
[0075] If the intelligent assistant has not been started yet, the AP starts the intelligent assistant service. Starting the intelligent assistant service by the AP may include but not limited to: initializing resources, loading services, starting services, preparing the user interface, initializing services and components. Exemplarily, initializing resources may include but not limited to: the AP allocates necessary resources to the intelligent assistant service, including the central processing unit (CPU) time slice, memory, and possible input / output (I / O) resources. Exemplarily, loading services may include but not limited to: the AP loads the code and resources of the service from storage into memory. Exemplarily, starting services may include but not limited to: the AP starts the entry point of the intelligent assistant service, such as the main Activity or a specific starting Service. Exemplarily, preparing the user interface may include but not limited to: the AP prepares and displays the user interface so that the user can see the status and feedback of the service.
[0076] If the intelligent assistant is already running in the background but is in a dormant or low-power state, the AP wakes up the intelligent assistant and restores it to the normal working state. The AP waking up the intelligent assistant and restoring it to the normal working state may include, but are not limited to: waking up the intelligent assistant, allocating computing resources, restoring services, and updating the user interface. Exemplarily, waking up the intelligent assistant may include, but is not limited to: the operating system converting the state of the intelligent assistant from a dormant or low-power state to a normal working state. Exemplarily, allocating computing resources may include, but is not limited to: the AP reallocating necessary computing resources such as the CPU and memory to ensure that the intelligent assistant can operate normally. Exemplarily, restoring services may include, but is not limited to: the AP restarting the services that stopped due to low power. Exemplarily, updating the user interface may include, but is not limited to: if the intelligent assistant has a user interface, the AP ensures that the interface can be normally displayed after waking up and updates any pending notifications or information. The AP restores the interaction functions between the user and the intelligent assistant, such as clicking, swiping, voice input, etc., so that the user can continue to use the intelligent assistant for various operations.
[0077] If some functions of the intelligent assistant are running in the background, the AP restores the complete functions of the intelligent assistant, including but not limited to: user interface display, background services. Exemplarily, restoring the main interface display may include, but is not limited to: the main interface of the intelligent assistant (such as the negative first screen, sidebar, etc.) being normally displayed on the screen, restoring the interaction functions between the user and the intelligent assistant, such as clicking, swiping, voice input, etc., and restoring data synchronization, such as user preferences, settings, history records, etc. Exemplarily, restoring various services running in the background of the intelligent assistant may include, but is not limited to: notification reminder, health monitoring, security protection, etc. It should be understood that the above is illustrated by taking the voice interaction module as the intelligent assistant as an example. After the ADSP wakes up the AP, the AP can refer to the above for waking up the voice interaction module. If the voice interaction module has not been started, the AP starts the voice interaction module. The AP starting the voice interaction module may include, but is not limited to: initializing resources, loading application programs, starting application programs, preparing the user interface, initializing services and components, and the specific content can be referred to the above. If the voice interaction module is already running in the background but is in a dormant or low-power state, the AP wakes up the voice interaction module and restores it to the normal working state, and the specific content can be referred to the above. If some functions of the voice interaction module are running in the background, the AP restores the complete functions of the voice interaction module, including but not limited to: user interface display, background services.
[0078] It should be understood that the above is an example illustration of the application scenario and does not make any limitation to the application scenario of the present application.
[0079] It can be understood that the visible interface elements such as the text and controls displayed in the above interface are only examples, and the present application does not limit parameters such as the display position, display style, and display size of the interface.
[0080] During the above-mentioned breath wake-up process, when the electronic device is in the screen-off state, the AP is in the sleep state, and the ADSP works in the low-power mode. It needs to continuously perform gesture detection, and switch to the normal working mode after detecting the raise hand gesture to collect audio signals, and then determine whether to trigger the breath wake-up, so as to decide whether to wake up the AP. After waking up the AP, the ADSP also needs to cache and forward the audio signal to the AP so that the AP can perform subsequent voice recognition and voice interaction responses. The ADSP has a low-power space and a non-low-power space. Since the space capacity of the low-power space is limited, the above-mentioned breath wake-up detection and triggering will also be subdivided into whether each step should be executed in the low-power space or the non-low-power space. Especially steps that require a large amount of memory space such as breath wake-up detection and audio data access and storage need to be arranged in the non-low-power space as much as possible, while gesture detection and starting breath wake-up can be arranged in the low-power space. Therefore, the actual deployment scenario is that when the electronic device is in the screen-off state, the ADSP continuously performs gesture detection in the low-power working mode. When the raise hand gesture is detected, it needs to switch to the normal working mode and continue for a period of time (such as 2 seconds, 3 seconds, etc.) to collect voice signals, and also determine whether to trigger the breath wake-up function based on the collected voice signals. After determining to trigger the breath wake-up function, subsequent steps are executed. This brings new problems. Because the posture of the electronic device in the screen-off state changes very frequently, it may cause the ADSP to switch to the normal working mode for a period of time each time the posture changes to try to collect voice signals for a sufficient duration and leave enough pause time for the user to input voice before, such as the fixed durations of 2 seconds, 3 seconds, etc. mentioned above. During the collection of voice signals, the ADSP is in the normal working mode, and the modules in both the non-low-power space and the low-power space are working, resulting in additional power consumption. The above-mentioned posture changes that do not trigger breath wake-up can include, for example, that the user may not speak after raising the electronic device (such as just breathing on the screen to wipe the screen), or the user may just flip the electronic device, or the user may pick up the electronic device and then quickly put it down. Other situations are not listed one by one. When the above actions occur, the ADSP of the electronic device will start the breath wake-up detection process because it detects the raise hand gesture, and start waiting for a period of time. At any time point during this waiting period, when a voice signal is received, it starts to continuously collect voice signals until the user stops input, and determines whether to trigger the breath wake-up function based on the collected voice signals. This results in that under the trigger of the above-mentioned posture changes that do not trigger breath wake-up, the ADSP wastes this part of the power consumption after switching to the normal working mode.
[0081] In view of the above problems, the present application provides a new voice interaction method. When collecting voice signals after detecting a raising hand gesture, it will continue to read the real-time gesture detection results and determine whether to terminate breath wake-up based on the new gesture detection results. This solution mainly decides whether to continue the detection process by combining the real-time gesture detection results after triggering the detection process of breath wake-up, so as to directly terminate the breath wake-up when a gesture that does not conform to breath wake-up occurs, without collecting audio signals and subsequent judgments. In the traditional solution, after detecting a raising hand gesture, the ADSP will be awakened to work in the normal working mode (not the low-power working mode) and last for a fixed duration (a preset fixed duration, such as 2 seconds or 3 seconds, etc.) to reserve enough pause time for the user to input voice, and once an audio signal is collected within this fixed duration, it will attempt to continuously collect audio signals and make subsequent judgments and awaken the AP based on the collected audio signals. These steps, including this fixed duration and the collection and subsequent judgments of audio signals, will still be continuously executed even when the above non-breath wake-up gesture occurs, resulting in waste of power. In response to this problem, the present application will continue to obtain real-time gesture detection results when starting to collect voice signals, so that unnecessary "idle running" processes can be terminated in a timely manner based on the new gesture detection results. That is to say, assuming that the above non-breath wake-up gesture is encountered again, it is not necessary to wait until the end of the fixed duration countdown, and it can be terminated in advance, thereby further reducing power consumption.
[0082] To facilitate the understanding of the solution of the present application and the traditional solution, the following will be described with specific examples. Scenario 1: Assume that the preset fixed duration is 2 s, the electronic device is a mobile phone, and the mobile phone is placed stationary on the table and in the screen-off state. At this time, the ADSP of the mobile phone operates in the low-power mode and continuously performs gesture detection. Assume that the user picks up the mobile phone on the table and puts it into the pocket while chatting with someone. Then, during the picking-up process, the gesture detection module in the low-power space of the ADSP can detect the hand-raising gesture. In response to this hand-raising gesture, the mobile phone wakes up the ADSP and makes it switch to the normal working mode and maintains the normal working mode for at least 2 seconds. The breath wake-up detection module in the non-low-power space will collect audio signals. In the above scenario, the mobile phone will surely start to collect the user's audio signals within the 2-second waiting duration. Assume that the audio signals start to be collected 1 second after the ADSP is woken up and are collected completely 4 seconds after being woken up. Then, in addition to the power consumption in the low-power space of the ADSP for more than 4 seconds, the power consumption within these 4 seconds is also increased, including the power consumption during the 1-second waiting duration and the power consumption for collecting audio signals for 3 seconds, as well as the computing power consumption when performing the subsequent step of determining whether the audio signals can be used to wake up the AP. Due to the above scenario, the user does not bring the mobile phone microphone close to the mouth, so the 3-second-long audio signals collected may be ordinary chatting sounds and do not meet the energy intensity requirements of the breath wake-up audio signals (the maximum intensity of the audio signals collected by the two microphones is still not enough, or the intensity difference between the audio signals of the two microphones in standby is not enough), thus failing to wake up the AP. Or, in the case where the AP can be woken up, since the input voice signals do not involve the user's intention after being recognized (such as intentions like answering questions or opening apps), this results in an invalid process starting from when the complete audio signals are collected.If the solution of the present application is adopted, in the above scenario, assume that the user picks up the mobile phone on the table and kicks it into the pocket while chatting with someone. Then, during the picking-up process, the gesture detection module in the low-power space of the ADSP can detect the hand-raising gesture. In response to this hand-raising gesture, the mobile phone wakes up the ADSP and makes it switch to the normal working mode, and maintains the normal working mode for at least 2 seconds. The breath wake-up detection module in the non-low-power space will collect audio signals, and the breath wake-up detection module in the non-low-power space continues to obtain real-time gesture detection results. In the above scenario, assume that when a new real-time gesture detection result is obtained at the 0.8th second after the ADSP is awakened, the new gesture detection result is a non-raising gesture (the posture of the mobile phone changes again due to kicking it into the pocket and is no longer in the raised state). Then, even in the case where the voice signal is collected starting from the 1st second as described above, the ADSP terminates the breath wake-up based on the new gesture detection result at the 0.8th second, and no longer needs to execute subsequent processes such as collecting audio signals. The actual power consumption includes the continuous power consumption of the ADSP in the low-power space for 0.8 seconds, and the power consumption of waiting to collect voice signals in the non-low-power space during these 0.8 seconds. It can be seen that in the same scenario, it is much lower than the power consumption of adopting the traditional solution. Scenario 2: Assume that the preset fixed duration is 2s, the electronic device is a mobile phone, and the mobile phone is in the screen-off state. At this time, the ADSP of the mobile phone works in the low-power mode and continuously performs gesture detection. Assume that the user flips the mobile phone and the user does not make any sound throughout the process, and the mobile phone is not close to the mouth. Also, assume that the user's flipping of the mobile phone is an instantaneous and fast action, and assume that the entire flipping action lasts for 0.2 seconds, with the first 0.06 seconds for raising the mobile phone and the last 0.14 seconds for flipping the mobile phone over. Then, during the process of raising the mobile phone before flipping it over, the gesture detection module in the low-power space of the ADSP can detect the hand-raising gesture. In response to this hand-raising gesture, the mobile phone wakes up the ADSP and makes it switch to the normal working mode, and maintains the normal working mode for at least 2 seconds. The breath wake-up detection module in the non-low-power space will collect audio signals. In the above scenario, since the user's flipping of the mobile phone is an instantaneous and fast action and the user does not make any sound, the ADSP will not collect any audio signals within the 2 seconds of maintenance. After 2 seconds, since no audio signals are collected, the ADSP stops the breath wake-up and switches back to the low-power working mode, only performing gesture detection. During this process, the power consumption of the ADSP includes, in addition to the power consumption for 2.06 seconds (gesture detection in the low-power space during the 2 seconds of maintaining the normal working mode and gesture detection in the low-power space in the low-power mode for 0.06 seconds), the power consumption of attempting to collect audio signals in the non-low-power space during the 2 seconds of maintenance.In the same scenario, if the solution of the present application is adopted, the ADSP will be awakened after the mobile phone detects a raising gesture at 0.06 seconds, but continue to receive the gesture detection results. Then, at 0.14 seconds in the normal working mode of the ADSP, the breath awakening will be directly terminated because a new gesture detection result (no longer a raising gesture) is received, and there is no need to wait for the 2-second countdown. During this process, the power consumption of the ADSP includes not only the power consumption in the low-power space for about 0.2 seconds, but also the power consumption in the non-low-power space for about 0.14 seconds. It can be seen that the power consumption is much lower than that in the same scenario when the traditional solution is adopted.
[0083] The above example is only to illustrate how the solution of the present application effectively reduces power consumption. There is no limitation on the specific values, and it is only for more intuitively comparing the reduction of power consumption. Other situations will not be listed one by one. In short, the solution of the present application continues to receive gesture detection results in the normal working mode of the ADSP, and determines whether to terminate the breath awakening based on the real-time gesture detection results, so as to eliminate the invalid "idle running" process caused by the raising gesture of non-breath awakening, thereby effectively reducing power consumption.
[0084] Figure 4 It is a schematic diagram of the software architecture of an electronic device according to an embodiment of the present application. As Figure 4 shown, the electronic device includes an application processor (AP) and an audio digital signal processor (ADSP). The AP can be an Android system and can include an application layer (APP layer), a framework layer (framework, FWK), and a hardware abstraction layer (hardware abstraction layer, HAL), etc.
[0085] The application layer may include applications (APPs). The application layer includes all the applications installed on the electronic device, which may include voice interaction applications, such as intelligent assistants. The intelligent assistant is used to provide voice interaction functions. The intelligent assistant can parse the user's voice commands and perform relevant operations according to the user's voice commands, so as to realize the voice interaction between the electronic device and the user. The framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The framework layer includes some predefined functions. The framework layer may include an audio trigger module (sound trigger, ST), an audio policy service module, an audio recording module, etc. The audio trigger module in the framework layer is used to control the startup of the voice interaction application in the application layer. The audio policy service module is used to establish a voice recognition channel with the voice interaction application. The audio recording module is used for audio recording. In some embodiments, the audio recording module can be implemented as AudioRecord, which is an API class provided by the Android framework layer for audio recording. The framework layer may also include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc. The hardware abstraction layer is the layer between the hardware and the software. The hardware abstraction layer may include an audio trigger module, etc. The audio trigger module in the hardware abstraction layer (HAL layer) can also be called an audio trigger driver (sound trigger hal, st-hal), and the hardware abstraction layer may also include an audio trigger pre-driver (primary-hal). The audio trigger module in the hardware abstraction layer is responsible for audio trigger functions, such as starting the breath wake-up recognition notification. The audio trigger module in the hardware abstraction layer is used to process the startup breath wake-up recognition notification and the hardware recognition of other sounds. Multiple drivers for driving the hardware to work may be installed in the hardware abstraction layer. It should also be understood that the above audio trigger driver (st-hal) and audio trigger pre-driver (primary-hal) can also be collectively referred to as an audio trigger module for triggering audio functions.
[0086] It should be noted that the application layer, the framework layer, and the hardware abstraction layer may also include other contents, which are not specifically limited herein. The embodiments of the present application only take the Android system as an example for illustration. In other operating systems (such as Windows system, IOS system, etc.), as long as the functions implemented by each functional module are similar to those in the embodiments of the present application, the solution of the present application can also be implemented.
[0087] Such as Figure 4As shown, the AP communicates and / or is electrically connected to the ADSP, and the ADSP communicates and / or is electrically connected to hardware devices such as microphones and sensors. The microphone may include a first microphone and a second microphone, and the first microphone and the second microphone may be respectively disposed at different positions of the electronic device, such as the top and the bottom. The sensor may include at least one of an acceleration sensor, a gyroscope, and an inertial detection unit, etc., for gesture detection.
[0088] The ADSP may include a low-power space (low-power memory space) and a non-low-power space (non-low-power memory space). The low-power memory space and the non-low-power memory space can be understood as different memory regions divided according to power consumption optimization requirements. The low-power memory space uses power consumption optimization techniques to reduce power consumption, while the non-low-power memory space focuses more on performance improvement but may have higher power consumption. The low-power memory space of the ADSP generally refers to the memory region designed to operate in the low-power mode of the ADSP. It can remain active when some parts of the processor enter the low-power state. The design of the low-power memory space of the ADSP allows specific tasks, such as processing audio data or sensor data, to continue to be processed when other parts of the system (such as the AP) are in the sleep state, thereby reducing the overall power consumption. The non-low-power memory space of the ADSP can be an ordinary memory region. They provide data storage and processing functions during system activity, but may not remain active when the ADSP enters the low-power mode. The non-low-power memory space of the ADSP can be used for services that need to continue running after the AP (application processor) enters the sleep state, such as voice wake-up, sensor data processing, step counting, etc. The low-power memory space of the ADSP can be the secondary cache of the general-purpose processor (CPU). The secondary cache can provide fast data access, and the secondary cache can be implemented by a low-power memory to reduce power consumption. The low-power memory space of the ADSP can be specifically implemented by a low-power memory. The low-power memory has low-power characteristics and is suitable for specific services that need to continue running when the AP is in the sleep state. The low-power memory may include, but is not limited to, static random-access memory (SRAM). As long as the SRAM remains powered on, the data stored in it can be constantly maintained. The SRAM does not require a refresh circuit to store the data inside it, which makes its power consumption relatively low during the working process. Due to the static storage characteristics of the SRAM, its power consumption is low when it is not performing read / write operations, and it is very suitable for the low-power memory space with strict power consumption requirements. In the low-power memory space of the ADSP, using SRAM can ensure stable data storage while reducing the overall power consumption.
[0089] Such as Figure 4As shown, the low-power memory space of the ADSP includes an audio area and a sensing area. The audio area and the sensing area can be understood as different working areas divided according to business requirements. The audio area is mainly responsible for audio-related services, such as tasks directly related to audio signals. The audio area can run algorithms related to audio processing. The sensing area is mainly responsible for tasks related to sensor data, such as the acquisition and processing of sensor data. The sensing area can run algorithms related to sensor data processing.
[0090] In the embodiment of the present application, the low-power memory space of the ADSP is divided into different physical areas, and the space for the sensor (i.e., the sensing area) and the space for the audio (i.e., the audio area) are separated. The sensing area can be implemented through a sensor low power island (slpi), and the audio area can be implemented through an audio low power island technology. Both the sensor low power island and the audio low power island can be implemented by using a low-power memory such as SRAM as the memory.
[0091] In the ADSP, the slpi can be understood as a hardware architecture specifically designed to process sensor data in the low-power mode. The slpi architecture is used to implement the low-power memory space related to sensor data, such as the sensing area. Specifically, it can include integrating modules that need to run continuously, such as the driver and algorithm of a specific service (such as sensor), into the slpi space, so that it can be powered separately by the slpi and continue to run when the AP is in the sleep state. In this way, the AP can enter the deep sleep state to reduce power consumption, while the slpi space is responsible for processing tasks that need to run continuously. Thus, the AP can enter the sleep state when not needed, thereby reducing the overall power consumption of the system.
[0092] In the slpi architecture, sensors (such as IMU, gyroscope sensors, and acceleration sensors, etc.), microphones, etc. can be electrically connected to the ADSP, so that the ADSP can process the data collected by the sensors. When the AP sleeps to reduce power consumption, the ADSP can process the sensor data, thereby achieving lower power consumption while ensuring that some functions are not affected by the AP's sleep.
[0093] Exemplarily, the above-mentioned sensors and microphones are mounted on the ADSP, such as the above-mentioned sensors and microphones are integrated in or connected to the same circuit board where the ADSP is located. In some embodiments, the driver programs of the sensors and microphones run on the ADSP. In other embodiments, the driver programs of the sensors and microphones run on the application processor (AP). The AP interacts with the sensors and microphones through a communication interface with the ADSP (such as I2C, SPI, UART, etc.).
[0094] In the ADSP, the audio low-power island is a functional area designed specifically for audio processing, used to process audio data in the low-power mode. It can continue to operate when the main processor is in the sleep state, thus reducing the power consumption of the entire system. This design allows audio data to be processed without the intervention of the CPU. For example, in scenarios such as waking up the voice interaction module, voice data can be directly transmitted without waking up the CPU.
[0095] The audio low-power island usually makes use of the hardware characteristics of the system. For example, it uses the low-power memory SRAM as the memory. Since SRAM does not need to be refreshed as frequently as dynamic random access memory (DRAM), its power consumption is relatively low. This space can be used to store service data that needs to continue to run after the AP goes to sleep, such as sensor data processing, step counting, voice wake-up, etc.
[0096] In the embodiments of this application, by placing some modules that need to run continuously in the breath wake-up module in the low-power memory space of the ADSP, it can be ensured that these modules can still operate normally when the AP is in the sleep state without having a significant impact on the overall power consumption. The partitioning of the low-power memory space of the ADSP allows audio and sensor data processing to still be handled by a dedicated DSP (such as the ADSP) when the main processor is inactive, thus ensuring low power consumption. In this way, the ADSP can continue to provide high-quality audio processing functions while maintaining low power consumption.
[0097] Optionally, place some modules that need to run continuously in the breath wake-up module in the low-power memory space of the ADSP, that is, place some code that needs to run continuously in the breath wake-up module in the low-power memory space of the ADSP.
[0098] It can be understood that taking the second memory as the low-power memory SRAM as an example, all the functions of the low-power memory SRAM can be implemented by the second memory. The embodiments of this application only take the low-power memory SRAM as an example for illustration, and do not mean that this application is limited thereto.
[0099] In the embodiments of this application, the non-low-power memory space of the ADSP can be specifically implemented by a non-low-power memory. The non-low-power memory can include but is not limited to: DRAM, double data rate synchronous dynamic random access memory (DDR SDRAM), or other types of memories.
[0100] DRAM is a dynamic random access memory, and the data stored in it needs to be refreshed periodically to maintain data integrity. In the non-low-power memory space of ADSP, since power consumption is not the primary consideration, DRAM with high integration and low cost can be used as the memory. The large capacity and relatively low power consumption of DRAM make it an ideal choice for the non-low-power memory space memory.
[0101] DDR is a synchronous dynamic random access memory with double data rate. It adopts double data rate transmission technology and can transmit data on both the rising edge and the falling edge of each clock cycle, thus achieving high-speed data transmission. DDR has a large storage capacity and supports multiple capacity specifications, which can meet the ADSP's demand for storing a large amount of data. DDR can provide large-capacity, high-speed and low-power data storage and transmission support for the non-low-power memory space.
[0102] It can be understood that the low-power memory space and the non-low-power memory space in ADSP can also be implemented by other memories. For example, the low-power memory space can also be implemented by low power double data rate synchronous dynamic random access memory (LPDDRSDRAM). According to the specific requirements and power consumption limitations of ADSP, appropriate memory types can be selected to implement the functions of the low-power memory space and the non-low-power memory space, and this application does not make specific limitations on this.
[0103] In the embodiment of this application, the low-power mode is a working mode designed by ADSP to reduce power consumption. When ADSP is in the low-power mode, in order to reduce the overall power consumption, the low-power memory space of ADSP will be selected. In the non-low-power mode, the non-low-power memory space of ADSP will be selected, and thus the best performance and bandwidth can be obtained.
[0104] Specifically, when executing the voice interaction method provided in the embodiment of this application and ADSP is in the low-power mode, the low-power memory space of ADSP can be used, rather than using the low-power memory space of ADSP, that is, the low-power memory in the low-power memory space is in the working state, while the non-low-power memory in the non-low-power memory space is in the low-power state. When executing the voice interaction method provided in the embodiment of this application and ADSP is in the non-low-power mode, the non-low-power memory space of ADSP can be used, and the low-power memory space of ADSP can also be continued to be used, that is, the relevant modules in the non-low-power memory space and the low-power memory space are all in the working state.
[0105] In the embodiments of the present application, the ADSP exiting the low-power mode and entering the low-power mode can be understood as a switching of the operating mode of non-low-power memories (such as DDR) in the non-low-power memory space. Taking the non-low-power memory DDR as an example, when the ADSP enters the low-power mode, the DDR is in a low-power state. When the DDR is in the low-power state, it can reduce power consumption by reducing the operating frequency, entering the low-power mode, etc. When the ADSP exits the low-power mode, the DDR enters the working state. When the DDR is in the working state, the DDR may operate at full speed or work according to the conventional memory access mode to meet the requirements of the ADSP for high performance and response speed.
[0106] In some embodiments, when the DDR is in the low-power state, the working current of the DDR is very small. The DDR may only need to maintain data without loss, but it cannot read or write. When the DDR is in the normal state, the working current of the DDR is very large. The working current of the DDR in the low-power state is less than the working current of the DDR in the normal state.
[0107] In some embodiments, when the DDR is in the low-power state, the working voltage of the DDR is very small. The DDR may only need to maintain data without loss, but it cannot read or write. When the DDR is in the normal state, the working voltage of the DDR is very large. The working voltage of the DDR in the low-power state is less than the working voltage of the DDR in the normal state.
[0108] It should be understood that the present application only takes DDR and DRAM as examples to illustrate the functions of non-low-power memories, but in practice, other non-low-power memories can also be used, without limitation.
[0109] In the above text, the non-main processor (other processors) is exemplified by the ADSP, and this processor includes a low-power space and a non-low-power space. If the solution of the present application is extended and applied to other similar scenarios, the ADSP can also be replaced by other non-main processors. For example, when similar improvements are desired in the field of image processing, the sensorhub processor, etc. can be used. However, when setting what functional modules and how to deploy them to the low-power space and the non-low-power space respectively, it will vary depending on the actual scenario to be applied, which will not be elaborated, and other cases will not be listed one by one.
[0110] Optionally, the low-power space of the ADSP can be a low-power memory space, and the non-low-power space of the ADSP can be a non-low-power memory space. The low-power memory space and the non-low-power memory space of the ADSP can be understood as different memory regions divided according to power consumption optimization requirements. The low-power memory space uses power consumption optimization techniques to reduce power consumption, while the non-low-power memory space focuses more on performance improvement, but may have higher power consumption. The relevant content of the low-power space, non-low-power space, low-power memory space, and non-low-power memory space of the ADSP can be referred to the above, and will not be elaborated here.
[0111] Exemplarily, the low-power space or the low-power memory space of the ADSP can select low-power storage technologies, such as low-power memories, such as low-power versions of static random access memories (SRAMs), or flash technologies with low-power characteristics. The relevant content of the low-power memory can be referred to the above, and will not be elaborated here.
[0112] Exemplarily, the non-low-power space or the non-low-power memory space of the ADSP can be implemented by non-low-power memories, such as DDR or DRAM, etc. The relevant content of the non-low-power memory can be referred to the above, and will not be elaborated here.
[0113] In the traditional breath wake-up scheme where the steps of waking up by breath are segmented and deployed in different power consumption spaces respectively, by placing the modules that need to run continuously in the breath wake-up module in the low-power memory space, it can be ensured that when the AP is in the sleep state, these modules can still operate normally without having a significant impact on the overall power consumption; by placing the remaining modules in the breath wake-up module in the non-low-power memory space to provide higher performance and bandwidth. The memory with high performance and bandwidth can process data faster, making the ADSP more rapid when processing large tasks, running complex programs, or executing multiple tasks. The improvement in speed can significantly improve the working efficiency of waking up the intelligent assistant, thereby enhancing the user experience.
[0114] Figure 4 This is an example of deploying different functional modules of the breath wake-up module in the ADSP in the above traditional scheme. Such as Figure 4 shown, the breath wake-up module of the electronic device includes a gesture detection module and a breath wake-up detection module. The gesture detection module is deployed in the low-power memory space of the ADSP, and the breath wake-up detection module is deployed in the non-low-power memory space of the ADSP. The gesture detection module is used to detect the motion state of the electronic device. Specifically, the ADSP runs the gesture detection module, obtains the motion data of the electronic device, and determines the motion state (i.e., posture change) of the electronic device according to the motion data of the electronic device.
[0115] Motion data may include, but is not limited to, IMU data. For example, motion data may include acceleration data in the x-axis direction, acceleration data in the y-axis direction, and acceleration data in the z-axis direction obtained by an acceleration sensor of an electronic device. For another example, motion data may include angular velocity data. Other cases are not listed one by one.
[0116] The gesture detection module may include a pose detection model. Taking the motion data as the acceleration data in the x-axis direction, the acceleration data in the y-axis direction, and the acceleration data in the z-axis direction of the electronic device as an example, since the absolute values of the acceleration data corresponding to the three coordinate axes of the x-axis, y-axis, and z-axis, d1, d2, and d3 can represent whether the electronic device is in a motion state, and p1, p2, and p3 can represent the amplitude of the motion of the electronic device, the difference values between d1 and p1, the difference values between d2 and p2, and the difference values between d3 and p3 can represent the motion state of the electronic device from other dimensions such as the smoothness of the motion. Therefore, the pose detection model can comprehensively judge whether the electronic device is lifted from multiple aspects such as whether the electronic device is in motion, the motion amplitude, and the smoothness of the motion through the above motion data, that is, judge whether the electronic device is in a lifted state or a lift event occurs, improving the accuracy of the judgment of the pose detection model.
[0117] The pose detection model can be various neural network models capable of performing pose detection, such as a trained convolutional neural network model or a deep learning model. It can be obtained as long as the neural network model is trained with sufficient labeled training data. The training data may include multiple sets of motion data and the pose labels corresponding to each set of motion data. It should also be noted that the solution of this application is for the detection of the raise hand gesture. Therefore, during the training of the pose detection model, only the detection of the raise hand gesture can be trained, and the poses of other non-raise hand gestures are marked as "non-raise hand gestures" to further reduce the requirements for the pose detection of the model. That is to say, the pose labels in the training data only include "raise hand gesture" and "non-raise hand gesture". After training with such training data (samples), the pose detection model will only recognize the raise hand gesture and the non-raise hand gesture, and there is no need to further subdivide the non-raise hand gesture into other gestures such as a flip gesture or a rotation gesture. Thus, to a certain extent, the model scale is further reduced and the detection accuracy is improved.
[0118] Since the pose detection model of this application is mainly used for gesture detection and is to be deployed in the low-power memory space of the ADSP, a lightweight neural network model can be selected to train the pose detection model and deploy it to the low-power memory space of the ADSP.
[0119] The gesture detection module will continuously perform gesture detection without interruption when the screen is off. When a raise hand gesture is detected, the breath wake-up detection module is triggered to perform subsequent detection.
[0120] The breath wake-up detection module is used to perform breath wake-up detection.
[0121] In the embodiment of the present application, when the gesture detection module detects that the electronic device is in the lifted state, it notifies the breath wake-up detection module to obtain motion data and / or voice data. The breath wake-up detection module determines whether to trigger the breath wake-up function according to the obtained motion data and / or voice data.
[0122] In the embodiment of the present application, the breath wake-up detection module may include a preliminary detection model, a voice detection model, a state detection model, and a voice-state detection model. The preliminary detection model is used to determine whether to detect the audio data received by the electronic device. The voice detection model is used to determine whether the audio data is a voice command sent by the user to the electronic device. The state detection model is used to determine whether the electronic device is in the lifted state when receiving the audio data. The voice-state detection model is used to detect that the audio data received by the electronic device is a voice command and the electronic device is currently in the lifted state.
[0123] The breath wake-up detection module may also wake up the voice interaction module according to other wake-up strategies, without limitation.
[0124] Such as Figure 4 As shown, the breath wake-up module of the electronic device may further include a breath wake-up start module.
[0125] The breath wake-up start module is used to start the breath wake-up detection module when the motion state of the electronic device is the lifted state.
[0126] The implementation logic of the breath wake-up start module may be to perform loop detection. Such as looping to detect whether the motion state of the electronic device is the lifted state until the motion state of the electronic device is detected as the lifted state. The breath wake-up start module may be deployed in the low-power memory space of the ADSP, so as to quickly respond to the motion state of the electronic device being the lifted state.
[0127] The gesture detection module may be deployed in the sensing area of the low-power memory space. The breath wake-up start module may be deployed in the audio area of the low-power memory space. The gesture detection module deployed in the sensing area may continue to process sensor data when the AP is in the sleep state, such as determining the motion state of the electronic device by detecting the motion data collected by the sensor. The breath wake-up start module deployed in the audio area may, when the AP is in the sleep state and in response to the motion state of the electronic device being the lifted state, start other modules of the breath wake-up module (such as the breath wake-up detection module), and then perform breath wake-up detection.
[0128] The breath wake-up module of the electronic device may further include a status forwarding module. When the gesture detection module detects that the electronic device is in the lifted state, the lifted state can be transmitted to the breath wake-up start module through the status forwarding module.
[0129] The status forwarding module can be implemented as a message transmission channel between the audio area and the sensing area. This message transmission channel is used to implement communication between processes or systems. Exemplarily, when the breath wake-up module is deployed on the ADSP, the status forwarding module can be implemented through message transmission channels such as qsocket and the Qualcomm messaging interface (QMI). When the breath wake-up module is deployed on other types of second processors, the status forwarding module can also be implemented through shared memory, etc. It can be understood that the present application does not make specific limitations on the specific implementation of the status forwarding module deployed on the second processor.
[0130] The status forwarding module can be deployed in the non-low-power memory space of the ADSP. The breath wake-up module of the electronic device may further include an audio data access module. The audio data access module is used to store the voice collected by the electronic device, so that after the voice interaction module is woken up subsequently, the voice in the audio data access module is transmitted to the voice interaction module for voice recognition.
[0131] When the gesture detection module detects that the electronic device is in the lifted state, it notifies the breath wake-up detection module to obtain motion data and / or voice, and notifies the audio data access module to obtain voice. That is, after detecting that the electronic device is in the lifted state, it starts to obtain the voice signals collected by multiple microphones, and copies the obtained voice signals into two copies, one is transmitted to the breath wake-up detection module, and the other is transmitted to the audio data access module.
[0132] Optionally, the audio data access module can be implemented as data access management (DAM) of audio data.
[0133] The audio data access module can be deployed in the non-low-power memory space of the ADSP. The audio data access module can be a module independent of the breath wake-up module.
[0134] The breath wake-up module of the electronic device may further include a breath wake-up pre-module. The breath wake-up pre-module is responsible for the pre-work of breath wake-up detection. When the breath wake-up start module detects that the electronic device is in the lifted state, the breath wake-up start module starts the breath wake-up pre-module.
[0135] When the breath wake-up startup module detects that the electronic device is in the lifted state, the breath wake-up startup module starts the breath wake-up pre-module. Then, the breath wake-up pre-module controls the ADSP to exit the low-power mode, obtains the motion data collected by the sensor, notifies the audio data access module to perform audio data caching, and starts the breath wake-up detection module.
[0136] The breath wake-up pre-module may include a low-power exit unit, a motion data acquisition unit, a notification unit, and a startup unit. The low-power exit unit is used to control the ADSP to exit the low-power mode. The motion data acquisition unit is used to obtain the motion data collected by the sensor. The notification unit is used to notify the audio data access module to perform audio data caching. The notification unit can notify the audio data access module to perform data caching by modifying the flag bit of the audio data access module. The startup unit is used to start the breath wake-up detection module.
[0137] The breath wake-up pre-module can be deployed in the low-power memory space and the non-low-power memory space of the ADSP. For example, the low-power exit unit is deployed in the low-power memory space of the ADSP. Another example is that the low-power exit unit is deployed in the audio area of the low-power memory space. Another example is that the motion data acquisition unit, the notification unit, and the startup unit can be deployed in the non-low-power memory space of the ADSP.
[0138] The breath wake-up module may further include a detection and processing module (not shown in Figure 4 ). This detection and processing module is used to perform corresponding operations according to the detection result of the breath wake-up detection module. When the breath wake-up detection module detects a successful breath wake-up, that is, it determines that the breath wake-up event is triggered, it executes the operations corresponding to the successful detection, such as waking up the AP and clearing the data and flag bits used for the current breath wake-up detection, etc. When the breath wake-up detection module detects a failed breath wake-up, that is, it determines that the breath wake-up event is not triggered, it executes the operations corresponding to the failed detection, such as clearing the data and flag bits used for the current breath wake-up detection, etc. and controlling the ADSP to enter the low-power mode.
[0139] Optionally, when the breath wake-up detection module detects that the breath wake-up event is triggered, it wakes up the AP and modifies the flag bit of the audio data access module to control the audio data access module to close the audio data caching.
[0140] Optionally, when the breath wake-up detection module detects that the breath wake-up event is not triggered, it controls the ADSP to enter the low-power mode and modifies the flag bit of the audio data access module to control the audio data access module to close the audio data caching.
[0141] When performing breath wake-up detection, the detection and processing module can first control the ADSP to enter the low-power mode. When performing breath wake-up detection, the detection and processing module can first control the audio data access module to close the audio data cache. The detection and processing module can be independent of the breath wake-up detection module, or can be used as the breath wake-up detection module or as a part of other modules. The detection and processing module can be deployed in the non-low-power memory space of the ADSP.
[0142] When the gesture detection module detects that the electronic device is in the lifted state, it can notify the breath wake-up detection module to obtain motion data and voices (such as the first voice and / or the second voice) through the above-mentioned state forwarding module, breath wake-up start module, and breath wake-up pre-module, and notify the audio data access module to obtain voices (such as the first voice and / or the second voice). In the embodiments of the present application, each module of the breath wake-up module is distributed and deployed in the ADSP. Specifically, the gesture detection module and the breath wake-up start module are deployed in the low-power memory space of the ADSP, while the state forwarding module, the breath wake-up pre-module, the breath wake-up detection module, and the audio data access module are deployed in the non-low-power memory space of the ADSP, reducing the dependence of the breath wake-up module on the low-power memory space of the ADSP.
[0143] The operation of complex programs implemented by the gesture detection module and the breath wake-up start module is not complex, and the occupancy of the low-power memory space is small. Even if the low-power memory space of the ADSP is limited, it can meet the requirement of deploying the gesture detection module and the breath wake-up start module in the low-power memory space of the ADSP.
[0144] The gesture detection module and the breath wake-up start module are deployed in the low-power memory space of the ADSP. Thus, the gesture detection module and the breath wake-up start module can detect the motion state of the electronic device in real time when the AP is in the sleep state, ensuring the real-time processing of the wake-up voice interaction module and meeting the high real-time requirement of wake-up.
[0145] The breath wake-up detection module is deployed in the non-low-power memory space of the ADSP. The non-low-power memory space is large enough to accommodate the breath wake-up detection module, and based on the higher performance and bandwidth provided by the non-low-power memory space of the ADSP, the detection efficiency of the breath wake-up detection module can be improved.
[0146] The low-power exit unit of the breath wake-up pre-module is deployed in the low-power memory space of the ADSP. Thus, when the breath wake-up start module detects the lifted state, it can quickly control the ADSP to exit the low-power mode.
[0147] The state forwarding module is deployed in the non-low-power memory space of the ADSP to quickly forward the lifted hand state based on the higher performance and bandwidth provided by the non-low-power memory space of the ADSP.
[0148] Deploy the motion data acquisition unit, notification unit, and startup unit of the breath wake-up pre-module in the non-low-power memory space of the ADSP, so as to quickly trigger the entry into the breath wake-up detection process based on the higher performance and bandwidth provided by the non-low-power memory space of the ADSP.
[0149] Deploy the audio data access module in the non-low-power memory space of the ADSP, so as to provide the efficiency of data caching based on the higher performance and bandwidth provided by the non-low-power memory space of the ADSP.
[0150] Taking the low-power memory of the low-power memory space as SRAM, the non-low-power memory of the non-low-power memory space as DRAM, and the motion data obtained by the gesture detection module as the IMU data collected by the IMU as an example.
[0151] It can be understood that taking the first memory as the non-low-power memory DRAM as an example, all the functions of the non-low-power memory DRAM can be realized by the first memory. Taking the second memory as the low-power memory SRAM as an example, all the functions of the low-power memory SRAM can be realized by the second memory. The embodiments of the present application only take the non-low-power memory DRAM and the low-power memory as SRAM as examples for illustration, and do not represent that the present application is limited thereto.
[0152] Figure 5 It is a schematic diagram of the execution process of the breath wake-up module in a voice interaction process according to an embodiment of the present application. Figure 5 It mainly gives examples of the steps executed by each module related to breath wake-up during voice interaction from each internal software architecture layer. Figure 5 It can also be regarded as an explanation of the breath wake-up process in the voice interaction solution of the present application from the perspective of the collaborative work of each internal module. Figure 5 For the relevant descriptions of the related application processor (AP), audio digital signal processor (ADSP), and each module therein, reference can be made to Figure 4 , which will not be elaborated here. Figure 5 Taking the low-power memory of the low-power memory space of the ADSP as SRAM, the non-low-power memory of the non-low-power memory space as DRAM, and the motion data obtained by the gesture detection module as the IMU data collected by the IMU as an example.
[0153] Such as Figure 5 As shown, the IMU of the electronic device collects IMU data in real time.
[0154] The gesture detection module obtains the IMU data collected by the IMU. When the gesture detection module detects a raise hand event according to the IMU data, it sends the raise hand event to the status forwarding module.
[0155] The status forwarding module determines that the electronic device is in a lifted state in response to a lift event. Specifically, the status forwarding module can obtain the lift event through inter-process or inter-system communication, and then determine that the electronic device is in a lifted state in response to the lift event. The status forwarding module sends the lifted state to the breath wake-up start module.
[0156] In response to the electronic device being in a lifted state, the breath wake-up start module triggers the start of the breath wake-up pre-module. When the breath wake-up pre-module starts, it performs the following operations: 1. Control the ADSP to exit the low-power mode; 2. Obtain IMU data and transmit the IMU data to the breath wake-up detection module; 3. Notify the audio data access module to perform audio data caching; 4. Start the breath wake-up detection module to perform breath wake-up detection.
[0157] After the breath wake-up detection module is started, the breath wake-up detection module obtains the audio data collected by the first microphone and the second microphone. The breath wake-up detection module also obtains the IMU data transmitted by the breath wake-up pre-module.
[0158] After the audio data access module is started, the audio data access module obtains the audio data collected by the first microphone and the second microphone and stores the obtained audio data; and, in response to the electronic device being in a lifted state, the audio data collected by the first microphone and the second microphone is also transmitted to the breath wake-up detection module.
[0159] When the breath wake-up detection module detects a breath wake-up event (that is, determines that the breath wake-up function can be triggered) based on the obtained audio data and / or motion data (such as IMU data), it wakes up the application processor AP. The wisdom assistant is started or woken up through the AP.
[0160] After the breath wake-up detection module detects a breath wake-up event, the breath wake-up detection module can also be used to control the audio data access module to close the data caching. After the wisdom assistant is started or woken up, the wisdom assistant can obtain the audio data cached by the audio data access module and perform speech recognition and corresponding operations.
[0161] Such as Figure 5As shown, the gesture detection module (the module in SRAM) will continuously detect and report the raise event only when the ADSP is in the low-power working mode. Specifically, the IMU continuously reports the detected motion data, and then the gesture detection module determines whether a raise event has occurred based on the IMU data. Once it is determined that a raise event has occurred, the gesture detection module will inform the breath wake-up start module of the current raise state via the status forwarding module, and then the breath wake-up start module further reports the raise state to the breath wake-up pre-module and the breath wake-up detection module (the module in DRAM) to wake up the ADSP to work in the normal working mode and trigger the steps of breath wake-up detection. After that, the breath wake-up detection module determines whether to trigger a breath wake-up event (that is, whether to start the breath wake-up function) based on the voice signals collected from the two microphones and the reported IMU data in the normal working mode of the ADSP. After determining to perform the breath wake-up function, the AP is woken up, and the AP reads the cached audio data from the audio data access module (the module in DRAM) of the ADSP. These cached audio data are synchronously sent to the breath wake-up detection module and the audio data access module after being collected by the two microphones, and the audio data used in the two modules are the same. Then the AP will perform subsequent voice interaction responses such as speech recognition, which will not be elaborated here.
[0162] However, since the gesture operations of users on electronic devices are diverse, the occurrence of a raise event does not necessarily be accompanied by breath wake-up. After the gesture detection module detects a raise event, the user may not further input a voice signal, resulting in the invalid wake-up of the ADSP and even the invalid wake-up of the AP. To address this problem, in this application, the gesture detection module continuously reports the raise event, even after the ADSP is woken up. Thus, the breath wake-up start module can decide whether to trigger a breath wake-up event based on the real-time raise state (raised or not raised), thereby effectively reducing power consumption. This way of determining whether to continue the subsequent process based on the new posture change can avoid the incorrect wake-up of the ADSP by a raise gesture (raise event) that is not intended for breath wake-up.
[0163] Figure 6 It is a schematic flowchart of a voice interaction method according to an embodiment of this application. The following Figure 6 describes each step shown.
[0164] S601. At a first moment, the electronic device is lifted, and the screen of the electronic device is in the off state. The power consumption of the electronic device changes from a first value to a second value. Before the first moment, the screen of the electronic device is in the off state, the electronic device is not lifted, and the second value is greater than the first value.
[0165] The first moment can be understood as the moment when the electronic device detects that it is lifted. For exampleFigure 3 At the moment when the electronic device is detected to be lifted during the time period of (b) in []. In some implementations, the first moment can be the moment when the electronic device detects a lift event (lift gesture). The moment before the first moment can be understood as the time when the electronic device is not lifted in the screen-off state. For example, Figure 3 Any moment during the time period of (a) in []. can be considered as the time when the electronic device is not lifted.
[0166] S602. At a second moment after the first moment, the electronic device is not lifted, and the screen of the electronic device is in the screen-off state. The power consumption of the electronic device changes from a second value to a first value.
[0167] That is to say, if the electronic device is still in the screen-off state and the lift state of the electronic device changes, the power consumption returns to the original lower value.
[0168] In one implementation, at a second moment after the first moment, when the electronic device is not lifted and the screen of the electronic device is in the screen-off state, and the power consumption of the electronic device changes from a second value to a first value, it can include: at the second moment, based on the electronic device transitioning from the lifted state to the non-lifted state, the power consumption of the electronic device changes from the second value to the first value. In this implementation, it explains why the power consumption value changes at the second moment, because the electronic device changes based on the change in the lift state in the screen-off state. It can also be understood that the power consumption change at the second moment occurs triggered by the transition of the lift state from lifted to non-lifted. Combining with the above, the first moment is when the lift state of the electronic device changes from non-lifted to lifted and increases, thus forming a complete solution for changing the power consumption based on the change in the lift state of the electronic device. In the traditional solution, there will be no such second moment as in this application (or it can be understood that the second moment in the traditional solution will not be like this application where the electronic device ends the process based on the change in the lift state). Once the electronic device is lifted in the traditional solution and the power consumption changes from the first value to the second value, it needs to wait at least for the following third moment (no appropriate voice data is collected within the preset duration) or the sixth moment (the entire voice interaction process is completed) to return to the first value again. Therefore, in comparison, this application achieves the effect of "timely stop-loss", being able to end the process in a timely manner after the lift state changes without wasting power by executing subsequent processes in vain.
[0169] In another implementation, the above method further includes: before the first moment, the first processor of the electronic device continuously detects a raising event in the low-power mode. The raising event is used to indicate whether the electronic device is raised. When a raising event is detected, it means the electronic device is raised, and when a non-raising event is detected, it means the electronic device is not raised; at the first moment, in response to detecting the raising event, the first processor switches from the low-power mode to the normal operating mode, and starts to determine whether to trigger a breath wake-up event in the normal operating mode, and continues to obtain the detection result of the raising event in the normal operating mode; at the second moment, in response to detecting a non-raising event, the first processor switches from the normal operating mode to the low-power mode, and continuously detects the raising event in the low-power mode. In this implementation, the first processor continuously detects the raising event in the low-power mode. Once the raising event is detected, it means the electronic device is raised, that is, the first moment arrives. The electronic device switches from the low-power mode to the normal operating mode in response to detecting the raising event, so that the power consumption increases from the previous first value to the subsequent second value starting from the first moment. From the first moment until the second moment, the first processor is in the normal operating mode. Until the second moment, when the electronic device detects a non-raising event again, it knows that the raised state has ended, and thus directly returns the first processor to the low-power mode based on this, rather than waiting for a sufficient duration and executing subsequent processes, thereby effectively reducing the power consumption.
[0170] The first processor can be an ADSP or other processors that can deploy a breath wake-up module and have a low-power memory space. The second processor can be an AP or other processors that can perform speech recognition and respond to the recognition result.
[0171] In another implementation, the above method further includes: between the first moment and the third moment, the electronic device remains raised, and the screen of the electronic device is in the off-screen state. The time interval between the first moment and the second moment is less than the time interval between the first moment and the third moment. The time interval between the first moment and the third moment is a preset time interval; at the third moment, the power consumption of the electronic device changes from the second value to the first value. In this implementation, the step of waiting for a fixed duration in the traditional solution is continued, so that on the basis of superimposing the solution of the present application, the timely end of the process can be ensured. That is, if the second moment does not occur, that is, the raised state of the electronic device does not change, the process ends after counting the preset time interval from the first moment, and the power consumption is reduced back to the first value.
[0172] In one example, at a third moment, the power consumption of the electronic device changes from a second value to a first value, including: at the third moment, based on the fact that the electronic device does not collect voice data that meets a first preset condition, the power consumption of the electronic device changes from the second value to the first value. The first preset condition includes: the signal strength of the voice data collected by the first microphone of the electronic device is greater than or equal to a first preset signal strength threshold, the signal strength of the voice data collected by the second microphone of the electronic device is greater than or equal to a second preset signal strength threshold, and the difference between the signal strength of the voice data collected by the first microphone and the signal strength of the voice data collected by the second microphone is greater than or equal to a preset signal strength difference threshold. In this example, the triggering factor for adjusting the power consumption at the third moment is given, that is, based on the fact that the electronic device does not collect voice data that meets the first preset condition. Combining with the above, there are two ways for this application to trigger "return to low power consumption", the lifted state is changed (or understood as the electronic device ends the lifted state) and no voice data that meets the first preset condition is collected within a preset time interval. The two triggering methods can be executed separately or superimposed. Only executing the first method is the solution of this application, only executing the second method is the traditional solution, and executing both methods has the best effect of reducing power consumption.
[0173] In one example, at the third moment, based on the fact that the electronic device does not collect voice data that meets the first preset condition, the power consumption of the electronic device changes from the second value to the first value, including: before the first moment, the first processor of the electronic device operates in a low-power mode; between the first moment and the third moment, the first processor operates in a normal working mode; the first processor continuously reads the voice data collected by the first microphone and the second microphone starting from the first moment; at the third moment, based on the fact that the electronic device is lifted and the electronic device does not collect voice data that meets the first preset condition, the first processor determines not to trigger the breath wake-up event, controls the first processor to switch from the normal working mode to the low-power mode, and continuously detects the lift event in the low-power mode. In this example, it mainly shows that the change in the power consumption value is brought about by the change in the working mode of the first processor. Before the first moment, the first processor is in the low-power mode, so the power consumption is the lowest first value. Between the first moment and the third moment, it is in the normal working mode, so the power consumption is the higher second value. At the third moment, since the voice data that meets the first preset condition is not collected and the breath wake-up event is not triggered, the first processor is made to work in the low-power mode again. Therefore, the power consumption changes from the first value to the second value and then back to the third value throughout the process.
[0174] It should be noted that in this example, the electronic device remains in the lifted state at the second moment (the electronic device is continuously lifted directly from the first moment to the third moment). Therefore, the process will not end and power consumption will not be reduced based on the change in the lifted state at the second moment in step S602. Instead, power consumption is reduced based on the fact that no appropriate voice data is collected at the third moment.
[0175] In one implementation, the above method further includes: at a fifth moment after the first moment, the electronic device turns on the screen, and the power consumption of the electronic device is a third value, where the third value is greater than the second value. In this implementation, the power consumption change brought about by successful breath wake-up is mainly presented. The moment when the interface 107 just lights up can be regarded as an example of the third moment.
[0176] In an example, at a fifth moment after the first moment, the electronic device turns on the screen, and the power consumption of the electronic device is a third value, where the third value is greater than the second value, including: at the first moment, the first processor of the electronic device switches from the low-power mode to the normal working mode, so that the power consumption of the electronic device increases from the first value to the second value; at a fourth moment after the first moment, based on the fact that the electronic device starts to collect voice data that meets the first preset condition, the first processor starts to continuously read the collected voice data; the first preset condition includes: the signal intensity of the voice data collected by the first microphone of the electronic device is greater than or equal to the first preset signal intensity threshold, the signal intensity of the voice data collected by the second microphone of the electronic device is greater than or equal to the second preset signal intensity threshold, and the difference between the signal intensity of the voice data collected by the first microphone and the signal intensity of the voice data collected by the second microphone is greater than or equal to the preset signal intensity difference threshold; at a fifth moment after the fourth moment, the first processor finishes reading the first voice data, and the first processor wakes up the second processor based on the fact that the first voice data meets the second preset condition, so that the second processor switches from the low-power mode to the normal working mode and performs voice recognition and response on the first voice data in the normal working mode, so that the power consumption of the electronic device increases from the second value to the third value. Meeting the second preset condition is used to indicate that the first voice data can be used to trigger a breath wake-up event. In this example, how the power consumption changes from the first moment to the fifth moment with the working mode of the processor is given. Before the first moment, both the first processor and the second processor are in the low-power mode. Between the first moment and the fourth moment, the first processor is in the normal working mode. Between the fourth moment and the fifth moment, the first processor is in the normal working mode. After the fifth moment, both the second processor and the first processor are in the normal working mode, so that the power consumption continues to increase in the case of successful wake-up.
[0177] In another example, the above method further includes: at the fifth moment, based on the fact that the first voice data does not meet the second preset condition, the first processor switches from the normal working mode to the low-power mode, so that the power consumption of the electronic device changes from the second value to the first value. In this example, when the first processor determines that the breath wake-up event is not triggered, it switches back to the low-power mode again.
[0178] In another example, the above method further includes: at the sixth moment after the fifth moment, the screen of the electronic device switches from the lit state to the off state, and the power consumption of the electronic device changes from the third value to the first value. In this example, when the screen of the electronic device is turned off again, the power consumption returns to the first value.
[0179] In another example, at the sixth moment after the fifth moment, the screen of the electronic device switches from the lit state to the off state, and the power consumption of the electronic device changes from the third value to the first value, including: at the sixth moment, based on the fact that the second processor has finished responding to the first voice data, both the first processor and the second processor switch to the low-power mode, so that the power consumption of the electronic device changes from the third value to the first value. In this example, after the second processor finishes responding, both processors are switched to the low-power mode, thereby reducing the power consumption to the first value.
[0180] In another example, the above method further includes: at the fourth moment, based on the fact that the electronic device has not collected voice data that meets the first preset condition, the first processor has not read voice data that meets the first preset condition, and the electronic device continues to collect voice data until the third moment, and the time interval between the third moment and the first moment is a preset time interval. In this example, if voice data has not been collected at the fourth moment, it will continue to be collected until the third moment arrives.
[0181] It should also be noted that in addition to the above first preset condition and second preset condition, a third preset condition can be set before the first preset condition. The third preset condition can be identity authentication, that is, voiceprint recognition. Only voice data that meets the voiceprint recognition can execute the subsequent process. For example, in the following Figure 9 is an example of voiceprint recognition first and then conversation recording. The start of the so-called conversation recording can be regarded as an example of the fourth moment in this article.
[0182] Figure 6 The method shown is mainly in the voice interaction process in the breath wake-up scenario. When the electronic device is lifted and then the lift ends, the power is quickly reduced back to the low-power state, thereby eliminating the unnecessary power consumption caused by non-breath wake-up raises, so as to reduce the overall power consumption. Figure 6 The method shown can be executed after Figures 1 - 2 the setup process shown ends, Figure 6 The method shown can be used forFigure 3 The scene shown Figure 6 The method shown can utilize Figure 4 and Figure 5 Each module shown to execute. For related content, refer to the above text and will not be elaborated here.
[0183] The complete process of voice interaction in this application involves three stages: the breath wake-up detection stage, the speech recognition stage, and the dialogue stage. In the breath wake-up detection stage, gesture detection and the detection of whether to trigger a breath wake-up event are mainly carried out. The speech recognition stage is mainly the stage of performing speech recognition after determining that a breath wake-up event is to be triggered. The dialogue stage is the stage of collecting audio data in the breath wake-up detection stage. Combining Figure 3 , the breath wake-up detection stage corresponds to Figure 3 in (a)-(b) and the stage after the user inputs speech; the speech recognition stage corresponds to interface 107, and the text information corresponding to the speech data input by the user is displayed in the interface through speech recognition; the dialogue stage can correspond to Figure 3 the stage where the user is speaking in (b). The breath wake-up detection stage is mainly executed by the ADSP, while the speech recognition stage and the dialogue stage are executed by the AP. The following will be described in combination with Figures 7 - 9 , where Figure 7 gives an example of the breath wake-up detection stage, Figure 8 gives an example of the speech recognition stage, Figure 9 and gives an example of the dialogue stage.
[0184] Figure 7 is a schematic diagram of the execution process of a breath wake-up detection in an embodiment of this application. Figure 7 Mainly explains the relevant steps in the breath wake-up detection stage during the entire voice interaction process.
[0185] S701. The breath wake-up module starts QMI to register the raise hand event.
[0186] It can be understood that the breath wake-up detection module starts to register the raise hand event with the status forwarding module (taking QMI as an example here). It should be understood that Figure 7 although QMI is taken as an example in
[0187] S702. QMI registers the raise hand event with the sensor server.
[0188] It should be understood that, as described above, QMI is a message transmission channel. Therefore, QMI can work under the scheduling of other modules. For example, a real-time operating system (RTOS) can respond to a request (notification, message) from the breath wake-up module to start the QMI registration of the raise hand event, and then start and register the raise hand event with the sensor server through QMI. For example, the real-time operating system can call a registration function to register the raise hand event. The registration function may include an event ID, a callback function, etc. The return value may be a status code indicating whether the registration is successful.
[0189] S703. The breath wake-up module obtains the returned registration result from the sensor server through QMI.
[0190] The returned registration result can be an identification information used to identify whether the raise hand event is successfully registered. It can be that the real-time operating system calls the registration function to start QMI to register the raise hand event with the sensor server. Then, the sensor server returns the registration result to the real-time operating system through QMI, and then it is returned to the breath wake-up module via the real-time operating system. After the breath wake-up module obtains the returned result, it determines whether the raise hand event is successfully registered based on the returned result.
[0191] S704. In the case of successful registration, the breath wake-up module closes the caching function of the audio data access module.
[0192] That is to say, after the breath wake-up module determines that the registration of the raise hand event is successful based on the obtained returned result in step S703, it can immediately close the caching function of the audio data access module.
[0193] The closing method can be achieved by modifying the flag bit of the audio data access module. The operation of modifying the flag bit can be executed by the breath wake-up pre-module (notification unit) in the breath wake-up module.
[0194] It should be understood that closing the cache is to prevent audio data from entering the cache module. Therefore, the default state of the audio data access module should be the state of closing the cache. It will only be opened when it is necessary to start recording audio data under the trigger of the raise hand event, so as to store the segment of audio data. When the breath wake-up event is triggered, the AP can read the segment of audio data from the audio data access module for subsequent voice recognition and other operations. For other situations, the caching function of the audio data should not and does not need to be opened.
[0195] Steps S701 - S704 can be understood as preparation steps, or regarded as "function initialization". After these preparations are completed, the detection of breath wake-up can be carried out.
[0196] The breath wake-up module controls the ADSP to enter the low-power mode and notifies QMI.
[0197] In one implementation, the breath wake-up module or the breath wake-up detection module of the breath wake-up module calls a function to notify the ADSP to enter the low-power mode. For example, it calls a function to notify the above-mentioned low-power island to enter the low-power mode.
[0198] S706. The breath wake-up module continuously obtains the detection result of the hand-raising event in the low-power mode.
[0199] Combined with the example of the software architecture above, the detection of the hand-raising event can be based on, for example, the IMU to detect the motion data of the electronic device and then report it to the gesture detection module of the breath wake-up module (this gesture detection module is deployed in the low-power memory space of the ADSP). The gesture detection module determines whether a hand-raising event occurs (it can be an identifier of a hand-raising event, used to identify two results: a hand-raising event and a non-hand-raising event). After detecting the hand-raising event, the gesture detection module further reports it to the status forwarding module of the breath wake-up module (this module can be deployed in the non-low-power memory space of the ADSP). The status forwarding module then sends the gesture detection result (whether a hand-raising event occurs) to the breath wake-up start module of the breath wake-up module (this module is deployed in the low-power memory space of the ADSP). Therefore, the so-called obtaining of the hand-raising event by the breath wake-up module in this step can be executed by the module of the breath wake-up module deployed in the low-power memory space.
[0200] S707. The sensor server reports the hand-raising event to the breath wake-up module.
[0201] That is to say, once the gesture detection module detects a hand-raising event or whether the electronic device is in a raised state, the sensor server can report the hand-raising event to the breath wake-up module via QMI. It can be expressed as "if (the sensor detects a hand-raising action) {exit the low-power state; the sensor notifies the audio of the hand-raising event}", where the audio can be regarded as the breath wake-up module here.
[0202] S708. When the breath wake-up module receives the trigger of the hand-raising event, it controls the ADSP to exit the low-power mode and notifies QMI.
[0203] That is to say, after step S708 is executed, the ADSP (or rather, the breath wake-up module therein) will work in the normal working mode (non-low-power working mode).
[0204] S709. When triggered by the hand-raising event, the breath wake-up module turns on the cache function of the audio data access module.
[0205] The S710, the breath wake-up module determines whether a breath wake-up event is triggered in the normal working mode.
[0206] Combined with the above, the motion data acquisition unit and the start unit of the breath wake-up pre-module of the breath wake-up module can run. The start unit starts the breath wake-up detection module. After the breath wake-up detection module is started, it acquires voice. The motion data acquisition unit transmits the acquired motion data to the breath wake-up detection module. The breath wake-up detection module can perform breath wake-up recognition based on the acquired voice and motion data. The breath wake-up detection module can determine whether a breath wake-up event is triggered based on the acquired voice and / or motion data.
[0207] The determination result must be one of the two possibilities of "trigger" and "not trigger", that is, one of the two situations of "yes" and "no". Then subsequent responses can be based on different determination results. For example, if the determination result is "trigger", step S716 is executed to wake up the AP for subsequent processing; if the determination result is "not trigger", steps S717 - S719 are executed to end the current round of wake-up process and continue to switch to the low-power mode to continuously acquire lift hand events.
[0208] In the traditional solution, the complete process of breath wake-up detection ends here. However, since not every lift hand event is executed by the user for breath wake-up, it may occur that until step S716 is executed and after the AP processor performs voice recognition, it is found that this lift hand event is not for performing breath wake-up voice interaction. For example, the user just picks up the phone and flips it while talking to someone else. In this case, since the phone will still be close to the mouth, voice data that meets the conditions will be collected, so it may occur that the AP does not find any questions or instructions until after voice recognition, resulting in a large waste of power consumption; it may also cause the power consumption during the period of making the judgment to be wasted in vain although it will return to the low-power mode because the determination result is "not trigger" and steps S707 - 719 are executed. For example, the user just picks up the phone and then quickly puts it down, resulting in not enough voice data being collected to trigger the breath wake-up event, thus being determined as "not trigger", but during this period, the ADSP is working in a non-low-power mode, bringing unnecessary power consumption waste.
[0209] To address the above problems, the present application adds a new process, that is, after the breath wake-up module exits the low-power mode in step S708, it still continues to continuously acquire lift hand events. Then the sensor server will still report non-lift hand events, so that the breath wake-up module can end the wake-up process in a timely manner (as early as possible) under the trigger of non-lift hand events during the execution of step S710, thereby quickly returning the ADSP to the low-power working mode and saving power consumption.
[0210] In addition, in the solution of this application, when step S710 is executed, the detection results of the comprehensively collected voice data and real-time hand-raising events (hand-raising events obtained during S710) are used to determine whether to trigger a breath wake-up event, so as to effectively reduce the power consumption in the case of non-breath wake-up by ending ineffective wake-up in a timely manner. In contrast, in the traditional solution, when step S710 is executed, only the hand-raising event reported in S707 (that is, the last hand-raising event reported before exiting the low-power mode) and the collected voice data are combined to determine whether to trigger a breath wake-up event, resulting in a series of power consumption as long as there is a hand-raising event (even if it is not a hand-raising event for breath wake-up), causing power consumption waste.
[0211] S711. During the execution of step S710, the sensor server reports non-hand-raising events to the breath wake-up module.
[0212] It should be understood that the non-hand-raising event here can be understood as the "negative" result of the hand-raising event. That is to say, the detection result of the hand-raising event may be one of the two indications of "it is a hand-raising event" and "it is not a hand-raising event", and the so-called "reporting non-hand-raising events" means that during the period when the breath wake-up module continuously obtains the detection results of the hand-raising event, it obtains a detection result and learns from it that there is another gesture of "it is not a hand-raising event" currently occurring. The so-called "non-hand-raising event" can be understood as other gestures that are not hand-raising events.
[0213] S712. Under the trigger of the non-hand-raising event, the breath wake-up module ends step S710, thereby ending the current round of breath wake-up process.
[0214] That is to say, after the breath wake-up module receives the "non-hand-raising event", it learns that a new posture change has occurred in the current electronic device, resulting in the end of the "hand-raising event" reported in step S707. Therefore, based on such a result, step S710 is ended in advance, thereby ending the current round of breath wake-up process in advance.
[0215] It should be understood that in order to ensure wake-up is possible, step S710 needs to last for a relatively long time. First, the user needs to wait until the microphone is close enough to the mouth before speaking, so a certain waiting delay needs to be reserved for the user. Second, the user may not speak immediately and may pause, such as swallowing saliva or clearing the throat, which will also bring possible waiting times. The above will result in a possible reserved duration before voice data is recorded. In addition, from the start of the user's speech to the end of the user's speech, there is also a need to pause and determine whether the user has finished recording voice data, and it is also necessary to decide whether to trigger the breath wake-up event based on the collected voice data and the detection result of the real-time hand-raising event. This brings the duration during which step S710 must be executed. The above possible reserved duration and the necessary execution duration combined are the execution duration of step S710. Therefore, it can be seen that the execution period of step S710 is a relatively long duration, generally several seconds. And the above reserved duration is often set to a fixed duration in traditional solutions, such as 2 seconds or 3 seconds in the above text, which will not be elaborated here. In this application, the step of continuously obtaining the detection result of the hand-raising event is still carried out during the execution of step S710, so that the posture change of the electronic device can be detected in time, that is, the interruption of the hand-raising event can be detected in time, so as to end the breath wake-up process and the subsequent AP processing process as soon as possible and save power consumption.
[0216] S713. When ending the current round of breath wake-up process, the breath wake-up module closes the caching function of the audio data access module.
[0217] S714. The breath wake-up module controls the ADSP to enter the low-power mode and notifies QMI.
[0218] This step can refer to step S705. The execution content and method of the two can be the same, except for the trigger condition (the execution time node).
[0219] S715. The breath wake-up module continuously obtains the detection result of the hand-raising event in the low-power mode.
[0220] This step can refer to step S706. The execution content and method of the two can be the same, except for the trigger condition (the execution time node).
[0221] S716. In the case of determining that the breath wake-up event is triggered, the breath wake-up module wakes up the application processor.
[0222] The subsequent execution process of the AP can refer to the above text and will not be elaborated here.
[0223] S717. When the determination result of step S710 is "no", that is, the breath wake-up event is not triggered, the breath wake-up module ends the current round of breath wake-up process, controls the ADSP to enter the low-power mode, and notifies QMI.
[0224] S718. When ending the breath wake-up process of the current round, the breath wake-up module closes the caching function of the audio data access module.
[0225] S719. The breath wake-up module continuously obtains the detection result of the raise hand event in the low power consumption mode.
[0226] This step can refer to steps S706 and S715. Their execution contents and methods can be the same, except for the trigger conditions (execution time nodes).
[0227] Figure 7 In the sensor server, it mainly receives the registration of the raise hand event of audio (which can be regarded as the breath wake-up module), detects the raise hand event, exits the low power consumption mode when detecting a raise hand action, and notifies audio of the raise hand event. Audio mainly starts the wake-up word-free function (starts the breath wake-up function), registers the raise hand event with the sensor and calls back the registration result, waits for the raise hand event after successful registration, and when the detection fails (that is, a non-raise hand event is detected), determines to exit the raise hand state, notifies the audio data access module (such as the DAM module) to close the caching function, interrupts the algorithm processing of the ADSP (that is, the step of judging whether to trigger the breath wake-up event), and makes the ADSP enter the low power consumption mode. Audio also needs to obtain the raise hand event in the low power consumption mode, exit the low power consumption mode when it is determined that a raise hand event has occurred, open the caching function and perform the discrimination of whether to trigger the breath wake-up event. When it is determined to trigger, stop caching and wake up the relevant APP of the AP for subsequent voice interaction processing. When it is determined not to trigger, close the caching and return to the low power consumption mode.
[0228] From Figure 7 It can be seen that the solution of this application mainly reduces unnecessary power consumption by continuously obtaining the raise hand event during the execution of judging whether to trigger the breath wake-up event and terminating the process based on the reported non-raise hand event.
[0229] Figure 8 is a schematic flowchart of a voice recognition stage in an embodiment of this application. Figure 8 The execution processes in the two scenarios of wake-up and non-wake-up are introduced separately. Figure 8 The process of
[0230] S801. When the intelligent assistant in the application layer of the AP fails to call the audio service module (audioserver) in the framework layer, the voice recognition is not started.
[0231] In an embodiment of the present application, after the intelligent assistant is started, the intelligent assistant attempts to call the audio service module. Only when the call is successful can the channel between the two be opened to start voice recognition. When the call fails, the channel cannot be opened and voice recognition cannot be started. Figure 8 In the figure, "×" indicates that the intelligent assistant fails to call the audio service module (i.e., the call fails).
[0232] It should be noted that when the intelligent assistant calls the audio service module, there are two situations: success and failure. When the audio channel of the audio service module is occupied, or when the intelligent assistant exits after successful startup, the intelligent assistant fails to call the audio service module. For example, after the intelligent assistant is started, if the recording function of the electronic device is turned on and the audio channel of the audio service module is occupied, the intelligent assistant may fail to call the audio service module. It can be seen that mainly because the audio service is not a dedicated module for the voice assistant and is not dedicated to voice recognition, conflicts with other services may occur, so the call may fail.
[0233] S802. The audio trigger module (soundtrigger) in the framework layer of the AP times out in monitoring and notifies the ADSP to return to the breath wake-up detection stage again.
[0234] The determination of the monitoring timeout here can be achieved by setting a time threshold or a timer. In an embodiment of the present application, when the intelligent assistant successfully calls the audio service module, the intelligent assistant sends a request to start voice recognition to the audio service module (for example, it can be represented as the state CaptureState). In response to the request to start voice recognition sent by the intelligent assistant, the audio service module sends a request to start voice recognition to the audio trigger module in the framework layer. The audio trigger module in the framework layer starts a timer to facilitate the audio trigger module to determine whether a request to start breath wake-up recognition is received within the timing time corresponding to the timer. If a request to start voice recognition is not received within the timing time, that is, the intelligent assistant fails to call the audio service module. If a request to start voice recognition is not received within the timing time, the monitoring times out.
[0235] When notifying the ADSP to restart breath wake-up detection, the function StartRecognition can be sent to the audio trigger module in the HAL layer of the AP, and then the audio trigger module in the HAL layer notifies the ADSP to perform a state switch for breath wake-up detection. For example, the status flag state can be changed from BUFFERING to ACTIVE (BUFFERING and ACTIVE respectively correspond to turning off and starting breath wake-up detection).
[0236] Scenario 1 (including steps S801 - S802) mainly targets the situation where voice recognition is not performed due to a call failure or a listening timeout, and the system reverts to the breath wake-up detection phase.
[0237] Scenario 2 (including steps S803 - S810) targets the situation where the call is successful and voice recognition can be performed.
[0238] S803. The intelligent assistant successfully calls the audio service module and starts voice recognition.
[0239] S804. The audio service module notifies the audio trigger module to start recording.
[0240] In the embodiment of the present application, when the audio service module is successfully called, the audio service module starts voice recognition.
[0241] It should be understood that the recording here is not the traditional recording. Instead, it obtains voice data from the audio trigger module in the HAL layer, and the audio trigger module in the HAL layer obtains the cached voice data from the audio data access module in the ADSP.
[0242] Starting the recording can be achieved through the function startinput.
[0243] S805. The audio trigger module in the framework layer notifies the ADSP to stop breath wake-up detection through the audio trigger module in the HAL layer.
[0244] It can be notified through the function StopRecognition.
[0245] Since the AP now needs to process the voice recognition of the audio data of the current round of breath wake-up, it is necessary to let the ADSP pause the continued breath wake-up detection first. However, it should be understood that step S805 may also not notify the ADSP to stop breath wake-up detection because corresponding policies can be set on the ADSP side to stop breath wake-up detection when the breath wake-up event is triggered, rather than waiting for the notification from the AP. Therefore, step S805 can just notify the audio trigger module in the HAL layer to stop breath wake-up detection, which can temporarily cut off the data interaction with the audio data access module in the ADSP.
[0246] S806. The audio trigger module in the framework layer reads the recording from the audio trigger module in the HAL layer and forwards it to the intelligent assistant.
[0247] When reading the recording LAB data, audioserver can request the audio trigger module in the HAL layer (the primary-hal therein) to read the data, and then st-hal obtains the recording data through a callback.
[0248] After that, the intelligent assistant will perform speech recognition on the recorded audio (voice data).
[0249] S807. After the intelligent assistant finishes speech recognition, it notifies the audio service module in the framework layer to stop recording.
[0250] The recording can be stopped by the function stopinput.
[0251] S808. The audio service module further notifies the audio trigger module in the framework layer to stop speech recognition.
[0252] In one example, it can be that audioserver sends a request to stop speech recognition to soundtrigger through the function setSundayTriggerCaptureState.
[0253] S809. The audio trigger module in the framework layer notifies ADSP to restart breath wake-up detection through the audio trigger module in the HAL layer.
[0254] It can be that soundtrigger sends a request to stop speech recognition to the audio trigger module (st-hal therein) in the HAL layer through the function StopRecognition. st-hal starts breath wake-up detection and notifies ADSP to restart breath wake-up detection.
[0255] Before Figure 8 the two scenarios mentioned above, a registration phase can also be included. In the registration phase, mainly the above-mentioned various modules are prepared so that the related functions of speech recognition in the breath wake-up scenario can be executed. It can be understood that the data path between the intelligent assistant and each module needs to be established so that the voice assistant app can perform speech recognition in the wake-word-free wake-up scenario. The registration phase needs to request ADSP to load the speech recognition model, start gesture detection, trigger breath wake-up events, read cache functions, report breath wake-up events, etc.
[0256] It should also be understood that Figure 8 only the intelligent assistant app is taken as an example, but in practice, other application programs can also be used, without limitation, as long as speech recognition can be performed.
[0257] Figure 9 is a schematic flowchart of a dialogue phase in an embodiment of the present application. Figure 9 It can be understood that it is executed in the audio acquisition phase of the dialogue phase in the breath wake-up detection phase, that is, it can be understood as the steps that can be executed during the execution of step S710.
[0258] S901. The audio service module (Audio service) in the framework layer of the AP determines that the voiceprint recognition is successful and starts the conversation recording.
[0259] After waking up the electronic device in the way of no wake-up word, the audio data input by the user is automatically subjected to voiceprint recognition. After verifying that the current user is a legitimate user based on the voiceprint recognition, it is automatically unlocked and the conversation recording is started. Thus, the voice control unlocking of the electronic device is realized, that is, the contactless unlocking of the electronic device is realized.
[0260] S902. The audio service module in the framework layer requests the audio trigger module in the hardware abstraction layer (HAL layer) of the AP to open the audio input stream.
[0261] Based on the successful voiceprint recognition, the audio service module requests to open the audio input stream. The request for stream establishment may include: the audio service module can initialize the audio input stream and concurrently return a handle to read the audio data.
[0262] In one example, the audio service module uses the function adev_open_input_stream(); audio_io_handle_thandle to send a request to the primary-hal to open the audio input stream.
[0263] Among them, the adev_open_input_stream() function is used to initialize an audio input stream and return a handle of the audio_io_handle_thandle type, which can be used to read the audio data, control the audio input stream, and configure the audio parameters.
[0264] S903. The audio trigger module in the HAL layer creates an audio stream.
[0265] In the embodiment of the present application, the hardware abstraction layer creates an audio stream in response to the request for opening the audio stream.
[0266] Exemplarily, it may be that the primary-hal creates an audio stream through the function CreateStreamln() and creates a class of StreamlnPrimary() in the method of CreateStreamln().
[0267] S904. The audio service module requests the audio trigger module in the HAL layer to start reading the audio data.
[0268] The request for starting to read the audio data can be sent through the function in_read().
[0269] S905. The audio trigger module in the HAL layer requests to read the audio data from the ADSP.
[0270] It can also be understood that the audio trigger module of the HAL layer controls the ADSP to read audio data.
[0271] S906. The ADSP obtains real-time audio data from the audio acquisition device (MIC).
[0272] In the embodiment of the present application, the ADSP can read real-time audio data from an audio acquisition device such as a microphone, and perform corresponding algorithm processing on the audio data in the ADSP. For example, in response to a request for audio data, the ADSP reads real-time audio data from an audio acquisition device such as a microphone.
[0273] It can be understood that under the trigger of step S707, the recording for breath wake-up detection starts. Once a voice signal input by the user is obtained, steps S901 - S905 are immediately executed to complete the preparation for recording, and S906 is executed to record and process the audio data. The so-called processing of the audio data by the ADSP is the related processing in step S710. For example, it may include judging whether the breath wake-up requirement is met based on the signal intensities of two microphones, judging whether to trigger a breath wake-up event, and synchronously caching to the audio data access module, etc., which will not be elaborated here.
[0274] To facilitate understanding of how the solution of the present application reduces power consumption, the following is combined with Figure 10 and specific examples for explanation.
[0275] Figure 10 is a power consumption comparison diagram of the solution of the embodiment of the present application and the traditional solution in the same interaction scenario.
[0276] As Figure 10 shown in (a) and (b) therein are the power consumption schematic diagrams of the traditional solution and the solution of the present application respectively. The horizontal axis is the natural time t, and the vertical axis is the power consumption P.
[0277] Suppose Figure 10 The scenario is as follows: The user picks up the stationary electronic device A from the desktop at time T1, and immediately puts it down after about 0.3 seconds (s). The user does not speak. Then, at time T2, the user flips the electronic device A, which lasts for about 0.3 s. The user does not speak. After that, at time T3, when chatting with someone, the user picks up the electronic device B, brings it close to the face, flips it over, and uses the small mirror on the back protective cover to take a picture of the face. Suppose the flip occurs at 0.5 s. At times T1, T2, and T3, the electronic device A is in the screen-off state, and the electronic device A has activated the breath wake-up function. The breath wake-up detection uses the ADSP. The gesture detection and the breath wake-up start detection are deployed in the low-power space of the ADSP, while the algorithm for breath wake-up detection and the audio data access module, etc., are deployed in the non-low-power space of the ADSP. The AP is used to perform voice recognition and subsequent responses. The waiting fixed duration for collecting voice data is 2 seconds. Figure 10In (a) and (b), they are the power consumptions in the same scenario described above respectively.
[0278] In the traditional solution, once a raise hand event is detected, the ADSP will be woken up to judge whether to trigger the breath wake-up event. During this period, the user's voice data will be collected. If the user's voice signal is collected within 2s, the user's voice data will continue to be collected and the judgment on whether to trigger the breath wake-up event will be made. And when it is determined to trigger, the AP will be woken up for subsequent processing such as speech recognition. If the required voice signal cannot be collected within 2 seconds continuously, it will return to the low power mode again.
[0279] As Figure 10 shown in (a), at time T1, the ADSP wakes up, the power rises from P1 to P2, and starts to attempt to collect the user's voice data. Since the user does not speak, after 2s, the ADSP determines not to trigger the breath wake-up event, and thus returns to the power P1 again; at time T2, the ADSP wakes up, the power rises from P1 to P2, and starts to attempt to collect the user's voice data. Since the user does not speak, after 2s, the ADSP determines not to trigger the breath wake-up event, and thus returns to the power P1 again; at time T3, the ADSP wakes up, the power rises from P1 to P2, and starts to attempt to collect the user's voice data. Since the user is chatting, when approaching the user's mouth, the voice data can be collected. Assuming that the voice data is continuously collected for 1.8s starting from 0.5s, the ADSP power further rises from P2 to P3 and lasts for 1.8s. After that, the ADSP determines that the breath wake-up event can be triggered based on the 1.8s duration of the voice data, and then further wakes up the AP for speech recognition, and the power further rises to P4 (the power of the ADSP + the power of the AP, but the power of the AP may be much higher than the power of the ADSP). Since the speech recognition of the AP needs to call the audio service module and so on and still needs to wait, it is assumed that the AP may still need to wait for a period of time for recognition. Assuming that after 2.3s, the recognized text information makes the AP confirm that it cannot respond, thus ending the current round of voice interaction. The AP returns to the low power mode and notifies the ADSP to also return to the low power mode and return to the low power mode again for breath wake-up detection. Therefore, Figure 10 The total power consumption in (a) is: P1 * total running duration + P2 * 2s + P2 * 2s + P2 * 0.5s + P3 * 1.8s + P4 * 2.3s.
[0280] It should also be understood that the difference between P3 and P2 may be almost negligible, because during these 1.8s and the previous 0.5s, the ADSP is working in the normal working mode and the AP has not been woken up yet. So whether there is a continuous increase in power consumption depends on the processing difficulty of the collected voice data.
[0281] In the solution of the present application, once a raise hand event is detected, the ADSP will be awakened to determine whether to trigger the breath wake-up event. During this period, the user's voice data will be collected. If a user voice signal is collected within 2 seconds, continuous collection of the user's voice data will start and a determination will be made on whether to trigger the breath wake-up event. When it is determined to trigger, the AP will be awakened for subsequent processing such as voice recognition. If no qualified voice signal is collected within 2 seconds, the system will return to the low-power mode again. However, on this basis, the detection result of the raise hand event will still be obtained. Once an indication of the end of the raise hand event is received, the judgment process and all subsequent processes will be directly terminated.
[0282] As Figure 10 As shown in (b) therein, at time T1, the ADSP is awakened, and the power rises from P1 to P2. An attempt is made to collect the user's voice data. Since the raise hand event ends after 0.3 seconds, the ADSP ends the process based on the detection result of the new raise hand event and thus returns to the power P1 again; at time T2, the ADSP is awakened, and the power rises from P1 to P2. An attempt is made to collect the user's voice data. Since the raise hand event ends after 0.3 seconds, the ADSP ends the process based on the detection result of the new raise hand event and thus returns to the power P1 again; at time T3, the ADSP is awakened, and the power rises from P1 to P2. An attempt is made to collect the user's voice data. Since the raise hand event ends after 0.5 seconds, the ADSP ends the process based on the detection result of the new raise hand event and thus returns to the power P1 again. Therefore, Figure 10 The total power consumption in (b) therein is: P1 * total running duration + P2 * 0.4s + P2 * 0.3s + P2 * 0.5s.
[0283] It can be seen from the comparison diagram that the solution of the present application "purifies" the raise hand event, eliminates the influence of the raise hand event not used for breath wake-up, and can effectively reduce the power consumption. However, it should be understood that Figure 10 This is only to illustrate how the solution of the present application reduces the power consumption in the same scenario, and there are no limitations on the specific scenario examples and specific values.
[0284] The above mainly introduced the method of the embodiments of the present application in combination with the accompanying drawings. It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence, these steps are not necessarily executed in the order shown in the figures. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps. Next, the device of the embodiments of the present application will be introduced in combination with the accompanying drawings.
[0285] Figure 11 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. Figure 11 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. As Figure 11 shown, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, an audio module 270, a speaker 270A, a receiver 270B, a microphone N70C, a headphone interface 270D, a sensor module 280, a camera 293, a display screen 294, etc. Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, an acceleration sensor 280E, etc.
[0286] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0287] Exemplarily, Figure 11The illustrated processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), an audio digital signal processor (ADSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0288] Among them, the controller may be the nerve center and command center of the electronic device 200. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0289] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0290] In the embodiments of the present application, mainly the AP and the ADSP cooperate with each other to complete the voice interaction in the breath wake-up scenario. Moreover, in view of the problem of limited low-power memory space of the ADSP in the present application, the breath wake-up module is subdivided into multiple sub-modules (such as a gesture detection module, a breath wake-up start module, etc.), and they are respectively deployed in the low-power memory space and the non-low-power memory space of the ADSP. In particular, the relevant algorithms for breath wake-up detection are deployed in the non-low-power memory space, so that even mid-range and low-end processors can deploy the breath wake-up detection scheme. On this basis, an additional screening condition is added when determining whether to trigger a breath wake-up event to eliminate the unnecessary power consumption caused by non-breath-wake-up raising-hand events, so as to further reduce the power consumption.
[0291] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0292] In some embodiments, the I2S interface may be used for audio communication. The processor 210 may include multiple groups of I2S buses. The processor 210 may be coupled to the audio module 270 through the I2S bus to enable communication between the processor 210 and the audio module 270.
[0293] In some embodiments, the audio module 270 may transmit an audio signal to the wireless communication module 260 through the I2S interface to implement the function of answering a call through a Bluetooth headset.
[0294] In some embodiments, the PCM interface may also be used for audio communication to sample, quantize, and encode an analog signal. The audio module 270 and the wireless communication module 260 may be coupled through the PCM bus interface.
[0295] In some embodiments, the audio module 270 may also transmit an audio signal to the wireless communication module 260 through the PCM interface to implement the function of answering a call through a Bluetooth headset. It should be understood that both the I2S interface and the PCM interface can be used for audio communication.
[0296] In some embodiments, the UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. The UART interface is typically used to connect the processor 210 and the wireless communication module 260. For example, the processor 210 communicates with the Bluetooth module in the wireless communication module 260 through the UART interface to implement the Bluetooth function. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 through the UART interface to implement the function of playing music through the Bluetooth headset.
[0297] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0298] The electronic device 200 realizes the display function through the GPU, the display screen 294, and the application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change the display information.
[0299] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. In some embodiments, the electronic device 200 may include 1 or N display screens 294, where N is a positive integer greater than 1.
[0300] In the embodiments of the present application, the display screen 294 can be used to display various interfaces, such as interface 101 - interface 108.
[0301] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 200 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0302] The NPU is a neural-network (NN) computing processor. By referring to the biological neural network structure, such as referring to the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 200 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0303] In the embodiments of the present application, it is mainly used for speech recognition and feedback of corresponding results to the user based on the speech recognition results.
[0304] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221. The internal memory 221 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 200 (such as audio data, a phone book, etc.). In addition, the internal memory 221 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0305] In the embodiment of the present application, the relevant model in the breath wake-up detection stage needs to be deployed in the non-low-power memory space of the ADSP, while other models such as the speech recognition model and the question-and-answer large model can be deployed in the memory space of the AP or the internal memory 221. The small program code of the relevant execution process can also be stored in the internal memory 221.
[0306] The electronic device 200 can implement audio functions through the audio module 270, the speaker 270A, the receiver 270B, the microphone N70C, and the application processor, etc. For example, music playback, recording, etc.
[0307] The audio module 270 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 can be disposed in the processor 210, or some functional modules of the audio module 270 can be disposed in the processor 210.
[0308] The speaker 270A, also called a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 200 can listen to music or hands-free calls through the speaker 270A.
[0309] The receiver 270B, also called a "handset", is used to convert an audio electrical signal into a sound signal. When the electronic device 200 answers a call or a voice message, the voice can be received by placing the receiver 270B close to the human ear.
[0310] The microphone N70C, also known as "microphone" and "transmitter", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak close to the microphone N70C with the mouth to input the sound signal into the microphone N70C. The electronic device 200 can be provided with at least one microphone N70C. In some other embodiments, the electronic device 200 can be provided with two microphones N70C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 200 can also be provided with three, four or more microphones N70C, which can collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0311] In the embodiment of the present application, the sound signal of the voice is collected by the microphone, and then the ADSP determines whether to trigger the breath wake-up event based on the collected sound signal, and after triggering, the AP performs voice recognition and further response on the collected sound signal.
[0312] The pressure sensor 280A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 280A can be disposed on the display screen 294. There are many types of pressure sensors 280A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates with conductive materials. When a force acts on the pressure sensor 280A, the capacitance between the electrodes changes. The electronic device 200 determines the intensity of the pressure according to the change of the capacitance. When a touch operation acts on the display screen 294, the electronic device 200 detects the intensity of the touch operation according to the pressure sensor 280A. The electronic device 200 can also calculate the position of the touch according to the detection signal of the pressure sensor 280A. In some embodiments, touch operations with the same touch position but different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0313] In the embodiment of the present application, touch screen operations such as clicks in the interfaces 101-108 can be collected by the pressure sensor 280A.
[0314] The gyroscope sensor 280B can be used to determine the motion posture of the electronic device 200. In some embodiments, the angular velocity of the electronic device 200 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 280B. The gyroscope sensor 280B can be used for anti-shake shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 280B detects the angle of jitter of the electronic device 200, calculates the distance that the lens module needs to compensate according to the angle, and makes the lens offset the jitter of the electronic device 200 through reverse movement to achieve anti-shake. The gyroscope sensor 280B can also be used for navigation and somatosensory game scenarios.
[0315] The acceleration sensor 280E can detect the magnitude of the acceleration of the electronic device 200 in various directions (generally three axes). When the electronic device 200 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0316] In the embodiments of the present application, the gyroscope sensor 280B, the acceleration sensor 280E, and / or other motion sensors can be used to detect the motion data of the electronic device, so as to determine whether a hand-lifting event has occurred based on the motion data.
[0317] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units, due to being based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.
[0318] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here.
[0319] The embodiments of the present application also provide an electronic device, which includes: a plurality of processors, a memory, and a computer program stored in the memory and executable on the plurality of processors. When the plurality of processors execute the computer program, the electronic device can implement the steps in any of the above methods.
[0320] The embodiments of the present application also provide a computer-readable storage medium storing a computer program, and when the computer program is executed by an electronic device, the steps in the above method embodiments can be implemented.
[0321] The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to the photographing device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a portable hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.
[0322] The embodiments of the present application provide a computer program product including a computer program, and when the computer program is executed by an electronic device, the steps in the above method embodiments can be implemented. The computer program includes computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc.
[0323] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0324] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0325] In the embodiments provided by the present application, it should be understood that the disclosed device / equipment and method can be implemented in other ways. For example, the device / equipment embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0326] The unit described as a separation component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0327] It should be understood that when used in the specification and appended claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0328] It should also be understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0329] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0330] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0331] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A voice interaction method, applied to an electronic device, characterized in that: include: Before a first moment, the first processor of the electronic device continuously detects a hand-raising event in a low power consumption mode, the hand-raising event being used to indicate whether the electronic device is lifted, wherein when a hand-raising event is detected, it indicates that the electronic device is lifted, and when a non-hand-raising event is detected, it indicates that the electronic device is not lifted, the power consumption of the electronic device is a first value, the second processor of the electronic device is not awakened, so that the screen of the electronic device is in an off state, the first processor is used to perform breath awakening detection, and the second processor is used to perform voice interaction response to voice data collected by the first processor when awakened by the first processor; At the first moment, the electronic device is lifted, and the second processor is not awakened, so that the screen of the electronic device is in an off state. In response to detecting the hand-raising event, the electronic device switches the first processor from a low power consumption mode to a normal working mode, and starts to determine whether a breath awakening event is triggered in the normal working mode, and continues to obtain the detection result of the hand-raising event in the normal working mode, so that the power consumption of the electronic device changes from the first value to a second value, and the second value is greater than the first value; At a second moment after the first moment, the electronic device is not lifted and the second processor is not awakened, so that the screen of the electronic device is in an off state. In response to detecting a non-hand-raising event, the electronic device switches the first processor from a normal working mode to a low power consumption mode based on the electronic device being converted from a lifted state to a not lifted state, and continues to detect hand-raising events in the low power consumption mode, so that the power consumption of the electronic device changes from the second value to the first value.
2. The method according to claim 1, characterized in that: The method further comprises: Between the first moment and the third moment, the electronic device remains lifted, the first processor operates in a normal working mode, and the second processor is not awakened so that the screen of the electronic device is in an off state, the time interval between the first moment and the second moment is less than the time interval between the first moment and the third moment, and the time interval between the first moment and the third moment is a preset time interval; At the third moment, the electronic device switches the first processor from a normal working mode to a low power consumption mode, and continues to detect hand raising events in the low power consumption mode, so that the power consumption of the electronic device changes from the second value to the first value.
3. The method according to claim 2, characterized in that At the third moment, the power consumption of the electronic device changes from the second value to the first value, including: At the third moment, based on the fact that the electronic device has not collected voice data that meets the first preset condition, the power consumption of the electronic device changes from the second value to the first value, and the first preset condition includes: the signal strength of the voice data collected by the first microphone of the electronic device is greater than or equal to the first preset signal strength threshold, the signal strength of the voice data collected by the second microphone of the electronic device is greater than or equal to the second preset signal strength threshold, and the difference between the signal strength of the voice data collected by the first microphone and the signal strength of the voice data collected by the second microphone is greater than or equal to the preset signal strength difference threshold.
4. The method according to claim 3, characterized in that At the third moment, based on the fact that the electronic device does not collect voice data that meets the first preset condition, the power consumption of the electronic device changes from the second value to the first value, including: Before the first moment, the first processor of the electronic device operates in a low power consumption mode; between the first moment and the third moment, the first processor operates in a normal working mode; The first processor continuously reads the voice data collected by the first microphone and the second microphone from the first moment; At the third moment, the first processor determines that the breath wake-up event is not triggered based on the fact that the electronic device is lifted up and the electronic device has not collected voice data that meets the first preset condition, controls the first processor to switch from the normal working mode to the low power consumption mode, and continues to detect the hand raising event in the low power consumption mode.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: At the fifth moment after the first moment, the second processor is awakened to light up the screen of the electronic device, and the second processor switches to a normal working mode so that the power consumption of the electronic device is a third value, and the third value is greater than the second value.
6. The method according to claim 5, characterized in that The step of turning on the screen of the electronic device at a fifth moment after the first moment, and the power consumption of the electronic device being a third value, wherein the third value is greater than the second value, comprises: At the first moment, the first processor of the electronic device switches from a low power consumption mode to a normal operating mode, so that the power consumption of the electronic device increases from the first value to the second value; At a fourth moment after the first moment, based on the electronic device starting to collect voice data that meets a first preset condition, the first processor starts to continuously read the collected voice data; the first preset condition includes: the signal strength of the voice data collected by the first microphone of the electronic device is greater than or equal to a first preset signal strength threshold, the signal strength of the voice data collected by the second microphone of the electronic device is greater than or equal to a second preset signal strength threshold, and the difference between the signal strength of the voice data collected by the first microphone and the signal strength of the voice data collected by the second microphone is greater than or equal to a preset signal strength difference threshold; At the fifth moment after the fourth moment, the first processor completes reading the first voice data, and the first processor wakes up the second processor based on that the first voice data satisfies the second preset condition, so that the second processor switches from the low power consumption mode to the normal working mode and performs voice recognition and response to the first voice data in the normal working mode, thereby increasing the power consumption of the electronic device from the second value to the third value. Satisfying the second preset condition is used to indicate that the first voice data can be used to trigger a breath wake-up event.
7. The method according to claim 6, characterized in that The method further comprises: At the fifth moment, based on the fact that the first voice data does not satisfy the second preset condition, the first processor switches from the normal working mode to the low power consumption mode, so that the power consumption of the electronic device changes from the second value to the first value.
8. The method according to claim 6, characterized in that The method further comprises: At a sixth moment after the fifth moment, the second processor switches from a normal operating mode to a low power consumption mode so that the screen of the electronic device switches from a bright screen state to an off screen state, and the power consumption of the electronic device switches from the third value to the first value.
9. The method according to claim 8, characterized in that At a sixth moment after the fifth moment, the screen of the electronic device changes from a light-on state to an off state, and the power consumption of the electronic device changes from the third value to the first value, including: At the sixth moment, based on the second processor completing the response to the first voice data, the first processor and the second processor both switch to a low power consumption mode, so that the power consumption of the electronic device changes from the third value to the first value.
10. The method according to claim 6, characterized in that The method further comprises: At the fourth moment, based on the fact that the electronic device has not collected voice data that meets the first preset condition and the first processor has not read voice data that meets the first preset condition, the electronic device continues to collect voice data until the third moment, and the time interval between the third moment and the first moment is the preset time interval.
11. A chip system, applied to electronic equipment, characterized in that: The chip system includes a first processor and a second processor, the first processor is used to perform breath awakening detection, and the second processor is used to respond to voice interaction to voice data collected by the first processor when awakened by the first processor; Before a first moment, the first processor continuously detects a hand-raising event in a low power consumption mode, the hand-raising event being used to indicate whether the electronic device is lifted, wherein when a hand-raising event is detected, it indicates that the electronic device is lifted, and when a non-hand-raising event is detected, it indicates that the electronic device is not lifted, the power consumption of the electronic device is a first value, and the second processor of the electronic device is not awakened, so that the screen of the electronic device is in an off state; At the first moment, the electronic device is lifted, and the second processor is not awakened, so that the screen of the electronic device is in an off state. In response to detecting the hand-raising event, the electronic device switches the first processor from a low power consumption mode to a normal working mode, and starts to determine whether a breath awakening event is triggered in the normal working mode, and continues to obtain the detection result of the hand-raising event in the normal working mode, so that the power consumption of the electronic device changes from the first value to a second value, and the second value is greater than the first value; At a second moment after the first moment, the electronic device is not lifted and the second processor is not awakened, so that the screen of the electronic device is in an off state. In response to detecting a non-hand-raising event, the electronic device switches the first processor from a normal working mode to a low power consumption mode based on the electronic device being converted from a lifted state to a not lifted state, and continues to detect hand-raising events in the low power consumption mode, so that the power consumption of the electronic device changes from the second value to the first value.
12. An electronic device, characterized in that: The electronic device includes a memory, multiple processors, a sensor, an audio acquisition device and a display screen, the sensor is used to collect motion data of the electronic device, the audio acquisition device is used to collect voice of the electronic device, the display screen is used to display an interface, the memory is used to store instructions, and the multiple processors are used to run the instructions stored in the memory, so that the electronic device executes the method as described in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Voice interaction method and related equipment
CN116229953A