Voice wake-up method, electronic device, and chip system
By detecting preset power-on conditions in the first system of the electronic device and powering on the second system, the problem of excessive voice wake-up delay is solved, and the target application is quickly started, and the user experience is improved.
Patent Information
- Application Number
- CN202011056925.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-29
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2040-09-29
AI Technical Summary
In the prior art, during the voice wake-up process of dual-system electronic devices, the second system needs to resume power-up for a long time, resulting in an excessive delay in voice wake-up, affecting the user experience.
When the electronic device is in the off-screen state, the first system remains powered on, detects preset power-on conditions and performs first voice recognition. At the same time, powering on the second system through the power management unit. The second system starts the target application when the power-on is completed and the voice recognition result meets the conditions.
Shorten or eliminate the power-on delay of the second system, reduce user waiting time, and improve user experience.
Smart Images

Figure CN114333854B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a voice wake-up method, electronic equipment, and chip system. Background Art
[0002] As a typical application of artificial intelligence, voice assistants are becoming increasingly popular in electronic devices. As the entry point to voice assistants, voice wake-up features, such as performance and latency, significantly impact the user experience. Voice wake-up involves pre-setting a wake-up word in an electronic device. When a user issues a voice command containing that word, the voice assistant application wakes up from its dormant state and responds, significantly improving the efficiency of human-computer interaction.
[0003] At present, for electronic devices with high power consumption requirements, dual systems can be used to improve the battery life of electronic devices. For different scenarios, electronic devices can choose the corresponding system to work to achieve the effect of reducing power consumption and improving battery life. Taking smart watches as an example, smart watches can be driven by the first system and the second system. For lightweight application scenarios (such as ordinary dial display, local music playback, etc.), smart watches can run the first system. For heavy-loaded application scenarios (such as 3D dials, WeChat third-party applications, etc.), smart watches can switch from the first system to the second system.
[0004] Typically, when a smartwatch is in the screen-off state, the first system is powered on and the second system is powered off. The voice wake-up process for a smartwatch involves the user sending a voice message to wake up the voice assistant app on the smartwatch. The first system then performs voice recognition on the message. After the first system successfully recognizes the message, power is restored to the second system. The second system then performs voice recognition on the message. After successful voice recognition, the second system launches the voice assistant app, completing the voice wake-up process.
[0005] However, during the voice wake-up process, restoring power to the second system takes a long time, typically around 200ms to 1000ms. This shows that while dual systems can reduce power consumption in electronic devices, launching the voice assistant application involves restoring power to the second system, which is powered off. The time required to restore power to the second system increases the latency of the voice wake-up function in the electronic device, resulting in excessive wake-up latency. Summary of the Invention
[0006] The present application provides a voice wake-up method, electronic device, and chip system, which can reduce the voice wake-up delay.
[0007] To achieve the above objectives, this application adopts the following technical solutions:
[0008] In a first aspect, an embodiment of the present application provides an electronic device, which includes a first system, a second system and a power management system. When the electronic device is in a screen-off state, the first system is in a power-on state and the second system is in a power-off state; the first system is used to light up the screen of the electronic device; the first system is also used to detect a preset power-on condition when receiving sound information, and perform a first voice recognition on the sound information; the first system is also used to power on the second system through a power management unit when the preset power-on condition is detected; the second system is used to start a target application when power-on is completed and the first voice recognition result meets the first preset condition.
[0009] When the first system in the above-mentioned electronic device receives sound information, it can not only perform a first voice recognition on the sound information, but also detect a preset power-on condition, and when the preset power-on condition is detected, the second system can be powered on. In this case, the first voice recognition process and the second system power-on process do not interfere with each other and are independent of each other, so they can be carried out simultaneously. Therefore, compared with the related art, the voice wake-up process of the electronic device can shorten or eliminate the delay caused by powering on the second system, thereby greatly reducing the voice wake-up delay. After the user issues a voice command containing the wake-up word, the electronic device can start the target application more quickly, reduce the user's waiting time, and improve the user experience.
[0010] In some embodiments, if the first system detects that the posture of the electronic device is a preset posture, the screen of the electronic device is turned on.
[0011] For example, the first system may include a sensor control center processor and an accelerometer / gyroscope unit. The accelerometer / gyroscope unit may be used to detect the posture of the electronic device. The accelerometer / gyroscope unit may transmit the detected posture of the electronic device to the sensor control center processor. When the posture of the electronic device reaches a preset posture, the sensor control center processor may transmit a screen-lighting signal to the display screen, requesting that the screen be illuminated.
[0012] For another example, the accelerometer and gyroscope unit may send a screen-lighting signal to the sensor control center processor when the electronic device is in a preset posture. The sensor control center processor requests the display screen to light up the screen according to the screen-lighting signal.
[0013] For another example, the accelerometer and gyroscope unit may send a screen-lighting signal to the display screen when the electronic device is in a preset posture. The display screen lights up the screen in response to the screen-lighting signal.
[0014] In some embodiments, if the first system detects that the user's body is close to the electronic device, the screen of the electronic device is lit.
[0015] For example, the first system may also include devices such as a proximity light sensor and an ultrasonic device. The first system may use the proximity light sensor and ultrasonic device to detect whether a user's limb is in proximity to the electronic device. If the proximity light sensor and ultrasonic device detects that the user's limb is in proximity to the electronic device, the proximity light sensor and ultrasonic device transmit a screen-lighting signal to the sensor control center processor. The sensor control center processor responds to the screen-lighting signal and requests the display screen to light up. Alternatively, the proximity light sensor and ultrasonic device transmit a screen-lighting signal to the display screen. The display screen responds to the screen-lighting signal and lights up the screen.
[0016] In some embodiments, if the first system detects that the user is looking at the screen of the electronic device, the screen of the electronic device is turned on.
[0017] For example, the first system may further include a camera that is in an active state when the electronic device is in the screen-off state. The first system may detect whether the user is looking at the screen of the electronic device through the camera. If the camera detects that the user is looking at the screen of the electronic device, it may send a screen-lighting signal to the sensor control center processor. The camera detecting that the user is looking at the screen of the electronic device indicates that the user is likely to use the electronic device. In response to the screen-lighting signal, the sensor control center processor requests the display screen to light up the screen.
[0018] In combination with the first aspect, in some embodiments, the above-mentioned preset power-on conditions include one or more of the following: the sound information contains human voice; there is a preset user behavior; the posture of the electronic device meets the preset posture condition.
[0019] The first system can detect, using the accelerometer-gyroscope unit, whether the posture of the electronic device satisfies a preset posture condition. For example, the preset posture condition can be: the electronic device maintains the preset posture for a time greater than a time threshold. Correspondingly, the first system can detect, using the accelerometer-gyroscope unit, whether the preset posture is maintained for a time greater than the time threshold.
[0020] For example, if the preset power-on condition includes the electronic device's posture meeting the preset posture condition, the first system detects whether the electronic device's posture meets the preset posture condition using the accelerometer-gyroscope unit. If the electronic device's posture meets the preset posture condition, the accelerometer-gyroscope unit may send a power-on signal to the power management unit. The power management unit, in response to the power-on signal, powers on the second system.
[0021] The above-mentioned time threshold can be: the length of time that the electronic device is in a preset posture corresponding to the user behavior of checking the time on the electronic device. The time threshold can be determined based on previously acquired information. For example, the above-mentioned first time threshold can be determined based on the length of time between the first time and the second time acquired in the past. The first time is the time when the electronic device turns on the screen, and the second time is the time when the first system performs voice activity detection on the acquired sound information after the electronic device turns on the screen and detects a human voice.
[0022] Exemplarily, the preset user behavior includes one or more of the following: the user's body approaching the electronic device; the user looking at the screen of the electronic device. For example, the first system may detect whether the user's body is approaching the electronic device using a proximity light sensor, an ultrasonic device, or other device. The first system may also detect whether the user is looking at the screen of the electronic device using a camera.
[0023] In combination with the first aspect, in some embodiments, the second system is specifically used to: perform second voice recognition on the sound information when power-on is completed and the first voice recognition result meets the first preset condition; and start the target application corresponding to the second voice recognition result when the second voice recognition result meets the second preset condition.
[0024] In combination with the first aspect, in some embodiments, the second system is further used to: generate a prompt message for prompting the user to re-perform voice wake-up when the second voice recognition result does not meet the second preset condition.
[0025] The prompt message can inform the user that the voice wake-up has failed and can remind the user to send a voice message containing the preset wake-up word again. The first system then performs a first voice recognition on the voice message sent again by the user. If the first voice recognition result meets the first preset condition, the second system performs a second voice recognition on the voice message sent again by the user. If the second voice recognition result meets the second preset condition, the second system launches the target application corresponding to the second voice recognition result.
[0026] In a second aspect, an embodiment of the present application provides a voice wake-up method, which is applied to an electronic device. The electronic device includes a first system, a second system and a power management unit. When the electronic device is in a screen-off state, the first system is in a power-on state and the second system is in a power-off state. The above method includes: the first system lights up the screen of the electronic device; if the first system receives sound information, it detects a preset power-on condition and performs a first voice recognition on the sound information; if the first system detects the preset power-on condition, it powers on the second system through the power management unit; if the power-on of the second system is completed and the first voice recognition result meets the first preset condition, the second system starts the target application.
[0027] In the above-mentioned voice wake-up method, when receiving sound information, the first system can not only perform the first voice recognition on the sound information, but also detect the preset power-on condition. Once the preset power-on condition is detected, the second system can be powered on. Therefore, compared with the related art, the above-mentioned voice wake-up method can shorten or eliminate the delay caused by powering on the second system, thereby greatly reducing the voice wake-up delay. After the user issues a voice command containing the wake-up word, the electronic device can quickly start the target application, reducing the user's waiting time and improving the user experience.
[0028] In one application scenario, the first system may light up the screen of the electronic device when detecting that the posture of the electronic device is a preset posture.
[0029] For example, the first system may include a sensor control center processor and an accelerometer / gyroscope unit. The accelerometer / gyroscope unit may be used to detect the posture of the electronic device. The accelerometer / gyroscope unit may transmit the detected posture of the electronic device to the sensor control center processor. When the posture of the electronic device reaches a preset posture, the sensor control center processor may transmit a screen-lighting signal to the display screen, requesting that the screen be illuminated.
[0030] For another example, the accelerometer and gyroscope unit may send a screen-lighting signal to the sensor control center processor when the electronic device is in a preset posture. The sensor control center processor requests the display screen to light up the screen according to the screen-lighting signal.
[0031] For another example, the accelerometer and gyroscope unit may send a screen-lighting signal to the display screen when the electronic device is in a preset posture. The display screen lights up the screen in response to the screen-lighting signal.
[0032] In another application scenario, the first system may light up the screen of the electronic device when it detects that the user's body is close to the electronic device.
[0033] For example, the first system may also include devices such as a proximity light sensor and an ultrasonic device. The first system may use the proximity light sensor and ultrasonic device to detect whether a user's limb is in proximity to the electronic device. If the proximity light sensor and ultrasonic device detects that the user's limb is in proximity to the electronic device, the proximity light sensor and ultrasonic device transmit a screen-lighting signal to the sensor control center processor. The sensor control center processor responds to the screen-lighting signal and requests the display screen to light up. Alternatively, the proximity light sensor and ultrasonic device transmit a screen-lighting signal to the display screen. The display screen responds to the screen-lighting signal and lights up the screen.
[0034] In yet another application scenario, the first system may light up the screen of the electronic device when detecting that the user is looking at the screen of the electronic device.
[0035] For example, the first system may further include a camera that is in an active state when the electronic device is in the screen-off state. The first system may detect whether the user is looking at the screen of the electronic device through the camera. If the camera detects that the user is looking at the screen of the electronic device, it may send a screen-lighting signal to the sensor control center processor. The camera detecting that the user is looking at the screen of the electronic device indicates that the user is likely to use the electronic device. In response to the screen-lighting signal, the sensor control center processor requests the display screen to light up the screen.
[0036] In conjunction with the second aspect, in some embodiments, the preset power-on condition is a condition that determines that the user is likely to perform voice wake-up. The preset power-on condition can be one or more of the following: the sound information contains a human voice; a preset user behavior occurs; or the posture of the electronic device meets a preset posture condition.
[0037] The first system can detect, using the accelerometer-gyroscope unit, whether the posture of the electronic device satisfies a preset posture condition. For example, the preset posture condition is that the electronic device maintains the preset posture for a time greater than a time threshold. Accordingly, the first system can detect, using the accelerometer-gyroscope unit, whether the preset posture is maintained for a time greater than the time threshold.
[0038] Illustratively, the aforementioned preset user behaviors include one or more of the following: the user's physical presence approaching the electronic device; the user's gaze at the screen of the electronic device. For example, the first system may detect whether the user's physical presence is approaching the electronic device using a proximity light sensor, ultrasonic device, or other device. The first system may also detect whether the user is gazed at the screen of the electronic device using a camera.
[0039] In combination with the second aspect, in some embodiments, if the second system is powered on and the first voice recognition result meets the first preset condition, the second system starts the target application, including: if the second system is powered on and the first voice recognition result meets the first preset condition, the second system performs a second voice recognition on the sound information; if the second voice recognition result meets the second preset condition, the second system starts the target application corresponding to the second voice recognition result.
[0040] In combination with the second aspect, in some embodiments, the above method also includes: if the second voice recognition result does not meet the second preset condition, the second system generates a prompt message for prompting the user to re-perform voice wake-up.
[0041] The prompt message can inform the user that the voice wake-up has failed and can remind the user to send a voice message containing the preset wake-up word again. The first system then performs a first voice recognition on the voice message sent again by the user. If the first voice recognition result meets the first preset condition, the second system performs a second voice recognition on the voice message sent again by the user. If the second voice recognition result meets the second preset condition, the second system launches the target application corresponding to the second voice recognition result.
[0042] In a third aspect, an embodiment of the present application provides a voice wake-up device, which is applied to an electronic device, and the electronic device includes a first system, a second system and a power management unit. When the electronic device is in the screen-off state, the first system is in the power-on state and the second system is in the power-off state. The voice wake-up device includes: a screen lighting unit for lighting up the screen of the electronic device. A detection and recognition unit for detecting a preset power-on condition when the first system receives sound information, and performing a first voice recognition on the sound information. A power-on unit for powering on the second system through the power management unit when the first system detects a preset power-on condition. An application startup unit for starting the target application when the power-on of the second system is completed and the first voice recognition result meets the first preset condition.
[0043] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising: one or more processors, a memory, and a display screen; the memory and the display screen are coupled to the one or more processors, the memory is used to store computer program code, and the computer program code includes computer instructions; when the one or more processors execute the computer instructions, the electronic device executes the method as described in any one of the first aspects.
[0044] In a fifth aspect, an embodiment of the present application provides a chip system, comprising a processor coupled to a memory, wherein the processor executes a computer program stored in the memory to implement the method described in any one of the first aspects. The chip system may be a single chip or a chip module composed of multiple chips.
[0045] In a sixth aspect, an embodiment of the present application provides a chip system, comprising a memory and a processor, wherein the processor executes a computer program stored in the memory to implement the method described in any one of the first aspects. The chip system may be a single chip or a chip module composed of multiple chips.
[0046] In a seventh aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the electronic device to execute any of the methods described in the first aspect above.
[0047] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method as described in any one of the first aspects is implemented.
[0048] It is understandable that the voice wake-up device described in the third aspect, the electronic device described in the fourth aspect, the chip systems described in the fifth and sixth aspects, the computer program product described in the seventh aspect, and the computer-readable storage medium described in the eighth aspect are all used to execute the method provided in the second aspect. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram of the voice wake-up process provided for related technologies;
[0050] Figure 2 for Figure 1 A schematic diagram of the delay of the voice wake-up process provided in the embodiment;
[0051] Figure 3 Schematic diagram of the voice wake-up process provided for related technologies;
[0052] Figure 4 for Figure 3 A schematic diagram of the delay of the voice wake-up process provided in the embodiment;
[0053] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0054] Figure 6 A schematic diagram of a scenario in which the smartwatch raises its wrist to light up the screen, provided in an embodiment of the present application;
[0055] Figure 7 A schematic diagram of the delay of the voice wake-up process provided in an embodiment of the present application;
[0056] Figure 8 A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0057] Figure 9 A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0058] Figure 10 A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0059] Figure 11 A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0060] Figure 12A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0061] Figure 13 A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0062] Figure 14 A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0063] Figure 15 A flowchart of the voice wake-up method provided in an embodiment of the present application;
[0064] Figure 16 A schematic diagram of the structure of a voice wake-up device provided in an embodiment of the present application;
[0065] Figure 17 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0067] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0068] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0069] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0070] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0071] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0072] In addition, the “plurality” mentioned in the embodiments of the present application should be interpreted as two or more.
[0073] The steps involved in the voice wake-up method provided in the embodiments of the present application are only examples. Not all steps must be performed, or not all information or content in the message is required. They can be increased or decreased as needed during use.
[0074] In the embodiments of the present application, the same step or steps or messages with the same function in different embodiments can be referenced to each other.
[0075] The business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field will know that with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0076] Before explaining the embodiments of the present application in detail, the application scenarios involved in the embodiments of the present application are first explained.
[0077] Voice wake-up refers to pre-setting a wake-up word in an electronic device. When a user issues a voice command containing that wake-up word, the voice assistant application is awakened from its dormant state and responds, significantly improving the efficiency of human-computer interaction. As the entry point to the voice assistant application, voice wake-up's performance and wake-up latency significantly impact the user experience. The optimal value for user experience of voice wake-up latency is generally 650ms.
[0078] Voice wake-up can be performed on electronic devices with dual systems (which can be called the first system and the second system). For lightweight application scenarios (such as ordinary dial display, local music playback, etc.), electronic devices can run the first system. For heavy-load application scenarios (such as 3D dials, WeChat third-party applications, etc.), electronic devices can switch from the first system to the second system.
[0079] The following describes the voice wake-up process in the related art by taking an electronic device as a smart watch, a first system as a sensor hub system, and a second system as an application processor (AP) system as an example.
[0080] When the smartwatch is in the screen-off state, the Sensor Hub system is powered on and the AP system is powered off. For example, the powered-off state may be that the CPU core of the AP system is powered off, i.e., the Suspend to RAM (STR) state. In another example, the powered-off state may be that the AP system is completely powered off, i.e., the Fast Suspend Resume (FSR) state.
[0081] Figure 1 A diagram of a voice wake-up process provided for related technologies. Figure 1 , after the Sensor Hub system collects sound information through the digital microphone (Digital Mic), it sends the sound information to the Sensor Hub system. The sound information may include: human voice information issued by the user to wake up the voice assistant application in the smart watch (for example: Hello Xiaoyi). The Sensor Hub system performs voice activity detection (Voice Activity Detection, VAD) on the sound information. If it is detected that the sound information contains a human voice, the sound information is subjected to a first voice recognition. After the first voice recognition result meets the preset conditions (that is, the first voice recognition is successful), the Sensor Hub system sends an instruction to the power management unit to power on the second system. The power management unit responds to the instruction and restores power to the AP system.
[0082] Restoring the AP system from the STR state means powering on the AP system's CPU core, while restoring the AP system from the FSR state means powering on the AP system's system on chip (SOC). After the AP system is powered on from the STR state / FSR state, the AP system performs a second voice recognition on the sound information. If the second voice recognition result meets the preset conditions (i.e., the second voice recognition is successful), the AP system launches the voice assistant application to complete voice wake-up.
[0083] Exemplarily, the first voice recognition and the second voice recognition are both used to identify whether the sound information contains a preset wake-up word. If the wake-up word is contained, the corresponding voice recognition is successful, and the next voice recognition is entered or the voice assistant application is started.
[0084] Among them, both the first voice recognition and the second voice recognition can be implemented by a voice recognition model, and / or a related voice recognition algorithm. Moreover, compared with the first voice recognition, the voice recognition model or voice recognition algorithm of the second voice recognition is generally a larger and more accurate model or algorithm. For example, in the case where the sound information contains a pseudo wake-up word that is similar to the preset wake-up word, the first voice recognition is likely to recognize that the sound information contains the preset wake-up word, while the second voice recognition can recognize that the sound information does not contain the threshold wake-up word. In this way, the second voice recognition can more accurately identify whether the sound information contains the wake-up word than the first voice recognition, thereby reducing the false wake-up rate.
[0085] Figure 2 for Figure 1 Schematic diagram of delay analysis of the voice wake-up process provided by the embodiment. Figure 2 The user sends the voice wake-up call from T0 to T2, and the Sensor Hub system continuously acquires the voice information from T0 to T2. The Sensor Hub system begins the aforementioned VAD processing at T0 and detects a human voice in the voice information at T1, meaning that the Sensor Hub system completes the VAD processing between T0 and T1. After completing the VAD processing, the Sensor Hub system begins the first voice recognition process on the acquired voice information. Because the user had not yet sent the complete voice wake-up call "Hello Xiaoyi" before T2, the first voice recognition process failed between T1 and T2. Only after the user sends the complete voice wake-up call "Hello Xiaoyi" at T2 does the Sensor Hub system complete the first voice recognition process between T2 and T3. The power management unit then powers on the AP system between T3 and T4. The CPU core power-up process takes approximately 200ms, and the SoC chip power-up process takes approximately 1000ms. After the AP system is powered on, the AP system performs a second voice recognition on the sound information between T4 and T5. After the second voice recognition is successful, the AP system starts the voice assistant application between T5 and T6 to complete the voice wake-up.
[0086] As can be seen, powering up the AP system takes approximately 200ms to 1000ms, which adds an additional 200ms to 1000ms of latency to the voice wake-up process. After the user speaks a complete voice command containing the wake-up word, it takes a long time for the smartwatch to launch the voice assistant application. This delay is far greater than the optimal value for voice wake-up latency, resulting in a poor user experience.
[0087] To solve the above problem of long voice wake-up delay, the related art provides another voice wake-up method. Figure 3 and Figure 4 , compared to Figure 1 The voice wake-up process in the embodiment removes the second voice recognition. After the AP system is powered on at time T4, the AP system directly starts the voice assistant application (the corresponding time is T4 to T7). It can be seen that the voice wake-up process in this embodiment can save the time T4 to T5 required for the second voice recognition and reduce the voice wake-up delay. However, omitting the second voice recognition will increase the false wake-up rate of voice wake-up, reduce the wake-up performance of the electronic device, and affect the user experience.
[0088] Based on the above problems, an embodiment of the present application provides an electronic device and a voice wake-up method applied to the electronic device. The electronic device includes a first system, a second system and a power management unit. When the electronic device is in the screen-off state, the first system is in the power-on state and the second system is in the power-off state. After the first system turns on the screen of the electronic device, if it receives sound information, it detects the preset power-on condition and performs a first voice recognition on the sound information. If the first system detects the preset power-on condition, it powers on the second system. If the power-on of the second system is completed and the first voice recognition result meets the preset condition, the second system starts the target application.
[0089] In the electronic device and voice wake-up method provided in the embodiments of the present application, the first system can simultaneously detect the preset power-on condition and perform the first voice recognition on the sound information. Furthermore, upon detecting the preset power-on condition, the second system can be powered on. The first voice recognition process and the second system power-on process do not interfere with each other and are independent of each other, so they can be performed simultaneously. Therefore, compared with related technologies, the above-mentioned electronic device and voice wake-up method can shorten or eliminate the voice wake-up delay corresponding to the above-mentioned T3-T4.
[0090] The electronic device provided in the embodiments of the present application is described in detail below.
[0091] The electronic devices involved in the embodiments of the present application may include but are not limited to smart watches, mobile phones, personal digital assistants (PDAs), tablet computers, portable devices (for example, portable computers), personal computers (PCs), and other devices that can have dual systems, and the embodiments of the present application are not limited to this.
[0092] Figure 5 This is a structural diagram of an electronic device 100 provided in an embodiment of the present application. Figure 5 The electronic device 100 may include a first system 110, a second system 120, and a power management unit 140. When the electronic device 100 is in the screen-off state, the first system 110 is in a power-on state, and the second system 120 is in a power-off state.
[0093] The first system 110 is used to light up the screen of the electronic device 100 .
[0094] In one possible manner, if the first system 110 detects that the posture of the electronic device 100 is a preset posture, the screen of the electronic device 100 is turned on.
[0095] See also Figure 5 The electronic device 100 further includes a display screen 130. The first system 110 may include a sensor control center processor 111 and an accelerometer and gyroscope unit 112. The accelerometer and gyroscope unit 112 may be used to detect the posture of the electronic device 100. The sensor control center processor 111 and the accelerometer and gyroscope unit 112 may be integrated into a single chip or device.
[0096] In one scenario, the accelerometer and gyroscope unit 112 may send the detected posture of the electronic device 100 to the sensor control center processor 111. When the posture of the electronic device 100 is a preset posture, the sensor control center processor 111 sends a screen-on signal to the display screen 130 to request the screen to be turned on.
[0097] In another scenario, when the posture of the electronic device 100 is a preset posture, the accelerometer and gyroscope unit 112 may send a screen-lighting signal to the sensor control center processor 111. The sensor control center processor 111 requests the display screen 130 to light up the screen according to the screen-lighting signal.
[0098] In another scenario, the accelerometer and gyroscope unit 112 may send a screen-lighting signal to the display screen 130 when the posture of the electronic device 100 is a preset posture. The display screen 130 lights up the screen in response to the screen-lighting signal.
[0099] For example, the accelerometer / gyroscope unit 112 may be a chip that integrates both accelerometer Acc and gyroscope Gyro functions. This chip may include an accelerometer, a gyroscope, and a processor. The accelerometer and gyroscope are used to monitor the posture of the electronic device 100 and transmit the monitoring data to the processor. The processor is configured to transmit the aforementioned screen-lighting signal to the sensor control center processor 111 or the display screen 130 when the electronic device 100 is in a preset posture.
[0100] like Figure 6 As shown, the smartwatch is in the screen-off state. After the user raises their wrist, the Sensor Hub processor responds to the screen-on signal to light up the screen and display the current time. If the electronic device 100 is a mobile phone or tablet computer, it can be set to be in the screen-off state when the mobile phone or tablet computer is placed horizontally. When the accelerometer-gyroscope unit 112 detects that the mobile phone or tablet computer is tilted or placed vertically, the mobile phone or tablet computer lights up the screen.
[0101] The preset posture can be set based on experience. For example, the preset posture can be determined based on the posture of the electronic device 100 when the user uses the electronic device 100. For example, the preset posture can be that the electronic device 100 is in a horizontal posture, or it can be that the electronic device 100 is in an inclined posture, which is not limited to this embodiment of the present application.
[0102] In yet another possible embodiment, if the first system 110 detects that the user's body is close to the electronic device 100 , the screen of the electronic device 100 is turned on.
[0103] For example, the first system 110 may further include a proximity light sensor, an ultrasonic device, or other devices. The first system 110 may use the proximity light sensor, ultrasonic device, or other devices to detect whether a user's limb is in proximity to the electronic device 100. If the proximity light sensor, ultrasonic device, or other devices detects that a user's limb is in proximity to the electronic device 100, the proximity light sensor, ultrasonic device, or other devices may send a screen-lighting signal to the sensor control center processor 111. In response to the screen-lighting signal, the sensor control center processor 111 may request the display screen 130 to light up the screen.
[0104] For example, if a proximity light sensor, ultrasonic device or other device detects that a user's body is close to the electronic device 100, a screen-lighting signal may be sent to the display screen 130. The display screen 130 responds to the screen-lighting signal and lights up the screen.
[0105] In yet another possible manner, if the first system 110 detects that the user is gazing at the screen of the electronic device 100 , the screen of the electronic device 100 is illuminated.
[0106] For example, the first system 110 may further include a camera, which is in an active state when the electronic device 100 is in the screen-off state. The first system 110 may detect whether the user is looking at the screen of the electronic device 100 through the camera. If the camera detects that the user is looking at the screen of the electronic device 100, it may send a screen-lighting signal to the sensor control center processor 111. The camera detecting that the user is looking at the screen of the electronic device 100 indicates that the user is likely to use the electronic device 100. In response to the screen-lighting signal, the sensor control center processor 111 requests the display screen 130 to light up the screen.
[0107] For example, when the camera detects that the user is looking at the screen of the electronic device 100, it can also send a screen-lighting signal to the display screen 130. In response to the screen-lighting signal, the display screen 130 lights up the screen.
[0108] After the first system 110 lights up the screen of the electronic device 100, if sound information is received, the first system 110 is further configured to detect a preset power-on condition and perform a first voice recognition on the sound information when the sound information is received.
[0109] The preset power-on condition is a condition that determines that the user is likely to perform voice wake-up. In some embodiments, the preset power-on condition may include one or more of the following: the sound information contains a human voice; a preset user behavior occurs; or the posture of the electronic device 100 meets a preset posture condition.
[0110] For example, the preset posture condition can be set based on experience to indicate that the user may need to use the electronic device 100. For example, the preset posture condition can be: the posture of the electronic device 100 is maintained in the preset posture for a time greater than a time threshold. In some embodiments, the first system 110 can detect whether the preset posture is maintained for a time greater than the time threshold using the accelerometer-gyroscope unit 112.
[0111] The time threshold may be: the duration of time the electronic device 100 is in a preset posture corresponding to the user behavior of checking the time through the electronic device 100. The time threshold may be determined based on previously acquired information. For example, the first time threshold may be determined based on the duration between the first time and the second time acquired previously. The first time is the time when the screen of the electronic device 100 is turned on, and the second time is the time when the first system detects a human voice by performing voice activity detection on the acquired sound information after the screen of the electronic device 100 is turned on.
[0112] Taking smartwatches as an example, from the perspective of voice wake-up, after the user raises their wrist to light up the smartwatch screen (referred to as raise wrist to light up) (without touch operation), user behaviors can be divided into two types: checking the time and voice wake-up. Among them, checking the time is a high-frequency scenario. After raising the wrist to light up the screen, the user's wrist posture will remain fixed for a period of time (for example, T hold1 ) to check the time. Voice wake-up is a low-frequency scenario. After raising the wrist to light up the screen, the user's wrist posture will remain fixed for a period of time (for example, T hold2 , and T hold2 >T hold1 ), and then the user speaks the wake-up word to start the voice assistant application. For example, the electronic device 100 lights up the screen at time T00 and detects a human voice at time T0, then the time length corresponding to T00 to T0 can be the above T hold2 The electronic device 100 can record multiple T hold2 , according to the multiple T hold2 The first time threshold is determined.
[0113] For example, the preset user behavior can be set based on experience, indicating that the user is likely to use the electronic device 100. For example, the preset user behavior can include one or more of the following: the user's body is close to the electronic device 100; the user is looking at the screen of the electronic device 100.
[0114] The first system 110 may detect whether the user's body is close to the electronic device 100 through a proximity light sensor, an ultrasonic device, or the like. The first system 110 may detect whether the user is looking at the screen of the electronic device 100 through a camera.
[0115] In some embodiments, the sensor control center processor 111 may include one or more processing units. For example, the sensor control center processor 111 may include a control unit, a digital signal processor (DSP), a voice activity detection unit, etc. The different processing units may be independent devices or integrated into one or more processors.
[0116] For example, the first system 110 may collect sound information using the digital microphone of the electronic device 100 and detect whether the sound information contains a human voice using the voice activity detection unit. If the sound information contains a human voice, it indicates that the user may want to wake up the electronic device 100 via a voice command. The sensor control center processor 111 may perform a first voice recognition on the sound information using the DSP unit. If the sound information does not contain a human voice, the voice activity detection unit continues to detect the sound information collected by the digital microphone.
[0117] It should be noted that when the voice activity detection unit detects whether the sound information contains a human voice, the sound information may not contain a complete wake-up word. At this time, the first system 110 can perform a first voice recognition on the currently collected sound information, but the first voice recognition result usually cannot meet the preset conditions, that is, the first voice recognition is unsuccessful. However, when the digital microphone collects sound information containing a complete wake-up word, and the wake-up word is consistent with the preset wake-up word, the first system 110's first voice recognition result for the sound information may meet the preset conditions, that is, the first voice recognition is successful.
[0118] In some embodiments, the first system 110 may perform first speech recognition on the sound information when the sound information contains a human voice. The first system 110 may detect information such as the frequency, energy, phase, and amplitude of the sound information using a relevant algorithm or model to determine whether the sound information contains a human voice.
[0119] For example, if the sound information does not contain a human voice, the first system 110 may not perform the first voice recognition on the sound information, but may continue to acquire sound information until the sound information contains a human voice, and then perform the first voice recognition on the sound information. This can reduce the energy consumption of the first system 110 and increase the battery life of the electronic device.
[0120] The first system 110 is further configured to power on the second system 120 via the power management unit 140 when a preset power-on condition is detected.
[0121] For example, when the first system 110 detects a preset power-on condition, it may send a power-on signal to the power management unit 140. The power management unit 140 powers on the second system 120 in response to the power-on signal.
[0122] The second system 120 is configured to start a target application when power-on is completed and the first voice recognition result meets a first preset condition.
[0123] In some embodiments, the target application may be an application corresponding to the first speech recognition result.
[0124] For example, the first preset condition may be that the first voice recognition result is that the sound information contains a preset wake-up word. In this case, the second system 120 may launch the target application corresponding to the first voice recognition result, that is, the second system 120 launches the target application corresponding to the preset wake-up word. For example, if the first voice recognition result is that the sound information contains the preset wake-up word "Hello Xiaoyi", the second system 120 launches the voice assistant application corresponding to the preset wake-up word "Hello Xiaoyi".
[0125] In some embodiments, when the second system 120 needs to perform second voice recognition on voice information, the target application can be the application corresponding to the second voice recognition result.
[0126] Specifically, the second system 120 is configured to: perform second voice recognition on voice information when the power-on is completed and the first voice recognition result meets the first preset condition; and if the second voice recognition result meets the second preset condition, start the target application corresponding to the second voice recognition result.
[0127] Exemplarily, the second preset condition can be that the second voice recognition result is that the voice information contains a preset wake-up word. At this time, the second system 120 can start the target application corresponding to the second voice recognition result, that is, the second system 120 starts the target application corresponding to the preset wake-up word. For example, if the second voice recognition result is that the voice information contains a preset wake-up word "Hello, Xiaoyi", the second system 120 starts the voice assistant application corresponding to the preset wake-up word "Hello, Xiaoyi".
[0128] It should be noted that the first preset condition and the second preset condition can be the same or different, and the embodiments of the present application do not limit this. Exemplarily, both the first preset condition and the second preset condition can be that the voice information contains a preset wake-up word, and the similarity between the wake-up word extracted from the voice information and the preset wake-up word can be different.
[0129] See Figure 5 The second system 120 may include an application processor 121. The application processor 121 may include one or more processing units. For example, the application processor 121 may include a central processing unit (CPU), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a memory, a video codec, a DSP, a baseband processor, and / or a neural network processor (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0130] Exemplarily, when the power-on of the second system 120 is completed and the first voice recognition result meets the first preset condition, the application processor 121 performs second voice recognition on the voice information through the DSP unit. If the second voice recognition result meets the second preset condition, the application processor 121 starts the target application corresponding to the second voice recognition result.
[0131] Optionally, the second system 120 is further configured to generate a prompt message for prompting the user to perform voice wake-up again if the second voice recognition result does not meet a second preset condition.
[0132] For ease of description, the following refers to the first voice recognition result that meets the first preset condition as first voice recognition success, the first voice recognition result that does not meet the first preset condition as first voice recognition failure, the second voice recognition result that meets the second preset condition as second voice recognition success, and the second voice recognition result that does not meet the second preset condition as second voice recognition failure.
[0133] The prompt message can inform the user that the voice wake-up failed and can prompt the user to send a voice message containing the preset wake-up word again. The first system 110 then performs a first voice recognition on the voice message sent again by the user. If the first voice recognition is successful, the second system 120 performs a second voice recognition on the voice message sent again by the user. If the second voice recognition is successful, the second system 120 launches the target application corresponding to the second voice recognition result.
[0134] Figure 7 for Figure 5 Schematic diagram of delay analysis of the voice wake-up process provided in the embodiment.
[0135] See also Figure 7 , the time when the user sends a sound message containing a complete preset wake-up word is T0-T2. The first system 110 performs VAD processing from T0 to T1 to detect whether the sound message contains a human voice. If the sound message contains a human voice, the first system 110 starts the first voice recognition of the sound message at T1. Since the sound message does not contain the complete wake-up word "Hello Xiaoyi" at this time, the first voice recognition fails. After the sound message containing the complete wake-up word "Hello Xiaoyi" is obtained at T2, the first system 110 completes the first voice recognition of the sound message from T2 to T3. After the first system 110 detects the preset power-on condition, the power management unit 140 powers on the second system 120 from T3' to T4'. After the power-on is completed, the second system 120 completes the second voice recognition of the sound message from T4' to T5'. After the second voice recognition is successful, the second system 120 starts the target application (e.g., a voice assistant application) from T5' to T6'.
[0136] T3' can be any time between T1 and T3. Figure 7 The fact that T3' is located between T2 and T3 is merely an example. In other embodiments, T3' may also be located between T1 and T2.
[0137] therefore, Figure 7 The voice wake-up process shown is relative to Figure 2 The voice wake-up process shown can reduce the delay caused by powering on the second system 120, that is, it can reduce Figure 2 The delay caused by the time length corresponding to T3 to T4 shown. If T4' time is T3 time or before T3 time, then Figure 7 The voice wake-up process shown can completely eliminate the delay caused by powering up the second system 120, that is, the wake-up delay can be reduced to Figure 2 The duration of T3 to T4 shown in the figure; if T4' time is after T3 time, then Figure 7 The voice wake-up process shown can reduce a portion of the delay caused by powering up the second system 120, that is, the wake-up delay can be reduced. Figure 7 The duration corresponding to T3'~T3 is shown.
[0138] For the case where the electronic device 100 only performs the first voice recognition on the sound information, remove Figure 7 The T4'~T5' part is sufficient. For the reduction of wake-up delay, you can refer to Figure 7 The embodiments of the present invention will not be described in detail here.
[0139] When the first system 110 in the electronic device 100 receives the sound information, it can not only perform the first voice recognition on the sound information, but also detect the preset power-on condition, and when the preset power-on condition is detected, it can power on the second system 120. In this case, the first voice recognition process and the second system power-on process do not interfere with each other and are independent of each other, so they can be carried out simultaneously. Therefore, compared with the related art, the voice wake-up process of the electronic device 100 can shorten or eliminate the delay caused by powering on the second system 120, that is, shorten or eliminate Figure 2 The durations T3 to T4 shown above can significantly reduce the voice wake-up delay. After the user issues a voice command containing the wake-up word, the electronic device 100 can quickly start the target application, reducing user waiting time and improving user experience.
[0140] See also Figure 5 The second system 120 may further include an internal memory 122 and an external memory interface 123. The application processor 121, the internal memory 122, and the external memory interface 123 may be integrated on a single chip or device. For example, the application processor 121, the internal memory 122, and the external memory interface 123 may be integrated on a SOC chip.
[0141] In some embodiments, the application processor 121 may include one or more interfaces. For example, the interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a secure digital input and output (SDIO) interface, a serial peripheral interface (SPI) interface, a mobile industry processor interface (MIPI), a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface. The application processor 121 may be coupled to the power management unit 140 via the SDIO interface, and coupled to the display screen 130 via the MIPI interface and the I2C interface.
[0142] In some embodiments, the sensor control center processor 111 may further include one or more interfaces. For example, the interfaces may include an I2C interface, an I2S interface, an SPI interface, an SDIO interface, and / or a MIPI interface. Specifically, the sensor control center processor 111 may be coupled to the application processor 121 via the I2S interface and the SDIO interface, coupled to the display screen 130 via the MIPI interface and the I2C interface, and coupled to the accelerometer and gyroscope unit 112 via the SPI interface.
[0143] The power management unit 140 is connected to the battery of the electronic device 100, the sensor control center processor 111, and the application processor 121. The power management unit 140 receives input from the battery and provides power to the sensor control center processor 111, the application processor 121, the internal memory 122, and the display screen 130.
[0144] It is understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The illustrated components can be implemented in hardware, software, or a combination of software and hardware. For example, the electronic device 100 may also include a universal serial bus (USB) interface, a battery, a mobile communication unit, an audio unit, a speaker, a receiver, a microphone, a button, a camera, and a subscriber identification module (SIM) card interface, a pressure sensor, an air pressure sensor, a magnetic sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.
[0145] The following is a detailed description of the voice wake-up method provided in the embodiments of the present application.
[0146] See also Figure 8 , Figure 8 This is a flow chart of a voice wake-up method provided in an embodiment of the present application. The voice wake-up method is applied to an electronic device, which includes a first system, a second system, and a power management unit. When the electronic device is in the screen-off state, the first system is powered on and the second system is powered off. The voice wake-up method includes steps 101 to 104.
[0147] Step 101: The first system lights up the screen of the electronic device.
[0148] The first system can monitor the posture of the electronic device and perform voice activity detection and voice recognition on the received sound information. For example, the first system can obtain the posture of the electronic device through the accelerometer and gyroscope unit, perform VAD detection on the received sound information, and perform first voice recognition on the sound information through the DSP unit.
[0149] In one application scenario, the first system may light up the screen of the electronic device when detecting that the posture of the electronic device is a preset posture.
[0150] In yet another application scenario, the first system may light up the screen of the electronic device when detecting that the user is looking at the screen of the electronic device.
[0151] In another application scenario, the first system may light up the screen of the electronic device when it detects that the user's body is close to the electronic device.
[0152] For the specific implementation of lighting up the screen of the electronic device, please refer to Figure 5 The relevant contents of the illustrated embodiment will not be repeated here.
[0153] Step 102: If the first system receives sound information, it detects a preset power-on condition and performs a first voice recognition on the sound information.
[0154] The preset power-on condition is a condition that determines that the user is likely to perform voice wake-up. The preset power-on condition may include one or more of the following: the sound information contains a human voice; a preset user behavior occurs; the posture of the electronic device meets a preset posture condition.
[0155] For example, the preset posture condition can be set based on experience to indicate that the user may need to use the electronic device. For example, the preset posture condition can be: the posture of the electronic device is maintained in the preset posture for a time greater than a time threshold.
[0156] For example, the preset user behavior can be set based on experience to indicate that the user may need to use the electronic device. For example, the preset user behavior can include one or more of the following: the user's body is close to the electronic device; the user is looking at the screen of the electronic device.
[0157] For detailed description of each item in the preset power-on conditions, please refer to Figure 5 The relevant contents of the illustrated embodiment will not be repeated here.
[0158] In some embodiments, step 102 may include: if the first system receives sound information, detecting a preset power-on condition, and performing first speech recognition on the sound information if the sound information contains a human voice. The first system may use a relevant algorithm or model to detect information such as frequency, energy, phase, and amplitude of the sound information to determine whether the sound information contains a human voice.
[0159] For example, if the sound information does not contain a human voice, the first system may not perform the first voice recognition on the sound information, but may continue to acquire sound information until the sound information contains a human voice, and then perform the first voice recognition on the sound information. This can reduce the energy consumption of the first system and increase the battery life of the electronic device.
[0160] Step 103: If the first system detects a preset power-on condition, the second system is powered on by the power management unit.
[0161] Step 104: If the second system is powered on and the first voice recognition result meets the first preset condition, the second system starts the target application.
[0162] In some embodiments, the target application may be an application corresponding to the first speech recognition result.
[0163] For example, the first preset condition may be: the first voice recognition result is that the sound information contains a preset wake-up word. In this case, the second system may start the target application corresponding to the first voice recognition result, that is, the second system starts the target application corresponding to the preset wake-up word. For example, if the first voice recognition result is that the sound information contains the preset wake-up word "Hello Xiaoyi", the second system will start the voice assistant application corresponding to the preset wake-up word "Hello Xiaoyi".
[0164] In some embodiments, when the second system needs to perform a second voice recognition on the sound information, the target application may be an application corresponding to the second voice recognition result.
[0165] Among them, if the second system is powered on and the first voice recognition result meets the first preset condition, the second system performs a second voice recognition on the sound information; if the second voice recognition result meets the second preset condition, the second system starts the target application corresponding to the second voice recognition result.
[0166] For example, the second preset condition may be: the second voice recognition result is that the sound information contains a preset wake-up word. In this case, the second system may start the target application corresponding to the second voice recognition result, that is, the second system starts the target application corresponding to the preset wake-up word. For example, if the second voice recognition result is that the sound information contains the preset wake-up word "Hello Xiaoyi", the second system starts the voice assistant application corresponding to the preset wake-up word "Hello Xiaoyi".
[0167] It should be noted that the first preset condition and the second preset condition may be the same or different, and this is not limited in the embodiments of the present application. For example, the first preset condition and the second preset condition may both be that the sound information contains a preset wake-up word, but the similarity between the wake-up word extracted from the sound information and the preset wake-up word may be different.
[0168] Optionally, if the second voice recognition result does not meet the second preset condition, a prompt message is generated to prompt the user to perform voice wake-up again.
[0169] For ease of description, the following refers to the first voice recognition result that meets the first preset condition as first voice recognition success, the first voice recognition result that does not meet the first preset condition as first voice recognition failure, the second voice recognition result that meets the second preset condition as second voice recognition success, and the second voice recognition result that does not meet the second preset condition as second voice recognition failure.
[0170] The prompt message can inform the user that the voice wake-up failed, allowing the user to send a voice message containing the preset wake-up word again. The first system then performs a first voice recognition on the voice message. If the first voice recognition is successful, the second system performs a second voice recognition on the voice message. If the second voice recognition is successful, the second system launches the target application.
[0171] For the delay analysis of the voice wake-up method in this embodiment, please refer to Figure 7 And related content will not be repeated here.
[0172] In the above voice wake-up method, when receiving the sound information, the first system can not only perform the first voice recognition on the sound information, but also detect the preset power-on condition, and when the preset power-on condition is detected, the second system can be powered on. Therefore, compared with the related art, the above voice wake-up method can shorten or eliminate the delay caused by powering on the second system, that is, shorten or eliminate Figure 2 The durations T3 to T4 shown in the figure can significantly reduce the voice wake-up delay. After the user issues a voice command containing the wake-up word, the electronic device can quickly launch the target application, reducing user waiting time and improving the user experience.
[0173] The following describes in detail several possible implementation methods of the above-mentioned preset power-on conditions, taking the target application as a voice assistant application as an example.
[0174] Example 1
[0175] The preset power-on condition is: the sound information contains human voice, and the electronic device is in the preset posture for a time greater than the time threshold. Figure 9 An embodiment corresponding to the preset power-on condition is described.
[0176] Figure 9 This is a flow chart of a voice wake-up method provided by an embodiment of the present application. Figure 9 , the method may include:
[0177] Step 201: Please refer to step 101 and will not be repeated here.
[0178] Step 202: If the first system receives the sound information, it performs voice activity detection on the sound information.
[0179] The first system may perform voice activity detection on the sound information through a voice activity detection unit. For example, the electronic device collects sound information through a digital microphone and sends the sound information to the voice activity detection unit, which performs voice activity detection on the sound.
[0180] Step 203: If the sound information includes a human voice, the first system activates the voice processing unit and detects the time for which the posture of the electronic device is maintained in a preset posture after the screen is turned on.
[0181] For example, if the voice activity detection unit detects that the sound information contains a human voice, it indicates that the user may want to wake up the electronic device through a voice command. In this case, the voice activity detection unit can send a start signal to the voice processing unit of the first system. The start signal is used to start the voice processing unit to prepare for the first voice recognition on the sound information.
[0182] It should be noted that the human voice may be a voice command for starting a voice assistant application, or it may be voice communication between users.
[0183] If the sound information contains a human voice, it indicates that the user may want to wake up the electronic device through voice commands. Furthermore, if the electronic device maintains a preset posture for longer than a time threshold after turning on the screen, it indicates that the user is likely to use an application on the electronic device. Therefore, based on whether the sound information contains a human voice and whether the preset posture is maintained for longer than a time threshold, a comprehensive prediction of the time to power on the second system can increase the accuracy of the prediction.
[0184] Step 204: The first system performs first speech recognition on the sound information through the speech processing unit.
[0185] Exemplarily, the speech processing unit may perform first speech recognition on the sound information through a relevant speech recognition model or speech recognition algorithm. The embodiment of the present application does not limit the specific method and process of the first speech recognition.
[0186] If the first voice recognition is successful, step 205 may be executed after step 204; if the first voice recognition is unsuccessful, step 204 may be executed again.
[0187] Step 205: If the holding time is greater than the time threshold, the first system powers on the second system through the power management unit.
[0188] For example, see Figure 7 , the CPU core or the SOC chip of the second system can be powered on during the time period T1 to T3. For example, if the sound information contains a human voice between T0 and T1, the first system can start detecting whether the preset posture is maintained for longer than a time threshold at time T1. If it is detected that the preset posture is maintained for longer than the time threshold, the first system can power on the CPU core or the SOC chip of the second system through the power management unit.
[0189] Among them, in the case where the holding time of the above-mentioned preset posture is less than or equal to the time threshold, since there is no high probability that the user needs to perform voice wake-up, in order to save energy consumption, the second system can be powered on first, but wait until the above-mentioned first voice recognition is successful, and then the first system powers on the CPU core of the second system or the SOC chip of the second system through the power management unit.
[0190] In some embodiments, after detecting that the sound information contains a human voice, the acceleration-gyroscope unit can be used to detect the holding time of the above-mentioned preset posture. The holding time can specifically be: the time from the time the screen is turned on to the time when the sound information is detected to contain a human voice. If the holding time is greater than the time threshold, the acceleration-gyroscope unit can send a power-on signal to the power management unit. In response to the power-on signal, the power management unit powers on the CPU core of the second system or the SOC chip of the second system. If the holding time is less than or equal to the time threshold, the acceleration-gyroscope unit does not send the power-on signal to the power management unit. Instead, the acceleration-gyroscope unit waits until the above-mentioned first voice recognition is successful, and then sends the power-on signal to the power management unit.
[0191] For details on the time threshold, please refer to the above-mentioned content and will not be elaborated here.
[0192] Step 206: The second system performs a second voice recognition on the sound information.
[0193] Exemplarily, the second system may perform a second voice recognition on the sound information through a relevant voice recognition model or voice recognition algorithm. The embodiment of the present application does not limit the specific method and process of the second voice recognition.
[0194] Generally, compared with the first voice recognition, the second voice recognition voice recognition model or voice recognition algorithm is generally larger and more accurate. For example, if the sound information contains a pseudo wake-up word that is similar to the preset wake-up word, the first voice recognition is likely to recognize that the sound information contains the preset wake-up word, while the second voice recognition can recognize that the sound information does not contain the threshold wake-up word. Therefore, the second voice recognition can more accurately identify whether the sound information contains the wake-up word, thereby reducing the false wake-up rate.
[0195] If the second voice recognition is successful, step 207 is executed. If the second voice recognition is unsuccessful, a prompt message is generated to prompt the user to perform voice wake-up again. After the user again sends a voice message containing the preset wake-up word, the first system performs the first voice recognition on the voice message sent by the user again. If the first voice recognition is successful, the second system 120 performs the second voice recognition on the voice message sent by the user again. If the second voice recognition is successful, the second system 120 launches the voice assistant application.
[0196] Step 207: After the second voice recognition is successful, the second system starts the voice assistant application.
[0197] The following describes the voice wake-up delay of this embodiment and Figure 2 The voice wake-up delay in the embodiment is compared. Among them, the time it takes for the user to speak the voice command containing the wake-up word is about 800ms, the time corresponding to T0~T1 is about 7ms, the time corresponding to T1~T2 is about 793ms, and the time corresponding to T2~T3 is about 100ms. The time required for the CPU core to power on is about 200ms, and the time required for the SOC chip to power on is about 1000ms. The time difference between starting to detect whether the holding time of the preset posture is greater than the time threshold at time T1 and detecting that the holding time is greater than the time threshold can be ignored.
[0198] For the case of powering up the CPU core, Figure 2 In the embodiment, the voice wake-up delay is the duration corresponding to T2 to T6. In this embodiment, the wake-up delay is the duration corresponding to T2 to T3 plus the duration corresponding to T4 to T6. Figure 2 The voice wake-up delay in the embodiment is reduced by the duration corresponding to T3 to T4 (ie, 200ms).
[0199] For the case of powering up the SOC chip, Figure 2 In the embodiment, the voice wake-up delay is the duration corresponding to T2 to T6. In this embodiment, the wake-up delay is the sum of the duration corresponding to T2 to T3, the duration corresponding to T4 to T6, and 107ms (the duration required for the SOC chip to power on, 1000ms, minus the duration corresponding to T1 to T3). Figure 2 The voice wake-up delay in the embodiment is reduced by about 900ms (1000ms-107ms=893ms).
[0200] Example 2
[0201] The preset power-on condition is: the sound information contains human voice and the user's body is close to the electronic device. Figure 10 An embodiment corresponding to the preset power-on condition is described.
[0202] Figure 10 This is a flow chart of another voice wake-up method provided by an embodiment of the present application. Figure 10 , the method may include:
[0203] Steps 301-302 refer to steps 201-202 and will not be described in detail here.
[0204] Step 303: If the sound information includes a human voice, the first system activates a voice signal processing unit of the first system and detects whether the user's body is close to the electronic device.
[0205] Based on the relevant content in step 203, it can be seen that if the sound information contains a human voice, it indicates that the user may want to wake up the electronic device through voice commands. Furthermore, if the user's body is close to the electronic device, it indicates that the user is likely to need to use an application of the electronic device. Therefore, based on whether the sound information contains a human voice and whether the user's body is close to the electronic device, comprehensively predicting the time to power on the second system can increase the accuracy of the prediction.
[0206] In some embodiments, a proximity light sensor or ultrasonic device can be used to detect whether a user's limb is near the electronic device. For example, a smartwatch can use a proximity light sensor or ultrasonic device to detect whether a user's arm, which is not wearing the smartwatch, is near the electronic device.
[0207] Step 304: Refer to step 204 and will not be repeated here.
[0208] Step 305: If it is detected that the user's body is close to the electronic device, the first system powers on the second system through the power management unit.
[0209] For example, see Figure 7 , the CPU core or the SOC chip of the second system can be powered on during the time range T1 to T3. For example, if the sound information contains a human voice during the time range T0 to T1, detection of whether a user's body is near the electronic device begins at T1. If the user's body is detected near the electronic device, the first system can use the power management unit to power on the CPU core or the SOC chip of the second system.
[0210] Among them, in the case where the user's body is not detected close to the electronic device, since there is no high probability that the user needs to perform voice wake-up, in order to save energy consumption, the second system can be powered on first. Instead, the first system can wait until the above-mentioned first voice recognition is successful, and then power on the CPU core of the second system or the SOC chip of the second system through the power management unit.
[0211] In one scenario, if the proximity light sensor or ultrasonic device detects a user's body approaching the electronic device, it can send a power-on signal to the power management unit. In response to the power-on signal, the power management unit powers on the CPU core of the second system or the SOC chip of the second system.
[0212] Steps 306 to 307 refer to steps 206 to 207 and will not be described in detail here.
[0213] Compared to the voice wake-up process in Example 1, this embodiment does not detect whether the preset posture is held for longer than a time threshold. Instead, it detects whether the user's body is close to the electronic device. The time taken for these two detections is similar, so the voice wake-up delay in this embodiment is basically the same as that in Example 1 and will not be further described here.
[0214] Example 3
[0215] The preset power-on condition is: the sound information contains human voice and the user is looking at the screen of the electronic device. Figure 11 An embodiment corresponding to the preset power-on condition is described.
[0216] Figure 11 This is a flow chart of another voice wake-up method provided by an embodiment of the present application. Figure 11 , the method comprising:
[0217] Steps 401-402 refer to steps 201-202 and will not be repeated here.
[0218] Step 403: If the sound information includes a human voice, the first system activates a voice signal processing unit of the first system and detects whether the user is looking at the screen of the electronic device.
[0219] Based on the relevant content in step 203, it can be seen that if the sound information contains a human voice, it indicates that the user may want to wake up the electronic device through voice commands. Furthermore, after lighting up the screen, the user looks at the screen of the electronic device, which indicates that the user is likely to use an application of the electronic device. Therefore, based on whether the sound information contains a human voice and whether the user is looking at the screen of the electronic device, the time to power on the second system can be comprehensively predicted to increase the accuracy of the prediction.
[0220] In some embodiments, a camera may be used to detect whether a user is looking at the screen of the electronic device. For example, a smartwatch may be used to detect whether a user is looking at the screen of the smartwatch.
[0221] Step 404: Refer to step 204 and will not be repeated here.
[0222] Step 405: If it is detected that the user is looking at the screen of the electronic device, the first system powers on the second system via the power management unit.
[0223] For example, see Figure 7, the CPU core or the SOC chip of the second system can be powered on during the time range T1 to T3. For example, if the sound information contains a human voice during the time range T0 to T1, the first system can detect whether the user is looking at the screen of the electronic device starting at T1. If the user is detected looking at the screen of the electronic device, the first system can power on the CPU core or the SOC chip of the second system through the power management unit.
[0224] Among them, in the case where the user is not detected looking at the screen of the electronic device, since there is no high probability that the user needs to perform voice wake-up, in order to save energy consumption, the second system may not be powered on first. Instead, the first system may wait until the above-mentioned first voice recognition is successful, and then power on the CPU core of the second system or the SOC chip of the second system through the power management unit.
[0225] In one scenario, if the camera detects that the user is looking at the screen of the electronic device, it can send a power-on signal to the power management unit. In response to the power-on signal, the power management unit powers on the CPU core of the second system or the SOC chip of the second system.
[0226] Steps 406 to 407 refer to steps 206 to 207 and will not be described in detail here.
[0227] Compared to the voice wake-up process in Example 1, this embodiment does not detect whether the preset posture is held for longer than a time threshold. Instead, it detects whether the user is looking at the electronic device's screen. The time taken for these two detections is similar, so the voice wake-up delay in this embodiment is basically the same as that in Example 1 and will not be further described here.
[0228] Example 4
[0229] The default power-on condition is: the sound information contains human voice. Figure 12 An embodiment corresponding to the preset power-on condition is described.
[0230] Figure 12 This is a flow chart of another voice wake-up method provided by an embodiment of the present application. Figure 12 , the method comprising:
[0231] Steps 501-502 refer to steps 201-202 and will not be repeated here.
[0232] Step 503: If the sound information includes a human voice, the first system starts the voice processing unit of the first system and powers on the second system through the power management unit.
[0233] For example, if the voice activity detection unit detects a human voice in the sound information, it indicates that the user may want to wake up the electronic device through a voice command. In this case, the voice activity detection unit can send a start signal to the voice processing unit of the first system. In response to the start signal, the voice processing unit prepares to perform a first voice recognition on the sound information. Furthermore, if the sound information contains a human voice, the first system can power on the second system through the power management unit.
[0234] Step 504: See step 204, which will not be repeated here.
[0235] Steps 505-506 refer to steps 206-207 and will not be described in detail here.
[0236] Compared to the voice wake-up process in Example 1, this embodiment does not detect whether the preset posture is held for longer than the time threshold. The time spent on detecting whether the preset posture is held for longer than the time threshold is almost negligible, so the voice wake-up delay in this embodiment is basically the same as the voice wake-up delay in Example 1 and will not be repeated here.
[0237] Example 5
[0238] The preset power-on condition is: the electronic device is in the preset posture for a time greater than the time threshold. Figure 13 An embodiment corresponding to the preset power-on condition is described.
[0239] Figure 13 This is a flow chart of another voice wake-up method provided by an embodiment of the present application. Figure 13 , the method comprising:
[0240] Step 601: Please refer to step 201 and will not be repeated here.
[0241] Step 602: The first system starts a voice processing unit and detects the time during which the posture of the electronic device is maintained in a preset posture.
[0242] Compared with Example 2, this embodiment can start the voice processing unit of the first system after the screen of the electronic device is turned on, and detect the time for which the posture of the electronic device is maintained in the preset posture after the screen is turned on.
[0243] Step 603: Please refer to step 204 and will not be repeated here.
[0244] It should be noted that the voice processing unit will perform the first voice recognition on the voice information only when the first system obtains the voice information. If the first system does not obtain the voice information, the voice processing unit may be in a waiting state.
[0245] Step 604: If the holding time is greater than the time threshold, the first system powers on the second system through the power management unit.
[0246] Among them, if the holding time is greater than the time threshold, it means that the user may want to start the voice assistant of the electronic device through voice commands. At this time, when the above holding time is greater than the time threshold, the first system can power on the CPU core of the second system or power on the SOC chip of the second system in advance through the power management unit. Among them, if the holding time is less than or equal to the time threshold, then after the first voice recognition is successful, the first system powers on the CPU core of the second system or powers on the SOC chip of the second system through the power management unit.
[0247] For steps 605 to 606, please refer to steps 206 to 207, which will not be described in detail here.
[0248] Compared to the voice wake-up process in Example 1, the first system in this embodiment does not need to detect whether the sound information contains a human voice. The time spent by the first system to detect whether the sound information contains a human voice is approximately 7ms, which is negligible. Therefore, the voice wake-up delay in this embodiment is basically the same as the voice wake-up delay in Example 1 and will not be further described here.
[0249] Example 6
[0250] The default power-on condition is: the user's body is close to the electronic device. Figure 14 An embodiment corresponding to the preset power-on condition is described.
[0251] Figure 14 This is a flow chart of another voice wake-up method provided by an embodiment of the present application. Figure 14 , the method comprising:
[0252] Step 701: Please refer to step 201 and will not be repeated here.
[0253] Step 702: The first system starts a voice processing unit and detects whether the user's body is close to the electronic device.
[0254] In some embodiments, a proximity light sensor or ultrasonic device can be used to detect whether a user's limb is close to the electronic device. For example, a smartwatch can use a proximity light sensor or ultrasonic device to detect whether a user's arm, which is not wearing the smartwatch, is close to the electronic device.
[0255] Step 703: Please refer to step 204 and will not be repeated here.
[0256] Step 704: If it is detected that the user's body is close to the electronic device, the first system powers on the second system through the power management unit.
[0257] The user's physical proximity to the electronic device indicates that the user may want to activate the electronic device's voice assistant through a voice command. In this case, upon detecting that the user's physical proximity to the electronic device, the first system can use the power management unit to power on the CPU core of the second system or pre-power on the SOC chip of the second system.
[0258] For steps 705 to 706, please refer to steps 206 to 207 and will not be repeated here.
[0259] Compared to the voice wake-up process in Example 1, the first system in this embodiment detects whether the user's limbs are close to the electronic device, rather than detecting whether the holding time is greater than the time threshold. The time spent on the two detections is almost the same and can be ignored. In addition, in this embodiment, the first system does not need to detect whether the sound information contains human voices, and the time spent by the first system to detect whether the sound information contains human voices is about 7ms. Therefore, the voice wake-up delay in this embodiment is basically the same as the voice wake-up delay in Example 1, and will not be repeated here.
[0260] Example 7
[0261] The preset power-on condition is: the user looks at the screen of the electronic device. Figure 15 An embodiment corresponding to the preset power-on condition is described.
[0262] Figure 15 This is a flow chart of another voice wake-up method provided by an embodiment of the present application. Figure 15 , the method comprising:
[0263] Step 801: Please refer to step 201 and will not be repeated here.
[0264] Step 802: The first system starts a voice processing unit and detects whether the user is looking at the screen of the electronic device.
[0265] In some embodiments, a camera may be used to detect whether a user is looking at the screen of an electronic device. For example, a smartwatch may be used to detect whether a user is looking at the screen of the smartwatch.
[0266] Step 803: Please refer to step 204 and will not be repeated here.
[0267] Step 804: If it is detected that the user is looking at the screen of the electronic device, the first system powers on the second system via the power management unit.
[0268] The user looking at the screen of the electronic device indicates that the user may want to activate the voice assistant of the electronic device through a voice command. In this case, upon detecting that the user is looking at the screen of the electronic device, the first system can use the power management unit to power on the CPU core of the second system or pre-power on the SOC chip of the second system.
[0269] For steps 805 to 806, please refer to steps 206 to 207 and will not be repeated here.
[0270] Compared with the voice wake-up process in Example 1, the first system in this embodiment detects whether the user is looking at the screen of the electronic device, rather than detecting whether the holding time is greater than the time threshold. The time spent on the two detections is almost the same and can be ignored. In addition, in this embodiment, the first system does not need to detect whether the sound information contains human voices. The time spent by the first system to detect whether the sound information contains human voices is about 7ms. Therefore, the voice wake-up delay of this embodiment is basically the same as the voice wake-up delay of Example 1, and will not be repeated here.
[0271] In the above embodiments 1 to 7, after the second system is powered on again, a second voice recognition operation is required for the sound information. However, the embodiments of the present application are not limited to this. For example, after the second system is powered on again, the voice assistant application corresponding to the first voice recognition result can be directly launched. The following describes several embodiments in detail.
[0272] Example 8
[0273] The preset power-on condition of this embodiment is the same as the preset power-on condition of embodiment 1.
[0274] The process of this embodiment differs from the process of embodiment 1 in that: after the second system is powered on again and the first voice recognition is successful, the second system does not need to perform a second voice recognition on the sound information, and directly starts the voice assistant application corresponding to the first voice recognition result.
[0275] Example 9
[0276] The preset power-on condition of this embodiment is the same as the preset power-on condition of embodiment 2.
[0277] The process of this embodiment differs from the process of embodiment 2 in that: after the second system is powered on again and the first voice recognition is successful, the second system does not need to perform a second voice recognition on the sound information, and directly starts the voice assistant application corresponding to the first voice recognition result.
[0278] Example 10
[0279] The preset power-on condition of this embodiment is the same as the preset power-on condition of embodiment 3.
[0280] The process of this embodiment differs from the process of embodiment 3 in that: after the second system is powered on again and the first voice recognition is successful, the second system does not need to perform a second voice recognition on the sound information, and directly starts the voice assistant application corresponding to the first voice recognition result.
[0281] Example 11
[0282] The preset power-on condition of this embodiment is the same as the preset power-on condition of embodiment 4.
[0283] The process of this embodiment differs from the process of embodiment 4 in that: after the second system is powered on again and the first voice recognition is successful, the second system does not need to perform a second voice recognition on the sound information, and directly starts the voice assistant application corresponding to the first voice recognition result.
[0284] Example 12
[0285] The preset power-on condition of this embodiment is the same as the preset power-on condition of embodiment 5.
[0286] The process of this embodiment differs from the process of embodiment 5 in that: after the second system is powered on again and the first voice recognition is successful, the second system does not need to perform a second voice recognition on the sound information, and directly starts the voice assistant application corresponding to the first voice recognition result.
[0287] Example 13
[0288] The preset power-on condition of this embodiment is the same as the preset power-on condition of embodiment 6.
[0289] The process of this embodiment differs from the process of Example 6 in that: after the second system is powered on again and the first voice recognition is successful, the second system does not need to perform a second voice recognition on the sound information, and directly starts the voice assistant application corresponding to the first voice recognition result.
[0290] Example 14
[0291] The preset power-on condition of this embodiment is the same as the preset power-on condition of embodiment 7.
[0292] The process of this embodiment differs from the process of embodiment 7 in that: after the second system is powered on again and the first voice recognition is successful, the second system does not need to perform a second voice recognition on the sound information, and directly starts the voice assistant application corresponding to the first voice recognition result.
[0293] Corresponding to the voice wake-up method in the above embodiment, Figure 16 The structure block diagram of the voice awakening device 900 provided in the embodiment of the present application is shown. For the sake of convenience, only the parts related to the embodiment of the present application are shown. The above-mentioned voice awakening device 900 is applied to Figure 5The electronic device includes a first system, a second system and a power management unit. When the electronic device is in a screen-off state, the first system is in a power-on state and the second system is in a power-off state.
[0294] See also Figure 16 The voice wake-up device 900 in the embodiment of the present application may include a screen lighting unit 901, a detection and identification unit 902, a power-on unit 903 and an application starting unit 904.
[0295] Among them, the screen lighting unit 901 is used to light up the screen of the electronic device. The detection and recognition unit 902 is used to detect the preset power-on condition when the first system receives sound information and perform the first voice recognition on the sound information. The power-on unit 903 is used to power on the second system through the power management unit when the first system detects the preset power-on condition. The application startup unit 904 is used to start the target application when the second system is powered on and the first voice recognition result meets the first preset condition.
[0296] In the above-mentioned voice wake-up device, the first system can simultaneously detect the preset power-on condition and perform the first voice recognition on the sound information, and when the preset power-on condition is detected, the second system can be powered on. Therefore, compared with the related art, the above-mentioned voice wake-up device can shorten or eliminate the time corresponding to the above-mentioned T3-T4, greatly shortening the voice wake-up delay. After the user issues a voice command containing the wake-up word, the above-mentioned voice wake-up device can quickly start the target application, reducing the user's waiting time and improving the user experience.
[0297] The present application also provides an electronic device. Figure 17 The electronic device 1000 may include: at least one processor 1010, a memory 1020, and a computer program stored in the memory 1020 and executable on the at least one processor 1010. When the processor 1010 executes the computer program, the steps in any of the above method embodiments are implemented, for example Figure 8 Steps 101 to 104 in the embodiment shown. Alternatively, when the processor 1010 executes the computer program, the functions of each unit in the above-mentioned device embodiments are realized, for example Figure 16 Functions of units 901 to 904 are shown.
[0298] For example, the computer program may be divided into one or more modules / units, one or more of which are stored in the memory 1020 and executed by the processor 1010 to complete the present application. One or more modules / units may be a series of computer program segments capable of completing specific functions, and the program segments are used to describe the execution process of the computer program in the electronic device 1000.
[0299] Those skilled in the art will understand that Figure 17 These are merely examples of electronic devices and do not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or a combination of certain components, or different components, such as input and output devices, network access devices, buses, etc.
[0300] The processor 1010 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0301] The memory 1020 can be an internal storage unit of the electronic device or an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. The memory 1020 is used to store computer programs and other programs and data required by the electronic device. The memory 1020 can also be used to temporarily store data that has been output or is about to be output.
[0302] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified into address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.
[0303] An embodiment of the present application also provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is executed on a computer or a processor, the computer or processor executes one or more steps in any of the above methods.
[0304] An embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer or a processor, enables the computer or processor to execute one or more steps in any of the above methods.
[0305] The present application also provides a chip system, which may include a memory and a processor, wherein the processor executes a computer program stored in the memory to implement one or more steps in any of the above methods. The chip system may be a single chip or a chip module composed of multiple chips.
[0306] The present application also provides a chip system, which may include a processor coupled to a memory, and the processor executes a computer program stored in the memory to implement one or more steps of any of the above methods. The chip system may be a single chip or a chip module composed of multiple chips.
[0307] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0308] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0309] Finally, it should be noted that the above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. An electronic device, characterized in that: The electronic device includes a first system, a second system, and a power management unit. When the electronic device is in a screen-off state, the first system is in a powered-on state, and the second system is in a powered-off state. The first system is used to light up the screen of the electronic device; The first system is further configured to detect a preset power-on condition upon receiving sound information and perform a first voice recognition on the sound information; The first system is further configured to power on the second system via the power management unit when the preset power-on condition is detected; The second system is configured to start a target application when power-on is completed and the first speech recognition result meets a first preset condition; The preset power-on condition includes: the posture of the electronic device satisfies a preset posture condition, the preset posture condition is: the posture of the electronic device is maintained in the preset posture for a time greater than a time threshold, the time threshold is determined based on historical time information of the electronic device maintaining the preset posture, the historical time information includes the time when the screen of the electronic device is turned on and the time when the voice information is received; The first system includes an acceleration gyro unit, and the acceleration gyro unit is used to detect the posture of the electronic device.
2. The electronic device according to claim 1, wherein The preset power-on condition may also include one or more of the following: The sound information includes human voice; There are preset user behaviors.
3. The electronic device according to claim 2, wherein: The preset user behavior includes one or more of the following: The user's body is close to the electronic device; The user looks at the screen of the electronic device.
4. The electronic device according to any one of claims 1 to 3, characterized in that: The second system is specifically used for: When power-on is completed and the first voice recognition result meets the first preset condition, performing a second voice recognition on the sound information; When the second voice recognition result meets a second preset condition, a target application corresponding to the second voice recognition result is started.
5. The electronic device according to claim 4, characterized in that The second system is also used for: If the second voice recognition result does not meet the second preset condition, a prompt message is generated to prompt the user to perform voice wake-up again.
6. A voice wake-up method, characterized in that: Applied to an electronic device, the electronic device includes a first system, a second system, and a power management unit. When the electronic device is in a screen-off state, the first system is in a powered-on state and the second system is in a powered-off state. The method includes: The first system lights up the screen of the electronic device; If the first system receives the sound information, detecting a preset power-on condition and performing a first voice recognition on the sound information; If the first system detects the preset power-on condition, powering on the second system through the power management unit; If the second system is powered on and the first voice recognition result meets the first preset condition, the second system starts the target application; The preset power-on condition includes: the posture of the electronic device satisfies a preset posture condition, the preset posture condition is: the posture of the electronic device is maintained in the preset posture for a time greater than a time threshold, the time threshold is determined based on historical time information of the electronic device maintaining the preset posture, the historical time information includes the time when the screen of the electronic device is turned on and the time when the voice information is received; The first system includes an acceleration gyro unit, and the acceleration gyro unit is used to detect the posture of the electronic device.
7. The method according to claim 6, characterized in that The preset power-on condition may also include one or more of the following: The sound information includes human voice; There are preset user behaviors.
8. The method according to claim 7, characterized in that The preset user behavior includes one or more of the following: The user's body is close to the electronic device; The user looks at the screen of the electronic device.
9. The method according to any one of claims 6 to 8, characterized in that If the second system is powered on and the first voice recognition result meets the first preset condition, the second system starts the target application, including: If the second system is powered on and the first voice recognition result meets the first preset condition, the second system performs a second voice recognition on the sound information; If the second voice recognition result meets a second preset condition, the second system starts a target application corresponding to the second voice recognition result.
10. The method according to claim 9, characterized in that The method further comprises: If the second voice recognition result does not meet the second preset condition, the second system generates a prompt message for prompting the user to perform voice wake-up again.
11. An electronic device, characterized in that: include: one or more processors, memory, and display screens; The memory and the display screen are coupled to the one or more processors, the memory being configured to store computer program codes, the computer program codes comprising computer instructions; When the one or more processors execute the computer instructions, the electronic device is caused to perform the method according to any one of claims 6 to 10.
12. A chip system, characterized in that: The chip system includes a processor coupled to a memory, and the processor executes a computer program stored in the memory to implement the method according to any one of claims 6 to 10.
Citation Information
Patent Citations
Switching control method and device for voice wake-up recognition
CN110223691A
Electronic APPARATUS AND METHOD FOR controlling the same
CN111402871A