Voice wake-up method and related device
The user gestures and voice activities are detected by combining sensors and microphones, and the voice interaction without wake-up words is realized, the cumbersome registration and high power consumption problems of fixed wake-up words are solved, and the efficiency and low power consumption of electronic devices are improved.
Patent Information
- Application Number
- CN202411978739.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-08-05
AI Technical Summary
The existing fixed wake-up word voice interaction method requires cumbersome registration steps and repeated wake-up word operation in a quiet environment, which affects the efficiency and popularity of smart voice interaction, and increases the power consumption of electronic devices and dependence on low-power space.
The sensor detects user gestures and collects sound data in combination with the microphone to determine whether the specified state of wake-up-free voice interaction is met, including posture and environmental conditions, and serially detects voice activity to activate voice assistants to reduce dependence on low-power space.
The voice interaction without wake-up words is realized, which reduces the power consumption of electronic devices, simplifies the registration process, improves the efficiency and response speed of voice interaction, and reduces the dependence on low-power space.
Smart Images

Figure CN120428848A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminals, and in particular to a voice wake-up method and related devices. Background Art
[0002] With the development of the terminal industry, electronic devices (such as smartphones, tablets, and wearable smart devices) are becoming increasingly popular in daily life. Intelligent voice interaction has become a common feature in daily life. Currently, a common intelligent voice interaction method is voice wake-up based on fixed wake-up words.
[0003] Using a fixed wake-up word to wake up the voice assistant is convenient. However, this intelligent voice interaction method requires registering a fixed wake-up word, which is a very cumbersome registration step. For example, you need to find a multi-level menu entry to register a fixed wake-up word. In addition, when the user interacts with the electronic device by voice, the user needs to repeat the wake-up word multiple times in a quiet environment and at a certain distance. This not only affects the promotion and popularization of the intelligent voice interaction function, but also affects the efficiency and response speed of the voice interaction between the electronic device and the user. Summary of the Invention
[0004] The present application provides a voice wake-up method and related devices, which reduce the power consumption of electronic devices. At the same time, since the implementation process of detecting user gestures and detecting voice activities in a low-power space is relatively simple, it can also reduce the dependence of electronic devices on low-power spaces.
[0005] In a first aspect, the present application provides a voice wake-up method, comprising: an electronic device collects first sound data through a microphone. When a first gesture is detected, the electronic device detects whether the first sound data has voice activity. If the first sound data has voice activity, the electronic device determines whether the state of the electronic device is a first state. The first state is a state corresponding to the electronic device performing voice interaction without a wake-up word. If the state of the electronic device is the first state, the electronic device starts a voice assistant. In this way, the power consumption of the electronic device can be reduced, and the dependence of the electronic device on low-power space can also be reduced.
[0006] In one possible implementation, the method further includes: the electronic device collects first sensor data. If the first sound data has voice activity, the electronic device determines whether the state of the electronic device is the first state, specifically including: if the first sound data has voice activity, the electronic device determines whether the state of the electronic device is the first state based on the first sound data and / or the first sensor data. In this way, it is possible to accurately determine whether the user has a need for wake-up word-free voice interaction.
[0007] In a possible implementation, the first state includes one or more of the following: the electronic device collects a human voice within a preset range from the microphone, the electronic device is in a preset posture, and the electronic device is in a preset environment.
[0008] In one possible implementation, if the first sound data has voice activity, the electronic device determines whether the state of the electronic device is the first state based on the first sound data and / or the first sensor data, specifically including: if the first sound data has voice activity, the electronic device uses the breath to wake up the first-level algorithm model, and determines whether the state of the electronic device is the first state based on the first sound data and / or the first sensor data. If the breath wakes up the first-level algorithm model and determines that the state of the electronic device is the first state, the electronic device uses the breath to wake up the second-level algorithm model and determines whether the state of the electronic device is the first state based on the first sound data and / or the first sensor data. If the breath wakes up the second-level algorithm model and determines that the state of the electronic device is the first state, the electronic device determines that the state of the electronic device is the first state. In this way, it is possible to accurately determine whether the user has the need for voice interaction without the wake-up word, and avoid the situation where the voice assistant is accidentally started.
[0009] In one possible implementation, the method further includes: the electronic device responding to the voice instruction in the first sound data through the voice assistant.
[0010] In one possible implementation, after the electronic device collects the first sound data through the microphone, the method further includes: the electronic device copies the first sound data and stores the copied first sound data in an audio circular buffer. The electronic device responds to the voice command in the first sound data through the voice assistant, specifically including: the electronic device obtains the copied first sound data in the audio circular buffer through the voice assistant, and recognizes the voice command in the copied first sound data. The electronic device executes the operation corresponding to the voice command through the voice assistant. In this way, the problem that the voice assistant cannot respond to the voice command because it cannot obtain the complete sound data can be avoided.
[0011] In one possible implementation, the primary and secondary breath wake-up algorithms are located in an application processor (AP). Alternatively, the primary breath wake-up algorithm is located in an audio digital signal processor (ADSP), while the secondary breath wake-up algorithm is located in the AP. This reduces the power consumption of the ADSP.
[0012] In a second aspect, the present application provides an electronic device, which includes: one or more processors, one or more sensors, and a memory. The memory is coupled to the one or more processors, which are coupled to the one or more sensors, and the memory is used to store computer program code, which includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute a method as described in any possible implementation of the first aspect. In this way, the power consumption of the electronic device is reduced. At the same time, since the implementation process of detecting user gestures and detecting voice activities in the low-power space is relatively simple, the dependence of the electronic device on the low-power space can also be reduced.
[0013] In a third aspect, the present application provides a chip system, which is applied to an electronic device, and the chip system includes one or more processors, which are used to call computer instructions to cause the electronic device to execute a method as described in any possible implementation of the first aspect. In this way, the power consumption of the electronic device is reduced. At the same time, since the implementation process of detecting user gestures and detecting voice activities in a low-power space is relatively simple, the electronic device's dependence on the low-power space can also be reduced.
[0014] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions that, when executed on an electronic device, cause the electronic device to perform a method as described in any possible implementation of the first aspect. This reduces the power consumption of the electronic device and, because detecting user gestures and voice activity in a low-power space is relatively simple, reduces the electronic device's reliance on the low-power space.
[0015] In a fifth aspect, the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the electronic device performs the method as described in any possible implementation of the first aspect. In this way, the power consumption of the electronic device is reduced. At the same time, since the implementation process of detecting user gestures and detecting voice activities in a low-power space is relatively simple, the electronic device's dependence on the low-power space can also be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0017] Figures 2A-2D A schematic diagram of a user interface for a set of wake-up word-free settings provided in an embodiment of the present application;
[0018] Figure 3A-Figure 3B A schematic diagram of an implementation method for waking up without a wake-up word provided in this application;
[0019] Figure 4 A flowchart illustrating a specific implementation of a voice wake-up method provided in an embodiment of the present application;
[0020] Figure 5 A flowchart of a method for processing sound data 1 provided in an embodiment of the present application;
[0021] Figure 6A A schematic diagram of a device architecture related to the voice wake-up method provided in an embodiment of the present application;
[0022] Figure 6B A schematic diagram of another device architecture related to the voice wake-up method provided in an embodiment of the present application;
[0023] Figure 7 A schematic diagram of the execution logic of a voice wake-up method provided in an embodiment of the present application;
[0024] Figure 8 A schematic diagram of the software architecture related to a voice wake-up method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following is a clear and detailed description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the text is only a description of the association relationship between related objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.
[0026] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0027] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.
[0028] In the embodiment of this application, Figure 1 The electronic device 100 shown is the electronic device in the embodiment of the present application.
[0029] The electronic device 100 can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, a smart home device and / or a smart city device. The embodiments of the present application do not impose any special restrictions on the specific type of the electronic device 100.
[0030] like Figure 1 As shown, the electronic device 100 may include a processor 101 , a memory 102 , a wireless communication module 103 (optional), a display screen 104 , a sensor module 105 , an audio module 106 , a speaker 107 and a microphone 108 .
[0031] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0032] The processor 101 may include one or more processor units. For example, the processor 101 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.
[0033] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.
[0034] Processor 101 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 101 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 101. If processor 101 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 101's latency, and thus improves system efficiency.
[0035] In some embodiments, the processor 101 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a USB interface.
[0036] The memory 102 is coupled to the processor 101 and is used to store various software programs and / or multiple sets of instructions. In a specific implementation, the memory 102 may include a volatile memory (volatile memory), such as a random access memory (RAM); it may also include a non-volatile memory (non-volatile memory), such as a ROM, a flash memory (flash memory), a hard disk drive (HDD) or a solid state drive (SSD); the memory 102 may also include a combination of the above types of memory. The memory 102 may also store some program codes so that the processor 101 can call the program code stored in the memory 102 to implement the implementation method of the embodiment of the present application in the electronic device 100. The memory 102 can store an operating system, such as an embedded operating system such as uCOS, VxWorks, RTLinux, etc.
[0037] The wireless communication module 103 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) and the like applied to the electronic device 100. The wireless communication module 103 can be one or more devices integrating at least one communication processing module. The wireless communication module 103 receives electromagnetic waves via an antenna, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 101. The wireless communication module 103 can also receive signals to be sent from the processor 101, frequency modulate and amplify them, and convert them into electromagnetic waves for radiation through the antenna. In some embodiments, the electronic device 100 can also transmit signals to the processor 101 through the Bluetooth module ( Figure 1 Not shown), WLAN module ( Figure 1 The device (not shown) transmits signals to detect or scan devices near the electronic device 100 and establishes a wireless communication connection with the nearby devices to transmit data. The Bluetooth module can provide solutions including one or more of classic Bluetooth (basic rate / enhanced data rate, BR / EDR) or Bluetooth low energy (Bluetooth low energy, BLE), and the WLAN module can provide solutions including one or more of Wi-Fi direct, Wi-Fi LAN, or Wi-Fi softAP.
[0038] The display screen 104 can be used to display images, videos, etc. The display screen 104 may include a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 104, where N is a positive integer greater than one.
[0039] The sensor module 105 may include a touch sensor 105A and an inertial sensor 105B. The touch sensor 105A may also be referred to as a "touch control device." The touch sensor 105A may be disposed on the display screen 104, and the touch sensor 105A and the display screen 104 form a touch screen, also known as a "touch screen." The touch sensor 105A may be used to detect touch operations applied thereto or in the vicinity thereof. The inertial sensor 105B may include an accelerometer and / or a gyroscope sensor, and / or other components for measuring the motion of the electronic device 100. The gyroscope sensor may also be used to determine the motion posture of the electronic device 100. In some embodiments, the electronic device 100 may use the gyroscope sensor to determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The accelerometer may be used to detect the magnitude of the acceleration of the electronic device 100 in various directions (generally the x, y, and z axes). When the electronic device 100 is stationary, the accelerometer may detect the magnitude and direction of gravity.
[0040] The audio module 106 can be used to convert digital audio information into an analog audio signal output, and can also be used to convert analog audio input into a digital audio signal. The audio module 106 can also be used to encode and decode audio signals. In some embodiments, the audio module 106 can also be provided in the processor 101, or some functional modules of the audio module 106 can be provided in the processor 101.
[0041] The speaker 107 , which may also be referred to as a “speaker,” is configured to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or a hands-free phone call through the speaker 107 .
[0042] The microphone 108, which can also be called a "microphone" or "microphone", can be used to collect sound signals in the surrounding environment of the electronic device, convert the sound signals into electrical signals, and then process the electrical signals through a series of processes, such as analog-to-digital conversion, etc., to obtain an audio signal in digital form that can be processed by the processor 101 of the electronic device. When making a call or sending a voice message, the user can speak close to the microphone 108 with their mouth to input the sound signal into the microphone 108. The electronic device 100 can be provided with at least one microphone 108. In some other embodiments, the electronic device 100 can be provided with two microphones 108, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three or more microphones 108 to collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0043] It should be noted that Figure 1 The electronic device 100 shown is only used to exemplarily explain the hardware structure of the electronic device provided in this application and does not constitute a specific limitation to this application.
[0044] Generally, the user can use a fixed wake-up word to wake up the voice assistant in the electronic device 100. For example, the user can say the fixed wake-up word "Hello YOYO" to the microphone on the electronic device 100 to wake up the voice assistant. Then, the voice assistant can respond to the user's voice command and perform the corresponding operation.
[0045] However, adopting this intelligent voice interaction method based on a fixed wake-up word requires registering the fixed wake-up word, and this registration step is very cumbersome. For example, it is necessary to find the multi-level menu entry for registering the fixed wake-up word. Moreover, when the user interacts with the electronic device by voice, the user needs to repeat the wake-up word many times in a quiet environment and at a certain distance, which not only affects the popularization and application of the intelligent voice interaction function but also affects the efficiency and response speed of the voice interaction between the electronic device and the user. Therefore, the electronic device 100 can adopt a wake-up word-free method to wake up the voice assistant, that is, the user does not need to say the fixed wake-up word and can also trigger the voice assistant to respond to the user's voice command.
[0046] Figures 2A-2D It is a schematic diagram of the user interface for a set of wake-up word-free settings provided in the embodiments of this application.
[0047] As Figure 2AAs shown, the electronic device 100 can display a desktop 210. The desktop 210 can display a page with application icons, which includes multiple application icons (for example, an icon for a weather application, an icon for a stock application, an icon for a calculator application, an icon 211 for a settings application, an icon for an email application, an icon for a theme application, an icon for a calendar application, an icon for a video application, etc.). A status bar is displayed in the upper part of the desktop 210. The status bar may include one or more indicators, wherein the one or more indicators may include one or more signal strength indicators of a mobile communication signal (also referred to as a cellular signal), a battery status indicator, a time indicator, a Wi-Fi signal indicator, and the like.
[0048] The electronic device 100 may receive a touch operation (eg, click) on the setting application icon 211. In response to the touch operation, the electronic device 100 may display the setting interface 220.
[0049] like Figure 2B As shown, the setting interface 220 may include one or more page options, such as application page options, battery page options, storage page options, security page options, privacy page options, healthy mobile phone usage page options, voice interaction service page options 221, accessibility page options, etc.
[0050] The electronic device 100 may receive a touch operation (eg, click) on the voice interaction service page option 221. In response to the touch operation, the electronic device 100 may display the voice interaction service page 230.
[0051] like Figure 2C As shown, the voice interaction service page 230 may include one or more wake-up setting options for voice interaction services, volume setting options and timbre setting options for voice interaction services, etc. The wake-up setting options for the above-mentioned one or more voice interaction services may include voice wake-up options, button wake-up options, headphone wake-up options and breath wake-up options 231, etc. Among them, in one possible implementation, the breath wake-up option 231 can be used to trigger the electronic device 100 to start the voice assistant in a breath wake-up manner. The breath wake-up manner is a breath triggering manner that only requires the electronic device 100 to be close to the user's mouth and speak to trigger the voice assistant to start and perform voice interaction. It can be understood that the breath wake-up manner is an exemplary wake-up word-free voice interaction manner.
[0052] When the user wants to start the breath wake-up mode, the electronic device 100 may receive a touch operation (eg, click) on the breath wake-up option 231. In response to the touch operation, the electronic device 100 may display the breath wake-up interface 240.
[0053] like Figure 2D As shown, the breath wake-up interface 240 may include description content of the breath wake-up function, a switch control 241, etc. Among them, the description content of the breath wake-up function is used to introduce the breath wake-up function, so that users can better understand the function of the function. The description content may be, for example, "lift up the phone, bring the bottom of the phone close to your mouth (within 5 cm), aim at the bottom microphone, and start your conversation journey." In some embodiments, the area below the breath wake-up function introduction may display the switch control 241 corresponding to the breath wake-up function. When the user clicks the switch control 241, the breath wake-up function of the breath-triggered voice interaction service can be activated.
[0054] Figure 3A-Figure 3B This application provides an implementation method for waking up without a wake-up word.
[0055] like Figure 3A As shown, in this wake-up method without a wake-up word, the user can lift the electronic device 100 and move the microphone of the electronic device 100 near the mouth, for example, about 7 cm close to the user's mouth. Then, the user can speak a voice command into the microphone. The electronic device 100 can collect the voice command issued by the user through the microphone, and start the voice assistant to respond to the voice command and perform the corresponding operation. For example, the user can lift the electronic device 100 and move the microphone of the electronic device 100 near the mouth, and speak the voice command "What's the weather like today?" into the microphone. When the electronic device 100 collects the voice command through the microphone, the electronic device 100 can start the voice assistant to respond to the voice command, query today's weather, and output the voice corresponding to the query result to the user through the speaker, such as "It's sunny today". In some possible implementations, if the intention is to trigger the voice assistant to wake up without a wake-up word, the angle between the screen of the electronic device 100 and the plane must be between -60 degrees and 60 degrees.
[0056] like Figure 3B As shown, the wake-up word-free wake-up method involves an audio digital signal processor (ADSP) and an application processor (AP) in the electronic device 100, wherein the ADSP includes an inertial measurement unit (IMU) data processing module and a voice module, and the AP may include a post-processing module.
[0057] Specifically, the IMU data processing module in the ADSP can receive sensor data acquired by sensors, which can be, for example, inertial sensor data (e.g., data detected by an accelerometer and / or data detected by a gyroscope). Furthermore, the voice module in the ADSP can receive sound data collected by a microphone, which can include a user's voice commands. The ADSP can determine whether the current state of the electronic device 100 matches a specified state based on the sound data and sensor data. The specified state can include at least one of the following: the motion state of the electronic device 100, the spatial posture of the electronic device 100, the distance between the microphone of the electronic device 100 and the sound source, etc. When the electronic device 100 determines that the current state of the electronic device 100 matches the specified state, the post-processing module in the AP can further set specified recognition parameters (e.g., voiceprint, sound source angle, etc.) using the large model, and output a judgment result based on the recognition parameters. The judgment result can be used to indicate whether the specified conditions for waking up the voice assistant without a wake-up word are currently met. Based on the judgment result, the electronic device 100 can determine whether to activate the voice assistant and respond to the voice command in the sound data.
[0058] It can be seen from the above process that when the electronic device 100 detects whether the user has the need for voice interaction without a wake-up word, the electronic device 100 needs to process the received data in parallel through the IMU data processing module and the voice module (for example, parallel processing of inertial sensor data and sound data, etc.). Therefore, the electronic device 100 needs to keep the IMU data processing module and the voice module in operation in real time, which consumes a lot of power. In addition, in general, the electronic device 100 usually sets the IMU data processing module and the voice module in a low-power space. Running the IMU data processing module and the voice module in parallel also makes the electronic device 100 more dependent on the low-power space.
[0059] Therefore, the present application provides a voice wake-up method, in which the electronic device 100 can detect whether the user has made a first gesture (for example, a tapping action, etc.) through a sensor, and collect surrounding sound data through a microphone. If the electronic device 100 determines that the user has made a first gesture, the electronic device 100 detects whether there is voice activity in the sound data when the user is speaking. If there is voice activity in the sound data, the electronic device 100 can determine whether the electronic device 100 currently matches the specified state 1 corresponding to the wake-up word-free voice interaction based on the sensor data and / or sound data. If the specified state 1 is matched, the electronic device 100 starts the voice assistant and responds to the voice instructions in the sound data through the voice assistant.
[0060] It can be seen from the process of the above-mentioned voice wake-up method that the electronic device 100 detects in a serial manner whether the user has the need for voice interaction without a wake-up word. For example, when the electronic device 100 determines that the user makes the first hand gesture, it starts detecting whether there is voice activity in the sound data 1. In this way, the module for detecting voice activity does not have to be in operation in real time, which can reduce the power consumption of the electronic device 100. At the same time, since the implementation process of the module for detecting user gestures and the module for voice activity detection in the low-power space is relatively simple, the dependence of the electronic device 100 on the low-power space can also be reduced.
[0061] Figure 4 This is a flowchart of a specific implementation method of a voice wake-up method provided in an embodiment of the present application.
[0062] like Figure 4 As shown, the specific implementation process of the voice wake-up method may include:
[0063] S401. The electronic device 100 collects sensor data 1 through one or more sensors and collects sound data 1 through a microphone.
[0064] The one or more sensors may include one or more of the following: an acceleration sensor, a gyroscope sensor, a pressure sensor, an ambient light sensor, etc. Sensor data 1 may include acceleration data, and in addition, sensor data 1 may also include one or more of the following: angular velocity data, ambient light intensity data, pressure data, etc.
[0065] The sound source orientation of the sound data 1 can be located in the surrounding environment of the electronic device 100. The sound data 1 can be collected by a microphone on the electronic device 100, or by multiple microphones on the electronic device 100. In one possible implementation, the electronic device 100 can collect the sound data 1 through two microphones. Among them, one microphone can be set at the end close to the front camera (i.e., the top), called the top microphone (referred to as the top microphone); one microphone can be set at the end close to the charging port (i.e., the bottom), called the bottom microphone (referred to as the bottom microphone).
[0066] S402 . The electronic device 100 determines whether a first gesture (eg, tapping) of the user is detected based on the sensor data 1 .
[0067] For example, the first gesture may be a tapping gesture, such as tapping the back cover, side frame, or display screen of the electronic device 100. In other words, this application does not limit the location of the tap. The tapping may be one, two, or three times, and this application does not limit the number of taps.
[0068] Specifically, the electronic device 100 can detect the change of acceleration on the three axes of X-axis, Y-axis and Z-axis through the acceleration sensor, thereby detecting the user's first gesture. Among them, the X-axis is the horizontal direction of the plane where the touch screen is located, the Y-axis is the vertical direction of the plane where the touch screen is located, and the Z-axis is the vertical direction running through the touch screen. In a possible scenario, when the electronic device 100 is placed on the desktop, the electronic device 100 is in a stationary state. At this time, the change in acceleration is zero, or close to zero, such as the change in acceleration of the X-axis, Y-axis and Z-axis is less than or equal to 0.1g, then it can be determined that the electronic device 100 is in a stationary state. When the user's finger taps a certain tapping point on the touch screen, under the action of mechanical force, the change in acceleration is not zero, or does not approach zero, such as the change in acceleration of the X-axis, Y-axis and Z-axis is greater than 0.1g, then the user's tapping gesture can be detected.
[0069] S403. When the user's first gesture is detected, the electronic device 100 detects whether the sound data 1 has voice activity.
[0070] When the user speaks, the sound data 1 collected by the electronic device 100 through the microphone will contain the voice activity caused by the user speaking. Therefore, the electronic device 100 can determine whether the user speaks by detecting whether there is voice activity in the sound data 1.
[0071] Specifically, in a possible implementation, the electronic device 100 may divide the sound data 1 into a plurality of short periods of sound data, and the length of each short period is generally 10ms to 30ms. Then, the electronic device 100 may calculate the energy and zero-crossing rate of each short period. If there is a short period of time whose energy is greater than a specified threshold and / or the zero-crossing rate is greater than a specified value A1, the electronic device 100 determines that there is voice activity in the sound data 1. The zero-crossing rate refers to the number of times the sound signal crosses the zero point within a period of time. When a person speaks, the vibration of the vocal cords will produce alternating positive and negative waveforms, and these waveforms will cross the zero point. Therefore, the zero-crossing rate can be used to indicate whether there is voice activity in the sound data 1.
[0072] S404. When the sound data 1 has voice activity, the electronic device 100 starts and determines whether the current state matches the specified state 1 corresponding to the wake-up word-free voice interaction based on the sound data 1 and / or sensor data 1 through the breath wake-up model.
[0073] Among them, the specified state 1 may include one or more of the following: human voices are collected within a preset range from the microphone, the electronic device 100 is in a preset posture (for example, the angle between the display screen of the electronic device 100 and the plane is at a preset angle such as -60 to 60 degrees, etc.), and the electronic device 100 is in a preset environment (for example, the electronic device 100 is in a bright light environment, etc.).
[0074] Among them, the breath awakening model can be used to determine whether the user has the intention to perform voice interaction without the wake-up word. Exemplarily, the breath awakening model can be based on neural network training such as recurrent neural network (RNN), convolutional neural network (CNN), Transformer model, end-to-end model, etc. In one possible implementation, the electronic device 100 can realize voice interaction without the wake-up word based on the interaction method of breath awakening. Among them, breath awakening refers to the user issuing a voice command by facing the electronic device 100 with his mouth and within a preset range of the microphone of the electronic device 100, triggering the voice assistant and the user to perform voice interaction. In this way, the user can put the electronic device 100 near his mouth and speak directly to the electronic device 100 to trigger the electronic device 100 to enter the working state of voice interaction without using a specific wake-up word or pressing a button.
[0075] In one possible implementation, when the user speaks at different distances from the microphone, different airflows will be formed on the microphone. For example, when the user speaks close to the microphone, if the speech content includes consonants such as "b, c, d, f, j, k, l, p, q, r, s, t, v, w, x, y, z", a popping sound will be caused on the microphone. In this way, the breath awakening model can identify the popping sound features in the sound data 1 for the input sound data 1. At the same time, the breath awakening model can identify the time delay of the user's speaking voice reaching different microphones (such as the top microphone and bottom microphone mentioned above) based on the sound data 1, thereby determining the distance between the user and the microphone when speaking. In this way, when the breath awakening model detects that the sound data 1 includes a human voice within the preset range close to the microphone, it determines the designated state 1 corresponding to the current matching wake-up word-free voice interaction.
[0076] In one possible implementation, the breath wake-up model can also determine whether the electronic device 100 has captured a human voice within a preset range of the microphone based on the sound data 1 and the pressure data. If a popping sound feature is detected in the sound data 1, and based on the time delay for the user's voice to reach different microphones and the pressure data generated by the user on the electronic device 100 when speaking that is greater than a preset pressure threshold, it is determined that the user's mouth is facing the electronic device 100 and speaking within a preset distance range from the microphone, then the designated state 1 corresponding to the current matching wake-up word-free voice interaction is determined.
[0077] In one possible implementation, the breath wake-up model can also determine whether the current state matches the specified state 1 corresponding to the wake-up word-free voice interaction based on the sound data 1, angular velocity and / or acceleration. For example, if a pop feature is detected in the sound data 1, and the time delay for the user's voice to reach different microphones is used to determine that the user's mouth is facing the electronic device 100 and is speaking within a preset distance range from the microphone. At the same time, the breath wake-up model determines that the angle between the display screen of the electronic device 100 and the plane is between -60 and 60 degrees based on the angular velocity and / or acceleration, and determines that the current state 1 corresponding to the wake-up word-free voice interaction is matched.
[0078] S405. When matching the specified state 1, the electronic device 100 starts the voice assistant and responds to the voice instructions in the sound data 1 through the voice assistant.
[0079] Specifically, when the microphone collects sound data 1, the electronic device 100 can determine whether the current state matches the specified state 1 based on the sound data 1, and can also detect the voice instructions in the sound data 1 through the voice assistant, so that the voice assistant responds to the voice instructions in the sound data 1 to perform corresponding operations and meet the user's intentions.
[0080] Figure 5 This is a flow chart of a method for processing sound data 1 provided in an embodiment of the present application.
[0081] like Figure 5As shown, the electronic device 100 can collect sound data 1 in real time through two microphones, such as microphone 1 (top microphone, also known as mic1) and microphone 2 (bottom microphone, also known as mic2). Then, the electronic device 100 can copy the real-time collected sound data 1 to obtain two copies of sound data 1. The electronic device 100 can use one copy of the sound data 1 to detect whether there is voice activity and whether it matches the specified state 1, and the other copied sound data 1 can be stored in the audio data circular buffer. The audio data circular buffer can save the copied sound data 1 based on the mechanism of circular buffering. When the electronic device 100 determines that it matches the specified state 1, the voice assistant is started. The voice assistant can obtain the buffered sound data 1 (i.e., the aforementioned copied sound data 1) from the audio data circular buffer. Then, the voice assistant can recognize the voice command in the copied sound data 1 and execute the corresponding operation in response to the voice command to meet the user's intention.
[0082] In a possible implementation, when the electronic device 100 does not detect the user's first gesture, the electronic device 100 can start execution from S401, and at this time, the detection of voice activity is not performed. Alternatively, the electronic device 100 can detect whether the sound data 1 includes a fixed wake-up word (e.g., "Hello YOYO"). If the sound data 1 includes the fixed wake-up word, the electronic device 100 starts the voice assistant and responds to the user's voice command through the voice assistant.
[0083] In a possible implementation, if the electronic device 100 does not detect voice activity in the sound data 1, the electronic device 100 can start execution from S401.
[0084] In a possible implementation, if the electronic device 100 determines that it does not match the specified state 1, the electronic device 100 can start execution from S401. Alternatively, the electronic device 100 can detect whether the sound data 1 includes a fixed wake-up word (e.g., "Hello YOYO"). If the sound data 1 includes the fixed wake-up word, the electronic device 100 starts the voice assistant and responds to the user's voice command through the voice assistant. At this time, it can be understood that the voice assistant no longer obtains the sound data 1 from the audio data circular buffer, but receives the user's voice data collected in real time by the microphone and recognizes the voice command from the voice data to respond to the voice command.
[0085] In one possible implementation, when the electronic device 100 determines that it matches designated state 1, it may send a breath wake-up event to the wake-up engine platform. The wake-up engine platform can register multiple applications. After these applications are registered, the wake-up engine platform can obtain the service status of the registered applications and identify the service status of the registered applications with corresponding voice event status flags. The service status of a registered application indicates the operating status of the registered application, for example, whether the registered application is a foreground application or a background application, and whether the focus of the user's operation is within the registered application. For example, for audio and video software, the audio or video can be set as a preset service status; for office software such as spreadsheets and documents, the cursor can be set to blink in the document, indicating "input is in progress," as a preset service status. In one example, when an application registers, the registration information sent by the application may include the current service status of the application. In other examples, the service status of the application may be reported periodically by the registered application, or the wake-up engine platform may notify the registered application to report the service status.
[0086] In response to a breath wake-up event, the wake-up engine platform can query each registered application's service status based on the corresponding voice event status flag to see if it meets its corresponding preset service status conditions. If so, the registered application that meets the preset service status conditions is determined to be the target application. The wake-up engine platform can then trigger the target application to retrieve the cached sound data 1 from the audio data circular buffer. The target application can then recognize the voice command in the sound data 1 and perform the corresponding operation in response to the voice command, satisfying the user's intention.
[0087] Each registered application can be either a foreground application or a background application running on an electronic device. Foreground applications can fully display all of the application's functions and occupy a large amount of system resources, while background applications can perform lightweight tasks and cannot realize all of the application's functions. In one example, only one target application can exist at a time. Therefore, when setting the preset service status conditions for each application, it is necessary to ensure that each preset service status condition is exclusive, that is, the preset service status conditions for each application can ensure that only one target application exists at a time.
[0088] Figure 6A A schematic diagram of a device architecture related to the voice wake-up method provided in an embodiment of the present application.
[0089] like Figure 6A As shown, the device architecture may include a system-on-chip (SOC) of the electronic device 100, two microphones (microphone 1 and microphone 2), and sensors (including the aforementioned acceleration sensor, gyroscope sensor, and pressure sensor, etc.).
[0090] The SOC may include an application processor AP and an advanced digital signal processor ADSP. The application processor AP may include a breath awakening algorithm model and a voice assistant. The breath awakening algorithm model may include a breath awakening first-level algorithm model (which may be referred to as a breath awakening first-level algorithm) and a breath awakening second-level algorithm model (which may be referred to as a breath awakening second-level algorithm). The breath awakening second-level algorithm model has a more complex network structure than the breath awakening first-level algorithm model, and the calculation accuracy is also higher. The advanced digital signal processor ADSP may include a low-power module LPI and an audio data circular buffer. The LPI may include a voice processing module (Audio PD) and a sensor data processing module (Sensor PD). The voice processing module may include a voice activity detection module, and the sensor data processing module may include a first gesture detection module. The data flow of the device architecture may include:
[0091] S1. Electronic device 100 collects sound data 1 through microphone 1 and microphone 2, and copies sound data 1 to obtain two copies of sound data 1. One copy of sound data 1 can be transmitted to the voice activity detection module, and the copied sound data 1 can be transmitted to the audio data circular buffer for caching.
[0092] The electronic device 100 may collect sensor data 1 through one or more sensors and transmit the sensor data to the sensor data processing module. The first gesture detection module in the sensor data processing module may determine whether a first gesture of the user is detected based on the acceleration data.
[0093] S2. If the first gesture detection module detects the first gesture of the user, the voice activity detection module is activated. The voice activity detection module detects whether the sound data 1 contains voice activity.
[0094] S3. If there is voice activity in sound data 1, the voice activity detection module activates the breath wakeup level 1 algorithm model. The breath wakeup level 1 algorithm model obtains sound data 1 from the voice activity detection module and / or obtains sensor data 1 from the sensor data processing module. Based on sound data 1 and / or sensor data 1, the breath wakeup level 1 algorithm model determines whether the current state matches designated state 1 corresponding to wake-up word-free voice interaction.
[0095] In a possible implementation, the breath awakening first-level algorithm model may output a confidence level. If the confidence level is greater than a specified value B1, it is determined that the current state matches the specified state 1; otherwise, it is determined that the current state does not match the specified state 1.
[0096] S4. When the breath wakeup algorithm model determines that the current state matches designated state 1, the breath wakeup algorithm model triggers the breath wakeup algorithm model to run, and sends the sound data 1 and / or sensor data 1 to the breath wakeup algorithm model for execution. The breath wakeup algorithm model can further and more accurately determine whether the current state matches designated state 1 corresponding to the wake-up word-free voice interaction based on the sound data 1 and / or sensor data 1.
[0097] In one possible implementation, the breath-based awakening secondary algorithm model can output a confidence score. If this confidence score is greater than a specified value C1, the current state is determined to match specified state 1; otherwise, the current state is determined to not match specified state 1. The specified value C1 can be greater than the specified value B1. In other words, the breath-based awakening secondary algorithm model has higher accuracy than the breath-based awakening primary algorithm model.
[0098] S5. If the breath awakening secondary algorithm model determines that the current state matches the specified state 1, the breath awakening secondary algorithm model starts the voice assistant.
[0099] In another possible implementation, the breath wakeup algorithm model may not be divided into a primary breath wakeup algorithm model and a secondary breath wakeup algorithm model. The electronic device 100 may directly use the breath wakeup algorithm model to determine whether the current state matches designated state 1 corresponding to the wake-up word-free voice interaction based on the sound data 1 and / or sensor data 1. If the breath wakeup algorithm model determines that the current state matches designated state 1, the voice assistant is activated.
[0100] S6. After the voice assistant is started, the cached sound data 1 (i.e., the aforementioned copied sound data 1) can be obtained from the audio data circular buffer. Then, the voice assistant can recognize the voice instructions in the copied sound data 1 and perform corresponding operations in response to the voice instructions to meet the user's intentions.
[0101] In a possible implementation, the audio data circular buffer may also be set in the application processor AP. In other words, this application does not limit the location of the audio data circular buffer.
[0102] Figure 6B A schematic diagram of another device architecture related to the voice wake-up method provided in an embodiment of the present application.
[0103] like Figure 6B shown, and Figure 6A The difference is that the breath awakening algorithm model can be set in the double data rate synchronous dynamic random access memory (DDR) in addition to the LPI in the ADSP ( Figure 6Bnot shown). Figure 6B Further description of the device architecture and data flow description can be found in Figure 6A The description in the illustrated embodiment will not be repeated here.
[0104] In one possible implementation, the breath wake-up algorithm model may not be divided into a breath wake-up primary algorithm model and a breath wake-up secondary algorithm model. The electronic device 100 may set the breath wake-up algorithm model on the DDR side of the ADSP. The electronic device 100 may directly use the breath wake-up algorithm model to determine whether the current state matches the specified state 1 corresponding to the wake-up word-free voice interaction based on the sound data 1 and / or sensor data 1. If the breath wake-up algorithm model determines that the current state matches the specified state 1, the voice assistant is started.
[0105] Figure 7 A schematic diagram of the execution logic of a voice wake-up method provided in an embodiment of the present application.
[0106] like Figure 7 As shown, combined Figure 5 、 Figure 6A and Figure 6B It can be seen that the voice wake-up method provided in the present application has an execution logic that the electronic device 100 first determines whether the user's tapping gesture (that is, the first gesture mentioned above) is detected. If the electronic device 100 detects the user's tapping action, it then performs voice activity detection (that is, detects whether there is voice activity in the sound data 1). If there is voice activity in the sound data 1, the electronic device 100 can start the breath wake-up first-level algorithm model, and judge based on the sound data 1 and / or sensor data 1 whether the current state 1 corresponding to the voice interaction without the wake-up word is matched. If the breath wake-up first-level algorithm model determines that the current state 1 is matched, the electronic device 100 starts the breath wake-up second-level algorithm model, and then further and more accurately judges based on the sound data 1 and / or sensor data 1 whether the current state 1 corresponding to the voice interaction without the wake-up word is matched.
[0107] In this way, the present application implements the voice wake-up method by means of a serial detection mechanism, which can effectively reduce the dependence on the LPI in the ADSP, save the power consumption of the electronic device 100, and ensure the accuracy of the user intention judgment.
[0108] Figure 8 A schematic diagram of the software architecture related to a voice wake-up method provided in an embodiment of the present application.
[0109] like Figure 8As shown, the software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to illustrate the software structure of the electronic device 100. The layered architecture divides the software into several layers. Each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, namely, from top to bottom, the application layer, the application framework layer, the system library layer, the hardware abstraction layer and the kernel layer.
[0110] like Figure 8 As shown, the application layer may include a series of application packages, such as camera, calendar, memo, browser, and voice assistant, etc. Among them, the voice assistant can be used to respond to the user's voice instructions to perform corresponding operations.
[0111] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include a window manager, a content provider, a telephony manager, a resource manager, a notification manager, etc., among which:
[0112] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0113] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0114] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).
[0115] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0116] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0117] In an embodiment of the present application, the application framework layer may further include a breath awakening algorithm model. The breath awakening algorithm model may determine, based on sound data 1 and / or sensor data 1, whether the current state matches a specified state 1 corresponding to a wake-up word-free voice interaction. In one possible implementation, the breath awakening algorithm model may also be located in the hardware abstraction layer or the system library layer.
[0118] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0119] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0120] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0121] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0122] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0123] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0124] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0125] A 2D graphics engine is a drawing engine for 2D drawings.
[0126] like Figure 8As shown, the hardware abstraction layer (HAL) is located between the kernel layer and the application framework layer, serving as a link between the two. Specifically, the HAL may include a voice activity detection module and a first gesture detection module. The voice activity detection module can detect whether there is voice activity in the sound data 1, and the first gesture detection module is used to determine whether the electronic device 100 has detected the user's first gesture. In one possible implementation, the voice activity detection module and the first gesture detection module may also be located in the application framework layer or the system library layer.
[0127] The kernel layer is the layer between hardware and software. It includes at least a display driver, a camera driver, a microphone driver, and a sensor driver. The sensor driver can be used to control one or more sensors to collect data. The microphone driver can be used to drive a microphone to collect sound data around electronic device 100.
[0128] Figure 8 The software structure shown is only used to exemplify the present application and does not constitute any limitation to the present application.
[0129] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0130] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program runs on a processor, it can implement the steps performed by the electronic device in the above-mentioned various method embodiments.
[0131] The present application also provides a chip system, which includes a processing circuit interface circuit, the interface circuit being configured to receive instructions and transmit them to the processing circuit, and the processing circuit being configured to execute the instructions so that the chip system implements the steps performed by the electronic device in any method embodiment of the present application. The chip system can be a single chip or a chip module composed of multiple chips.
[0132] The term "user interface (UI)" in the specification and drawings of this application refers to the media interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface of an application is a source code written in a specific computer language such as Java and Extensible Markup Language (XML). The interface source code is parsed and rendered on the terminal device, and finally presented as content that the user can recognize, such as pictures, text, buttons and other controls. Controls, also known as widgets, are the basic elements of the user interface. Typical controls include toolbars, menu bars, text boxes, buttons, scroll bars, pictures and text. The properties and contents of controls in the interface are defined by tags or nodes, such as XML through <textview> 、 <imgview> 、 <videoview>Nodes such as <head> and <body> are used to specify the controls contained in the interface. A node corresponds to a control or attribute in the interface, and the node is presented as user-visible content after parsing and rendering. In addition, many applications, such as hybrid applications, usually also contain web pages in their interfaces. A web page, also known as a page, can be understood as a special control embedded in the application interface. A web page is a source code written in a specific computer language, such as hypertext markup language (HTML), cascading style sheets (CSS), JavaScript (JS), etc. The web page source code can be loaded and displayed as user-recognizable content by a browser or a web page display component with similar functions to a browser. The specific content contained in a web page is also defined by tags or nodes in the web page source code, such as HTML through <body>. 、 、 <video> 、 <canvas>To define the elements and attributes of a web page.
[0133] A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operations that uses graphics. It can be an icon, window, control, or other interface element displayed on the display of an electronic device. Controls can include icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, and other visual interface elements.
[0134] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).
[0135] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0136] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.< / canvas> < / video> < / videoview> < / imgview> < / textview>
Claims
1. A voice wake-up method, characterized in that: include: The electronic device collects first sound data through a microphone; When the first gesture is detected, the electronic device detects whether the first sound data has voice activity; If the first sound data has voice activity, the electronic device determines whether the state of the electronic device is a first state; wherein the first state is a state corresponding to the electronic device performing voice interaction without a wake-up word; If the state of the electronic device is the first state, the electronic device starts the voice assistant.
2. The method according to claim 1, characterized in that The method further comprises: The electronic device collects first sensor data; If the first sound data contains voice activity, the electronic device determines whether the state of the electronic device is the first state, specifically including: If the first sound data contains voice activity, the electronic device determines whether the state of the electronic device is a first state based on the first sound data and / or the first sensor data.
3. The method according to claim 2, characterized in that The first state includes one or more of the following: the electronic device collects a human voice within a preset range from the microphone, the electronic device is in a preset posture, and the electronic device is in a preset environment.
4. The method according to claim 2, characterized in that If the first sound data contains voice activity, the electronic device determines, based on the first sound data and / or the first sensor data, whether the state of the electronic device is a first state, specifically including: If the first sound data has voice activity, the electronic device determines whether the state of the electronic device is a first state based on the first sound data and / or the first sensor data by using a breath awakening first-level algorithm model; If the breath awakening primary algorithm model determines that the state of the electronic device is the first state, the electronic device uses the breath awakening secondary algorithm model to determine whether the state of the electronic device is the first state based on the first sound data and / or the first sensor data; wherein the breath awakening secondary algorithm model has a higher accuracy than the breath awakening primary algorithm model; If the breath awakening secondary algorithm model determines that the state of the electronic device is the first state, the electronic device determines that the state of the electronic device is the first state.
5. The method according to claim 1, wherein The method further comprises: The electronic device responds to the voice command in the first sound data through the voice assistant.
6. The method according to claim 5, characterized in that After the electronic device collects the first sound data through the microphone, the method further includes: The electronic device copies the first sound data and stores the copied first sound data in an audio circular buffer; The electronic device responds to the voice command in the first sound data through the voice assistant, specifically including: The electronic device obtains the copied first sound data in the audio circular buffer through the voice assistant, and recognizes the voice command in the copied first sound data; The electronic device executes the operation corresponding to the voice instruction through the voice assistant.
7. The method according to claim 4, characterized in that The breath awakening level 1 algorithm and the breath awakening level 2 algorithm are located in an application processor AP; or, the breath awakening level 1 algorithm is located in an audio digital signal processor ADSP, and the breath awakening level 2 algorithm is located in the AP.
8. An electronic device, characterized in that: The electronic device includes: one or more processors, one or more sensors and a memory; the memory is coupled to the one or more processors, the one or more processors are coupled to the one or more sensors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method described in any one of claims 1-7.
9. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the processor is used to call computer instructions to enable the electronic device to execute the method as described in any one of claims 1-7.
10. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 7.
11. A computer program product, characterized in that The device comprises a computer program, which, when executed by a processor, causes the electronic device to perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice-based man-machine interaction method and system, electronic equipment and medium
CN114067779A
Voice interaction method, electronic equipment and computer readable storage medium
CN118057527A
Voice assistant wake-up method and wake-up device
CN118057805A
Interaction method and device, electronic equipment and readable storage medium
CN119007717A
Low power detection of a voice control activation phrase
US20160253997A1