Device wake-up methods, devices, equipment, storage media and program products
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-10
- Publication Date
- 2026-08-14
AI Technical Summary
然而这种方式的误唤醒率较高
[0018]本公开实施例提供一种设备唤醒方法、装置、设备、存储介质及程序产品。该方法中,设备持续监听声音信号,在未监听到人声信号之前始终保持低功耗状态运行。设备在监听到声音信号中存在人声信号后,获取图像,基于图像以及监听到的人声信号执行双因子(即:注视方向和唇动)唤醒验证流程;在两个因子对应的唤醒条件均满足时(即:用户的注视方向朝向设备,且唇动特征与人声信号匹配),激活设备的语音互动模块。这样,便确保仅用户注视设备且存在语音输入时才激活设备的语音互动模块,使得设备从低功耗模式切换至语音互动模式。由于借助双因子(注视方向和唇动)进行唤醒验证,设备被唤醒的时机更加准确,设备的误唤醒率降低,同时设备的整体功耗水平也得以降低。
Smart Images

Figure CN122575360A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of smart device technology, and in particular to a device wake-up method, apparatus, device, storage medium, and program product. Background Technology
[0002] Smart devices are becoming increasingly prevalent in people's lives. For example, intelligent robots can perform various actions in response to user commands; in-vehicle intelligent devices can control the operation of vehicle components such as air conditioning and ambient lighting in response to user commands. These smart devices generally remain in standby mode before responding to user commands and are awakened based on preset logic.
[0003] In related technologies, wake words are pre-configured for smart devices to enable automatic wake-up. Smart devices can determine whether to wake up by detecting whether the user's voice contains the wake word. However, this method has a relatively high false wake-up rate. Summary of the Invention
[0004] In view of this, the present disclosure provides a device wake-up method, apparatus, device, storage medium and program product.
[0005] Firstly, it provides methods for waking up the device, including: Listening to sound signals; In response to the detection of a human voice signal in the sound signal, an image is acquired, and the user's gaze direction and lip movement characteristics are determined based on the image; In response to the user's gaze direction being toward the device and the lip movement characteristics matching the human voice signal, the device's voice interaction module is activated.
[0006] Optionally, the method further includes: When the number of users is one, the voice interaction module is controlled to treat the user as the interaction object; When there are multiple users, the voice interaction module is controlled to select the user with the highest score as the interaction target, wherein the score is used to characterize the user's willingness to interact with the device via voice.
[0007] Optionally, the score is calculated based on the user's sound source location, the user's gaze direction, and the distance between the user and the device.
[0008] Optionally, the method further includes: If the user's gaze direction deviates from the device or the lip movement characteristics do not match the human voice signal, the audio signal continues to be monitored.
[0009] Optionally, the user's gaze direction is towards the device, including: the angle between the user's gaze direction and the optical axis direction of the device's camera is less than an angle threshold; and / or The matching of lip movement features with the human voice signal includes: the matching degree between the user's lip opening and closing features and the energy features of the human voice signal is greater than a matching threshold.
[0010] Optionally, in response to detecting the presence of a human voice signal in the sound signal, acquiring an image and determining the user's gaze direction and lip movement features based on the image includes: Based on the image and the human voice signal, determine the deviation between the user's sound source location and facial orientation; In response to the deviation being less than or equal to a deviation threshold, the user's gaze direction and lip movement characteristics are determined based on the image.
[0011] Optionally, the method further includes: In response to the detection of a human voice signal in the sound signal, at least one of the light intensity signal and the image is acquired; In response to determining that the device meets the visual judgment failure state based on at least one of the light intensity signal and the image, the sound source location of the human voice signal is obtained; In response to the sound source location of the human voice signal satisfying the sound source validity condition and the human voice signal containing speech keywords, the voice interaction module of the device is activated.
[0012] Optionally, the method further includes: In response to the fact that the user's lip movement features are not visible in the image, the user's gaze direction and the sound source location of the human voice signal are obtained; In response to the user's gaze being directed toward the device and the sound source location of the human voice signal satisfying the sound source validity condition, the device's voice interaction module is activated.
[0013] Optionally, the device is a smart device with walking capabilities, and the method further includes: When the device is in a walking state, in response to detecting a human voice signal in the sound signal, it continues to monitor the sound signal; or When the device is in a walking state, in response to the detection of a human voice signal in the sound signal, the device is controlled to stop walking and the device wake-up verification process is executed.
[0014] Secondly, a device wake-up device is provided, comprising: The monitoring module is used to monitor sound signals; The determination module is used to acquire an image in response to the detection of a human voice signal in the sound signal, and to determine the user's gaze direction and lip movement characteristics based on the image; An activation module is used to activate the device's voice interaction module in response to the user's gaze direction being toward the device and the lip movement characteristics matching the human voice signal.
[0015] Thirdly, an electronic device is provided, including a processor and a memory, the memory storing at least one computer instruction, wherein when the processor executes the at least one computer instruction, the electronic device performs the device wake-up method as described above.
[0016] Fourthly, a storage medium is provided that stores at least one computer instruction, which, when executed by a processor of a computer device, causes the computer device to perform the device wake-up method as described above.
[0017] Fifthly, a computer program product is provided, comprising at least one computer instruction that, when executed by a processor of a computer device, causes the computer device to perform the device wake-up method as described above.
[0018] This disclosure provides a device wake-up method, apparatus, device, storage medium, and program product. In this method, the device continuously monitors sound signals and maintains a low-power operation until a human voice signal is detected. Upon detecting a human voice signal within the sound signal, the device acquires an image and performs a two-factor (i.e., gaze direction and lip movement) wake-up verification process based on the image and the detected human voice signal. When the wake-up conditions corresponding to both factors are met (i.e., the user's gaze direction is towards the device, and the lip movement feature matches the human voice signal), the device's voice interaction module is activated. This ensures that the device's voice interaction module is activated only when the user is looking at the device and there is voice input, allowing the device to switch from a low-power mode to a voice interaction mode. Because wake-up verification is performed using two factors (gaze direction and lip movement), the timing of device wake-up is more accurate, the false wake-up rate is reduced, and the overall power consumption of the device is also reduced. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart of a device wake-up method provided in this disclosure embodiment; Figure 2 A flowchart of another device wake-up method provided in this disclosure embodiment; Figure 3 A schematic diagram illustrating the logic of calculating a score, provided for an embodiment of this disclosure; Figure 4 A flowchart illustrating another device wake-up method provided in this disclosure embodiment; Figure 5 A flowchart of another device wake-up method provided in this disclosure embodiment; Figure 6 A flowchart of another device wake-up method provided in this disclosure embodiment; Figure 7 A flowchart of another device wake-up method provided in this disclosure embodiment; Figure 8 A flowchart of a device wake-up method provided in this disclosure embodiment; Figure 9 A flowchart of a device wake-up method provided in this disclosure embodiment; Figure 10 A logical schematic diagram of a device wake-up method provided in an embodiment of this disclosure; Figure 11 A schematic diagram of a lip-movement-based wake-up verification provided for an embodiment of this disclosure; Figure 12 A logical schematic diagram of a device wake-up method provided in an embodiment of this disclosure; Figure 13 This is a structural diagram of a device wake-up device provided in an embodiment of the present disclosure; Figure 14 This is a structural diagram of an electronic device provided in an embodiment of the present disclosure.
[0021] The accompanying drawings have illustrated specific embodiments of this disclosure, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this disclosure to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0022] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0023] Unless otherwise defined, all technical terms used in the embodiments of this disclosure have the same meaning as commonly understood by one of ordinary skill in the art.
[0024] To make the technical solutions and advantages of this disclosure clearer, the embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings.
[0025] Firstly, embodiments of this disclosure provide a device wake-up method. This method is executed by the device, specifically by a controller integrated within the device, or by a dedicated controller independent of the device body (e.g., a smart remote control communicatively connected to the device body). This method can be implemented through hardware, software, or a combination of both. The method of this disclosure can be applied to various electronic devices, such as in-vehicle terminals, home display devices, robotic vacuum cleaners, bipedal robots, or quadrupedal robots, etc.
[0026] refer to Figure 1 The device wake-up method of this disclosure includes the following steps.
[0027] Step 101: Listen to the audio signal.
[0028] In this step, the device continuously monitors the audio signal and operates in the default low-power mode to reduce standby power consumption. To enable the monitoring function, the device can have a microphone array (e.g., a microphone array with four or more units) to ensure continuous Voice Activity Detection (VAD).
[0029] Optionally, the device's microphone array can operate at a lower sampling rate in the aforementioned low-power mode to reduce standby power consumption. This lower sampling rate could be, for example, 8 kHz.
[0030] Step 102: In response to the detection of human voice signals in the sound signal, acquire an image and determine the user's gaze direction and lip movement characteristics based on the image.
[0031] Optionally, the device may include a voice detection module with a built-in voice detection algorithm to identify and judge the sound signals detected by the microphone array and detect whether a human voice signal is present. This voice detection algorithm can be implemented using machine learning, deep neural networks, etc., and its implementation principle can be found in relevant technologies; no specific limitations are made here.
[0032] In this step, after detecting a human voice signal from the audio signal, the device acquires an image through its camera module (that is, the device activates its vision module at this time, which may include an RGB camera, etc.). Subsequently, the device determines the user's gaze direction and lip movement characteristics based on the image. The image here can be composed of one or more image frames.
[0033] In some embodiments, the device's controller or processing unit may incorporate a face gaze direction detection algorithm or a facial landmark detection algorithm to determine the user's gaze direction in the image by performing corresponding detection on the acquired image. The device's controller or processing unit may also incorporate a face lip feature detection algorithm or a facial landmark detection algorithm to determine the user's lip features in the image by performing corresponding detection on the acquired image. The face gaze direction detection algorithm, face lip feature detection algorithm, or facial landmark detection algorithm can be referenced from related technologies, and this disclosure does not impose specific limitations.
[0034] In some embodiments, the device can communicate with a cloud service platform, upload acquired images to the cloud service platform, request the cloud service platform to identify the user's gaze direction and lip features, and obtain the identification results from the cloud service platform. The cloud service platform may be equipped to run one or more of the aforementioned detection algorithms to help the device determine the user's gaze direction and lip features in a timely manner.
[0035] Step 103: In response to the user's gaze direction being towards the device and the lip movement characteristics matching the human voice signal, the device's voice interaction module is activated.
[0036] In this step, the device executes a two-factor wake-up verification process based on gaze direction and lip movement features, that is, determining whether the target user's gaze direction is towards the device, and determining whether the lip movement features match the human voice signal; and in response to the result of both factor wake-up verifications being "yes", the device's voice interaction module is activated.
[0037] Here, the human voice signal used can be obtained by intercepting the sound signal monitored by the device, and the interception time period can be consistent with the previous image acquisition time period. For example, the aforementioned image and the human voice signal here are respectively the image and human voice signal acquired within a certain duration (e.g., 1 second, 2 seconds, or 3 seconds) from the time it is determined that a human voice signal is detected in the sound signal.
[0038] In some embodiments, the user's gaze direction is towards the device, including: the angle between the user's gaze direction and the optical axis of the device's camera is less than an angle threshold. The threshold in this disclosure can be preset or temporarily set as needed. Optionally, the angle threshold is 20 degrees. It should be noted that the device can determine the user's gaze direction and its angle with the camera's optical axis by performing corresponding detection on the image using a built-in face gaze direction detection algorithm or face key point detection algorithm; or it can obtain the aforementioned angle through interaction with a cloud service platform.
[0039] In some embodiments, where the user's gaze is directed toward the device, the process further includes: the duration for which the angle between the user's gaze direction and the optical axis of the device's camera is less than an angle threshold reaches a duration threshold. This allows for further improvement in the accuracy of device wake-up through a gaze-direction-based wake-up verification process. Optionally, the duration threshold can be 0.5s, 0.6s, 0.7s, 0.8s, 0.9s, 1s, etc., and its specific duration can be adjusted according to the required wake-up sensitivity and accuracy of the device. Higher required wake-up sensitivity allows for a smaller duration threshold; higher required wake-up accuracy allows for a larger duration threshold.
[0040] In some embodiments, matching lip movement features with a human voice signal includes: the matching degree between the target user's lip opening and closing features and the energy features of the human voice signal is greater than a matching threshold. That is, the aforementioned lip movement features can be lip opening and closing features, and the corresponding signal features of the human voice signal can be energy features. Since the user's lip opening and closing often corresponds to the moment of sound production, and there is generally a strong correlation between lip movement opening and closing features and the energy of the sound signal, the matching degree can be calculated based on the lip opening and closing features and the energy features of the human voice signal.
[0041] Optionally, the device can use built-in facial landmark detection algorithms to detect facial regions, lip movement landmarks, or regions of interest in the mouth in an image, detect information such as lip movement amplitude, and obtain the timing signal of lip opening and closing; then, the matching degree is calculated based on the timing signal of lip opening and closing.
[0042] Optionally, the signal characteristics corresponding to the human voice signal can be specifically the energy envelope of the human voice signal.
[0043] Optional, see reference Figure 11 The aforementioned matching degree can be the time correlation coefficient between the timing signal of lip opening and closing and the energy envelope of the human voice signal. Matching lip movement features with the human voice signal includes: the time correlation coefficient between the timing signal of lip opening and closing and the energy envelope of the human voice signal is greater than or equal to a correlation coefficient threshold. This correlation coefficient threshold can be, for example, 0.6, 0.7, 0.8, 0.9, etc., and can be specifically set according to the required wake-up sensitivity and wake-up accuracy of the device. The higher the required wake-up sensitivity, the smaller the correlation coefficient threshold can be; the higher the required wake-up accuracy, the larger the correlation coefficient threshold can be. Optionally, the correlation coefficient in this disclosure can be calculated using the Pearson method.
[0044] In some embodiments, the device has a built-in voice interaction module. After the aforementioned steps automatically activate the voice interaction module, the device can respond promptly to the user's voice input and interact with the user via voice. The voice interaction module can be an intelligent module integrating multiple functions such as display, semantic recognition, and voice output.
[0045] For example, in a device with a display screen, the voice interaction module may include the display screen or be signal-connected to it. When interacting with the user, the voice interaction module can illuminate the display screen and display corresponding images on the screen (such as an LCD screen), such as an animation of someone listening, an animation of someone speaking, or an image with corresponding facial expressions while speaking. The voice interaction module may also include a microphone array or be signal-connected to it to output sound signals or receive further user voice input.
[0046] The voice interaction module can have automatic speech recognition (ASR) units, language model inference units, natural language processing (NLP) units, and text-to-speech (TTS) units, or be connected to cloud-based speech recognition units, natural language processing units, language model inference units, and speech synthesis units to achieve full-link configuration of voice interaction.
[0047] In some embodiments, when the device is a smart device such as a bipedal robot or a quadrupedal robot, the voice interaction module can also be connected to the movable parts of the device to further enhance the interaction effect with the help of the movable parts of the device. For example, the limbs of the quadrupedal robot can move accordingly to simulate limb movements.
[0048] In other words, depending on the integration level of the voice interaction module and the richness of the device's own functional components, the device can achieve different levels of voice interaction effects. For example, for bipedal or quadrupedal robots, a multimodal response of voice, facial expressions, and body movements can be output after semantic processing.
[0049] In some embodiments, when the voice interaction module is activated, the sampling frequency of the microphone array can be increased to enable the device to respond to user voice input in a timely manner. For example, the sampling frequency of the microphone array can be increased to 16 kHz.
[0050] In summary, according to the method of this disclosure, the device continuously monitors sound signals and maintains a low-power operation until a human voice signal is detected. Upon detecting a human voice signal within the sound signal, the device acquires an image and performs a two-factor (i.e., gaze direction and lip movement) wake-up verification process based on the image and the detected human voice signal. When the wake-up conditions corresponding to both factors are met (i.e., the user's gaze direction is towards the device, and the lip movement feature matches the human voice signal), the device's voice interaction module is activated. This ensures that the device's voice interaction module is activated only when the user is looking at the device and there is voice input, allowing the device to switch from low-power mode to voice interaction mode. Because wake-up verification is performed using two factors (gaze direction and lip movement), the timing of device wake-up is more accurate, the false wake-up rate is reduced, and the overall power consumption of the device is also reduced.
[0051] Furthermore, the above method actually implements a three-level progressive activation logic, and the device executing this method can, for example, integrate a fusion decision module (which can exist in software form). (See reference) Figure 10 and 12 First, the first level is low-power passive monitoring. During this stage, the device does not activate the camera or turn on the screen, with an overall power consumption increase of less than 50 mW. Second, the second level is two-factor authentication. This authentication is activated after the first level confirms the presence of human voice, and both factors must pass authentication simultaneously to proceed to the next level; otherwise, the device reverts to the first level. Finally, the third level is full-link activation. In this stage, the device activates all modules related to voice interaction, and the device engages in voice interaction with the user. When the interaction ends or there is no interaction within a predetermined time (e.g., 5 minutes), the device reverts to the first level. This three-level progressive activation effectively reduces the device's operating power consumption and lowers the false wake-up rate.
[0052] In some embodiments, reference Figure 2 The method disclosed herein also includes: Step 104: When there is only one user, control the voice interaction module to use that user as the interaction target; Step 105: When there are multiple users, the voice interaction module controls the user with the highest score as the interaction target, where the score is used to characterize the user's willingness to interact with the device via voice.
[0053] Steps 104 and 105 are not executed simultaneously or in a specific order, but rather one of them is executed at any given time based on the number of users. In the aforementioned steps 102-103, since there may be one or more users in the image, the number of users whose gaze is directed towards the device and whose lip movement features match the voice signal may also be one or more. When there are multiple users whose gaze is directed towards the device and whose lip movement features match the voice signal, it is necessary to use a score to determine the interaction target of the voice interaction module to avoid interaction confusion in multi-person dialogue scenarios.
[0054] In some embodiments, when one or more users are continuously present in the captured image frames, these users are extracted, and their gaze direction and lip movement features are obtained; users who only appear in some image frames can be ignored. In this way, users with potentially strong interactive intentions can be identified.
[0055] In other words, in this embodiment, when only one user passes the two-factor authentication, that user can be directly used as the interaction target. However, when there are multiple users, it is necessary to determine the score for each user to ascertain the strength of each user's willingness to interact with the device via voice, and then select the user with the highest willingness as the interaction target. This avoids the confusion of multi-user conversations and allows the device to accurately interact with the user with the highest willingness to interact.
[0056] The specific method for obtaining the above scores can be set according to needs or scenarios, as long as it corresponds to the user's willingness to interact.
[0057] In some embodiments, the score is calculated based on the user's sound source location, the user's gaze direction, and the distance between the user and the device. This ensures that the score accurately represents the user's willingness to interact.
[0058] Optional, see reference Figure 3 The score can be calculated based on the deviation between the user's sound source location and the optical axis of the device's camera, the deviation between the user's gaze direction and the camera's optical axis, and the distance between the user and the device.
[0059] Optionally, based on the different parameters' ability to represent user interaction intentions, different weights can be assigned to the aforementioned parameters, and a weighted sum of the different parameters can be performed to calculate the initial score. The reciprocal of the initial score is then used as the final score. For example, the initial score = 0.45. Distance +0.35 The deviation value corresponding to the gaze direction is +0.2. The deviation values corresponding to the sound source location are 0.45, 0.35, and 0.2, which are weight values. These weight values can be adjusted according to needs or experience, or obtained through calibration.
[0060] Alternatively, in some embodiments, the voice interaction module controls the user with the lowest score as the interaction target. This score is negatively correlated with the user's willingness to interact with the device via voice. In this case, the initial score can be directly used as the aforementioned score.
[0061] To perform the above calculations in a timely manner, the device's microphone array can execute a sound source localization algorithm when it detects a human voice signal in the audio signal. This algorithm obtains the azimuth angle of the sound source corresponding to the human voice signal (with an accuracy of, for example, less than or equal to 15°) and records it in an internal buffer. The azimuth angle here can be the angle between the line connecting the sound source azimuth to the device's camera and the camera's optical axis. This angle represents the deviation value corresponding to the aforementioned sound source azimuth. The deviation value between the user's gaze direction and the camera's optical axis can be characterized by the angle between the gaze direction and the camera's optical axis, and its acquisition method can be referred to the aforementioned embodiment.
[0062] Optionally, the device's camera module may include a depth camera, which is used to acquire the distance between the device and the user. In other embodiments, radar, infrared, or other methods may also be used to measure the distance between the device and the user; this disclosure does not impose specific limitations.
[0063] This embodiment presents an arbitration method for situations involving multiple users. Specifically, when multiple users meeting the wake-up criteria are detected, the user with the strongest willingness to interact with the device can be determined by combining parameters such as the sound source location, the distance between the user and the device, and the viewing angle deviation. Other users will not trigger voice interaction, thus avoiding chaotic multi-user voice interaction. This embodiment is particularly applicable in home scenarios (with multiple users) or scenarios where other devices are broadcasting voice messages (e.g., a television playing audio), effectively reducing the false wake-up rate in these scenarios.
[0064] In some embodiments, reference Figure 4 The method in this disclosure embodiment further includes: Step 106: In response to the user's gaze direction deviating from the device or the lip movement characteristics not matching the human voice signal, continue to monitor the sound signal.
[0065] In this step, if the verification process corresponding to at least one of the two factors fails (the gaze direction deviates from the device or the lip movement feature does not match the human voice signal), the device continues to listen to the sound signal without performing operations such as activating the voice interaction module, so that the power consumption of the device continues to remain at a low level.
[0066] In some embodiments, reference Figure 5 In step 102, in response to the detection of human voice signals in the audio signal, an image is acquired, and the user's gaze direction and lip movement features are determined based on the image, including: Step 1021: In response to the detection of human voice signals in the sound signal, determine the location of the sound source corresponding to the human voice signals.
[0067] In this step, when the device detects a human voice signal, it can use a pre-built sound source localization algorithm to determine the location of the sound source corresponding to the human voice signal. The sound source localization algorithm in this disclosure can be a localization algorithm based on time delay difference, a beamforming method, a high-resolution spectrum estimation method, a machine learning method, etc. For specific details, please refer to relevant technologies. This disclosure does not make any specific limitations.
[0068] Step 1022: In response to the sound source location corresponding to the human voice signal satisfying the sound source validity condition, acquire an image, and determine the user's gaze direction and lip movement characteristics based on the image.
[0069] Optionally, the sound source validity condition is that the sound source of the human voice signal is located within a certain range of the device.
[0070] In this step, the device can determine whether the sound source location meets the sound source validity condition to determine whether the sound source of the human voice signal is located within a certain range of the device. When the sound source location is within a certain range of the device, it can be considered that the human voice signal may be emitted by the user towards the device. At this time, an image is acquired to further determine whether the user is a user with the intention to interact; otherwise, the human voice signal can be considered to be only a background sound signal. Here, the certain range of the device can be set according to the wake-up sensitivity required by the device, such as the shooting range of the device's camera, or the central area of the camera's shooting range, etc. Optionally, the deviation angle between the line connecting the sound source location and the device's camera and the optical axis of the camera can be calculated. When the deviation angle is less than or equal to the deviation angle threshold (the sound source location meets the sound source validity condition), the device acquires an image and executes the subsequent process.
[0071] Step 1023: In response to the fact that the location of the sound source corresponding to the human voice signal does not meet the sound source validity condition, continue to monitor the sound signal.
[0072] In this step, if the location of the sound source corresponding to the human voice signal does not meet the sound source validity condition, the human voice signal is considered to be only the background sound in the environment where the device is located. The device is then controlled to continue listening to the sound signal, maintain low power consumption operation, and not execute wake-up verification processes such as image acquisition.
[0073] In other words, steps 1021-1023 above are actually a pre-verification process performed using the sound source location before executing the formal two-factor wake-up verification process. This pre-verification process does not require activating the device's visual module, but can be completed solely based on the device's microphone component and built-in algorithms. It can filter out the user's voice signal from outside the device's range in advance, effectively reducing the probability of the device being falsely woken up, and at the same time, effectively reducing the device's operating power consumption.
[0074] In some embodiments, reference Figure 6 Step 102 also includes: Step 601: Based on the image and voice signal, determine the deviation between the user's sound source location and facial orientation.
[0075] In this step, before executing the formal wake-up verification process, the deviation between the user's facial orientation and the user's sound source location is determined. The user's facial orientation can be determined by the device based on an image, while the target user's sound source location can be determined by the device based on the human voice signal.
[0076] Optionally, the device may have a built-in face orientation detection algorithm to directly determine the user's facial orientation based on the image. The facial orientation can also be characterized by the angle of the face relative to the camera's optical axis. The face orientation detection algorithm can be based on facial landmark detection algorithms; specific details can be found in related technologies, and this application does not impose any specific limitations.
[0077] Optionally, the device can determine the location of the user's sound source by referring to the previous embodiments, which can also be characterized by the angle of the sound source location relative to the optical axis of the camera.
[0078] Based on this, the deviation between the user's sound source location and facial orientation can be determined. This deviation represents the difference between the sound source location and the user's actual facial orientation. The greater the deviation, the more likely the user is communicating with other nearby users rather than intending to interact with the device.
[0079] Step 602: In response to the deviation being greater than the deviation threshold, continue listening to the sound signal.
[0080] In this step, if the deviation exceeds the deviation threshold, it can be assumed that the user does not intend to interact with the device via voice. At this point, the voice signal is identified as a side conversation, and the device is controlled to continue operating in low-power mode without executing subsequent processes such as image acquisition and two-factor wake-up verification. This avoids waking up the device's voice interaction module, thereby reducing the false wake-up rate and lowering the device's operating power consumption.
[0081] Step 603: In response to the deviation being less than or equal to the deviation threshold, determine the user's gaze direction and lip movement features based on the image.
[0082] In this step, when the deviation is less than or equal to the deviation threshold, it can be assumed that the user may intend to interact with the device via voice, thereby controlling the device to execute the subsequent wake-up verification process to ensure timely response to the user's voice input.
[0083] Optionally, when the user's facial orientation is characterized by the angle of the face relative to the camera's optical axis, and the sound source location is characterized by the angle of the sound source's location relative to the camera's optical axis, the aforementioned deviation threshold can be 30°, 45°, 60°, etc., specifically set according to the required wake-up sensitivity. The higher the required wake-up sensitivity, the larger the aforementioned deviation threshold can be.
[0084] Steps 601-603 also provide a method for pre-verification before executing the formal wake-up verification process. Compared with the previous steps 1021-1023, steps 601-602 here not only utilize the sound source location but also need to use the device's camera module to acquire images to determine the user's facial orientation. However, the pre-verification effect is superior, avoiding false wake-ups caused by more difficult-to-judge side-talk scenarios. Thus, even in scenarios with complex side-talk, the device can maintain low-power operation without frequently executing complex wake-up verification processes.
[0085] In some embodiments, reference Figure 7 The method disclosed herein also includes: Step 701: In response to the detection of a human voice signal in the sound signal, acquire at least one of the light intensity signal and the image.
[0086] In this step, after detecting a human voice signal, the device can obtain the ambient light intensity around the device using its integrated light intensity sensor, or from a light intensity sensor connected to it but independent of the device, and also use the device's camera module to acquire an image. The light intensity signal and image can be used to determine whether the device meets the visual judgment failure state. When the device meets the visual judgment failure state, the wake-up verification process involved in step 103 may not execute normally or the execution result may always be that the voice interaction module cannot be activated, even though there may actually be normal user voice interaction input. Therefore, this embodiment can be used to perform wake-up verification in such edge scenarios.
[0087] Step 702: In response to determining that the device meets the visual judgment failure state based on at least one of the light intensity signal and the image, the sound source location of the human voice signal is obtained.
[0088] In this step, the device determines whether the current device meets the visual judgment failure state based on at least one of the light intensity signal and the target image; if it does, the location of the sound source of the human voice signal is obtained.
[0089] Optionally, when the light intensity signal is below a light intensity threshold, the device is determined to be in a visual judgment failure state. This light intensity threshold can be pre-calibrated or set by the user; it represents the ambient light intensity at which the captured image is not clear enough to support subsequent visual judgments (such as gaze direction and lip movement determination). In other words, when the light intensity signal is below the light intensity threshold, it is difficult to extract valid information (such as the user's gaze direction and lip movements) from the image obtained by the device's camera module, therefore the device is considered to be in a visual judgment failure state.
[0090] Optionally, if the extraction of the user's gaze direction or lip movement based on the image fails, the device can be determined to be in a visual judgment failure state. This determination logic can be executed independently of the preceding light intensity determination logic; both can be executed, or one can be selected for execution. This determination logic can be used in edge scenarios such as when the ambient light intensity is strong but the user is in a strong backlit position, making it difficult to extract useful information from the captured image.
[0091] Step 703: In response to the sound source location of the human voice signal satisfying the sound source validity condition and the human voice signal containing speech keywords, the voice interaction module of the device is activated.
[0092] In this step, a corresponding downgraded two-factor wake-up verification process is provided for scenarios where the device is in a visual judgment failure state. Specifically, the device's voice interaction module can be activated when the sound source location of the human voice signal meets the sound source validity condition and the human voice signal includes voice keywords. In other words, when visual judgment cannot be performed, it can automatically downgrade to a two-factor wake-up verification method based on sound source location and voice keywords, avoiding the inability to wake up the device when visual judgment fails.
[0093] In some embodiments, reference Figure 8 The method disclosed herein also includes: Step 801: In response to the fact that the user's lip movement features are not visible in the image, obtain the user's gaze direction and the sound source location of the human voice signal.
[0094] In this step, if the extraction of the user's lip movement features based on the image fails, such as failure to extract the lip region of the face or failure to detect key lip regions, the user's gaze direction can be obtained while abandoning the acquisition of lip movement features, as well as the sound source location of the human voice signal. A degraded alternative two-factor wake-up verification method based on gaze direction and sound source location can then be executed.
[0095] In this method, the wake-up verification process corresponding to the gaze direction and the wake-up verification process corresponding to the sound source location can be referred to in the previous embodiments, and will not be repeated here.
[0096] Step 802: In response to the user's gaze direction being towards the device and the sound source location of the human voice signal meeting the sound source validity condition, the voice interaction module of the device is activated.
[0097] In this step, when the user's gaze direction meets the gaze condition and the sound source location of the human voice signal meets the sound source validity condition, the wake-up verification process for both the gaze direction and sound source location factors is considered to have passed, and the device's voice interaction module can be activated at this time.
[0098] Optionally, the angle threshold in the wake-up verification process corresponding to the gaze direction in this scenario can be relaxed, for example, from 20° to 25°, to achieve adaptive threshold updates.
[0099] Step 803: In response to the user's gaze direction deviating from the device or the sound source location of the human voice signal not meeting the sound source validity condition, continue to monitor the sound signal.
[0100] In this step, if the user's gaze direction deviates from the device or the sound source location of the human voice signal does not meet the sound source validity condition, it is considered that at least one of the wake-up verification processes for the two factors of gaze direction and sound source location has failed. At this time, the sound signal continues to be monitored, but the device's voice interaction module is not activated, so that the device continues to operate at low power.
[0101] In other words, this embodiment provides a degraded two-factor wake-up verification process for edge scenarios where lip movements are not visible, ensuring a low false wake-up rate even when the user's lip movements are not visible. This embodiment can be applied to wake-up of devices in scenarios such as when the user is wearing a mask.
[0102] In some embodiments, the device disclosed herein is a walking intelligent device (e.g., a robotic vacuum cleaner, a bipedal robot, or a quadrupedal robot). Reference Figure 9 The method disclosed herein also includes: Step 901: When the device is in walking mode, in response to the detection of a human voice signal in the sound signal, continue listening to the sound signal; or Step 902: When the device is in walking mode, in response to the detection of human voice signal in the sound signal, control the device to stop walking and execute the device wake-up verification process.
[0103] In this embodiment, the device can obtain its current motion state and determine whether it is in a walking state. Specifically, the device can control its own motion state using an internally integrated motion controller and obtain its current motion state from the motion controller in real time or periodically.
[0104] In step 901, when it is determined that it is in a walking state, even if human voice is heard, it continues to operate in low power mode and continue to listen to the sound, while suppressing the subsequent wake-up verification process, so as to avoid unnecessary wake-up caused by the vibration of various components in the walking state of the device, and reduce the false wake-up rate.
[0105] In step 902, once it is determined that the device is in a walking state and a human voice is detected, the device is controlled to stop walking in order to respond to the user's request. After stopping walking, the device can execute a wake-up verification process, which can be referred to in the previous embodiments. In this way, the vibration of various components (such as inertial sensors) during the device's walking state can also be avoided from affecting the accuracy of wake-up, ensuring a low false wake-up rate.
[0106] In some embodiments, the method further includes: ignoring the sound source location corresponding to the human voice signal when it contradicts the user's gaze direction determined based on the image. This contradiction includes situations where the deviation between the angle between the sound source location and the user's gaze direction and the camera's optical axis exceeds a threshold. In other words, when the sound source location contradicts the visual direction, the visual gaze direction is used as the standard, and the sound source location is only used as an auxiliary reference.
[0107] It should be noted that while the preceding descriptions of wake-up methods in various edge scenarios through multiple embodiments do not imply that these embodiments must be executed independently. Depending on the device's operational requirements, the execution logic in the aforementioned embodiments can coexist, be executed synchronously or sequentially, be partially executed, or be executed using strategies selected by the user. These embodiments enable a smooth degradation to a conservative control strategy based on low-level sensor fusion when the core sensor module of the device experiences partial failure.
[0108] To make the technical effects of the disclosed solution clearer, the following provides some operating performance indicators of the device when simply waking it up with keywords in the disclosed solution and related technologies.
[0109] Table 1. Comparison of the technical effects of the disclosed solution and related technical solutions.
[0110] As can be seen, the method of this disclosure can effectively reduce the false wake-up rate, and at the same time, it is more accurate in responding to edge scenarios such as dialect wake-up and multi-person scene targeted wake-up, while the device has low standby power consumption. The method of this disclosure can overcome the pain points of related technologies such as insufficient device intelligence, rigid interaction, inability to adapt to complex home scenarios, and lack of security protection, and improve the naturalness, security, and personalized experience of the interaction between smart devices and users.
[0111] Secondly, refer to Figure 13This disclosure provides a device wake-up device, including the following modules: The monitoring module 1301 is used to monitor sound signals; The determination module 1302 is used to acquire an image in response to the detection of human voice signals in the sound signal, and to determine the user's gaze direction and lip movement characteristics based on the image. The activation module 1303 is used to activate the voice interaction module of the device in response to the user's gaze direction being towards the device and the lip movement characteristics matching the human voice signal.
[0112] In some embodiments, the apparatus further includes: The control module 1304 is configured to: when there is only one user, control the voice interaction module to use the user as the interaction object; when there are multiple users, control the voice interaction module to use the user with the highest score as the interaction object, wherein the score is used to characterize the user's willingness to interact with the device via voice.
[0113] In some embodiments, the score is calculated based on the user's sound source location, the user's gaze direction, and the distance between the user and the device.
[0114] In some embodiments, the monitoring module 1301 is further configured to continue monitoring the sound signal in response to the user's gaze direction deviating from the device or the lip movement characteristics not matching the human voice signal.
[0115] In some embodiments, the user's gaze direction is towards the device, including: the angle between the user's gaze direction and the optical axis direction of the device's camera is less than an angle threshold; and / or the lip movement feature matches the human voice signal, including: the matching degree between the user's lip opening and closing feature and the energy feature of the human voice signal is greater than a matching threshold.
[0116] In some embodiments, the determining module 1302 is further configured to: determine the deviation between the user's sound source orientation and facial orientation based on the image and the human voice signal; and determine the user's gaze direction and lip movement characteristics based on the image in response to the deviation being less than or equal to a deviation threshold.
[0117] In some embodiments, the apparatus further includes: The acquisition module 1305 is configured to acquire at least one of the light intensity signal and the image in response to detecting the presence of a human voice signal in the sound signal; and to acquire the sound source location of the human voice signal in response to determining that the device meets the visual judgment failure state based on at least one of the light intensity signal and the image. The activation module 1303 is also used to activate the voice interaction module of the device in response to the fact that the sound source location of the human voice signal meets the sound source validity condition and that the human voice signal contains voice keywords.
[0118] In some embodiments, the acquisition module 1305 is further configured to: in response to the user's lip movement features being invisible in the image, acquire the user's gaze direction and the sound source location of the human voice signal; The activation module 1303 is also used to activate the voice interaction module of the device in response to the user's gaze direction being toward the device and the sound source location of the human voice signal satisfying the sound source validity condition.
[0119] In some embodiments, the device is a smart device capable of walking, and the listening module 1301 is further configured to: when the device is in a walking state, in response to detecting a human voice signal in the sound signal, continue listening to the sound signal; or The control module 1304 is also configured to, when the device is in a walking state, in response to detecting a human voice signal in the sound signal, control the device to stop walking and execute the device wake-up verification process.
[0120] It should be noted that the above module division is merely exemplary, and the device may have different module divisions based on the required functions. The device in this embodiment has the same or similar technical concept as the foregoing method embodiments, and the functions and execution details of each module can be referred to the method embodiments.
[0121] This disclosure provides a device wake-up apparatus. The device continuously monitors sound signals and operates in a low-power state until a human voice signal is detected. Upon detecting a human voice signal within the sound signal, the device acquires an image and performs a two-factor (i.e., gaze direction and lip movement) wake-up verification process based on the image and the detected human voice signal. When both wake-up conditions corresponding to these two factors are met (i.e., the user's gaze is directed towards the device, and the lip movement features match the human voice signal), the device's voice interaction module is activated. This ensures that the device's voice interaction module is activated only when the user is looking at the device and there is voice input, allowing the device to switch from low-power mode to voice interaction mode. Because wake-up verification is performed using two factors (gaze direction and lip movement), the timing of device wake-up is more accurate, the false wake-up rate is reduced, and the overall power consumption of the device is also reduced.
[0122] Thirdly, this disclosure provides an electronic device. (See reference...) Figure 14 The electronic device 1400 disclosed herein includes a processor 1401 and a memory 1402. The memory 1402 stores at least one computer instruction. When the processor 1401 executes at least one computer instruction, it causes the computer device 1400 to perform the device wake-up method as described above.
[0123] Fourthly, this disclosure provides a storage medium storing at least one computer instruction, which, when executed by a processor of a computer device, causes the computer device to perform the aforementioned device wake-up method.
[0124] Fifthly, this disclosure provides a computer program product including at least one computer instruction, which, when executed by the processor of a computer device, causes the computer device to perform the aforementioned device wake-up method.
[0125] For the technical details and effects of the embodiments of the second, third, fourth and fifth aspects, please refer to the embodiments of the first aspect, and this disclosure will not repeat them in too much detail.
[0126] In this disclosure, the terms “first” and “second” are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term “multiple” means two or more, unless otherwise expressly defined.
[0127] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only.
[0128] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A device wake-up method, characterized in that, include: Listening to sound signals; In response to the detection of a human voice signal in the sound signal, an image is acquired, and the user's gaze direction and lip movement characteristics are determined based on the image; In response to the user's gaze direction being toward the device and the lip movement characteristics matching the human voice signal, the device's voice interaction module is activated.
2. The method according to claim 1, characterized in that, The method further includes: When the number of users is one, the voice interaction module is controlled to treat the user as the interaction object; When there are multiple users, the voice interaction module is controlled to select the user with the highest score as the interaction target, wherein the score is used to characterize the user's willingness to interact with the device via voice.
3. The method according to claim 2, characterized in that, The score is calculated based on the user's sound source location, the user's gaze direction, and the distance between the user and the device.
4. The method according to claim 1, characterized in that, The method further includes: If the user's gaze direction deviates from the device or the lip movement characteristics do not match the human voice signal, the audio signal continues to be monitored.
5. The method according to any one of claims 1-4, characterized in that, The user's gaze direction is towards the device, including: the angle between the user's gaze direction and the optical axis direction of the device's camera is less than an angle threshold; and / or The matching of lip movement features with the human voice signal includes: the matching degree between the user's lip opening and closing features and the energy features of the human voice signal is greater than a matching threshold.
6. The method according to any one of claims 1-4, characterized in that, The step of responding to the detection of a human voice signal in the sound signal, acquiring an image, and determining the user's gaze direction and lip movement features based on the image includes: Based on the image and the human voice signal, determine the deviation between the user's sound source location and facial orientation; In response to the deviation being less than or equal to a deviation threshold, the user's gaze direction and lip movement characteristics are determined based on the image.
7. The method according to any one of claims 1-4, characterized in that, The method further includes: In response to the detection of a human voice signal in the sound signal, at least one of the light intensity signal and the image is acquired; In response to determining that the device meets the visual judgment failure state based on at least one of the light intensity signal and the image, the sound source location of the human voice signal is obtained; In response to the sound source location of the human voice signal satisfying the sound source validity condition and the human voice signal containing speech keywords, the voice interaction module of the device is activated.
8. The method according to any one of claims 1-4, characterized in that, The method further includes: In response to the fact that the user's lip movement features are not visible in the image, the user's gaze direction and the sound source location of the human voice signal are obtained; In response to the user's gaze being directed toward the device and the sound source location of the human voice signal satisfying the sound source validity condition, the device's voice interaction module is activated.
9. The device according to any one of claims 1-4, characterized in that, The device is a smart device with walking capabilities, and the method further includes: When the device is in a walking state, in response to detecting a human voice signal in the sound signal, it continues to monitor the sound signal; or When the device is in a walking state, in response to the detection of a human voice signal in the sound signal, the device is controlled to stop walking and the device wake-up verification process is executed.
10. A device wake-up device, characterized in that, include: The monitoring module is used to monitor sound signals; The determination module is used to acquire an image in response to the detection of a human voice signal in the sound signal, and to determine the user's gaze direction and lip movement characteristics based on the image; An activation module is used to activate the device's voice interaction module in response to the user's gaze direction being toward the device and the lip movement characteristics matching the human voice signal.
11. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing at least one computer instruction, and the processor, when executing the at least one computer instruction, causes the electronic device to perform the device wake-up method according to any one of claims 1-9.
12. A storage medium storing at least one computer instruction, which, when executed by a processor of a computer device, causes the computer device to perform the device wake-up method according to any one of claims 1-9.
13. A computer program product comprising at least one computer instruction, which, when executed by a processor of a computer device, causes the computer device to perform the device wake-up method according to any one of claims 1-9.