Methods, apparatuses, devices, media, and products for waking up a device

By evaluating user audio input in real time within the wearable device and transmitting it to a second device in advance for secondary verification, the problem of user experience latency caused by low Bluetooth transmission efficiency is solved, enabling faster device wake-up and more efficient human-computer interaction.

CN122195515APending Publication Date: 2026-06-12BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411823578.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In wearable devices, traditional wake-up solutions suffer from user experience latency and low interaction efficiency due to the low transmission efficiency of Bluetooth.

Method used

By evaluating user audio input in real time on the first device, a portion of the audio input is sent to the second device in advance for secondary verification, reducing audio transmission latency and using pre-wake events to advance the start time of audio transmission.

Benefits of technology

It shortens the device wake-up time, improves the user experience, and enhances the smoothness and responsiveness of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195515A_ABST
    Figure CN122195515A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to methods, apparatuses, devices, media and products for waking up a device. The method includes determining, at a first device, a received audio input from a user. The method also includes determining a target evaluation result that the received audio input represents a portion of a wake-up command of the first device. The method further includes providing, in response to the target evaluation result being greater than a threshold result, the received audio input and a subsequent audio input corresponding to another portion of the wake-up command to a second device for determining whether to wake up the first device. The method also includes waking up the first device in response to receiving an indication from the second device to wake up the first device. Through the method, the time of audio sending is controlled, the sending process is advanced, the delay of the link is optimized, the interaction delay perceived by the user is effectively reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure generally relate to the field of device management, and more specifically to methods, apparatus, devices, media, and products for waking up devices and processing messages. Background Technology

[0002] Currently, voice interaction technology is increasingly prevalent in smart devices, especially wearable devices such as Bluetooth headsets or smart speakers. For users, the smoothness of human-computer interaction is paramount. Therefore, improving the communication efficiency between users and smart devices in human-computer interaction scenarios has become a key research focus for many developers. Developers are committed to continuously enhancing the user experience in human-computer interaction by improving the efficiency of interaction between smart devices and users.

[0003] With the continuous development of wearable devices, reducing the wake-up speed of wearable devices is often the main research method in human-computer interaction. For users, wearable devices can respond quickly, and the time users wait to wake up the device is shorter, which can significantly enhance the overall user experience in the human-computer interaction process. Summary of the Invention

[0004] Embodiments of this disclosure provide methods, apparatus, devices, media, and products for waking up a device.

[0005] According to a first aspect of this disclosure, a method for waking up a device is provided. The method includes determining, at a first device, received audio input from a user. The method further includes determining, in response to the received audio input representing a portion of a wake-up command for the first device, the target evaluation result representing a portion of the wake-up command. The method further includes, in response to the target evaluation result being greater than the threshold result, providing the received audio input and subsequent audio input representing another portion of the corresponding wake-up command to a second device for determining whether to wake up the first device. The method further includes, in response to receiving an instruction from the second device to wake up the first device, waking up the first device.

[0006] According to a second aspect of this disclosure, an apparatus for waking up a device is provided. The apparatus includes an audio input receiving module configured to determine, at a first device, received audio input from a user; a wake-up command evaluation module configured to determine a target evaluation result indicating that the received audio input represents a portion of a wake-up command for the first device; an audio input forwarding module configured to, in response to a target evaluation result greater than a threshold result, provide the received audio input and subsequent audio input corresponding to another portion of the wake-up command to a second device for determining whether to wake up the first device; and a device wake-up module configured to wake up the first device in response to receiving an instruction from the second device to wake up the first device.

[0007] In a third aspect of this disclosure, an electronic device is provided, including at least one processor; and a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to the first aspect of this disclosure.

[0008] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0009] In a fifth aspect of this disclosure, a computer program product is provided. This computer program product includes a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0010] It should be understood that the content described in this section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0012] Figure 1 The illustration shows a schematic diagram of an example environment in which some embodiments of the present disclosure may be implemented;

[0013] Figure 2 The illustration shows a schematic diagram of an example method for waking up a device according to some embodiments of the present disclosure;

[0014] Figure 3 The illustration shows a schematic diagram of an example process of device and application interaction flow according to some embodiments of the present disclosure;

[0015] Figure 4 The illustration shows a schematic block diagram of an apparatus for waking up a device according to some embodiments of the present disclosure;

[0016] Figure 5 A schematic block diagram of an example device suitable for implementing various embodiments of the present disclosure is illustrated. Detailed Implementation

[0017] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0019] For example, upon receiving a user's proactive request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0022] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0023] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0024] Typically, in human-computer interaction, the device is activated and interaction is triggered by the user actively speaking a wake word. However, Bluetooth headsets, limited by computing resources, can only have a small built-in wake-up model. To ensure a high wake-up rate and a low false wake-up rate, a larger secondary model needs to be placed within the mobile application for secondary verification. Therefore, after the wake-up event is detected on the headset, the audio is transmitted to the application via Bluetooth. After the application processes the audio, it sends the conclusion of whether the wake-up was successful back to the headset. The headset can play a notification tone to let the user know that the wake-up was successful. However, due to the limitations of Bluetooth transmission efficiency, there is a significant delay in transmitting the audio to the mobile verification process. For example, due to Bluetooth bandwidth limitations, transmitting an average of 1 second of 4-syllable wake word audio to the application via Bluetooth takes approximately 180ms. This increases the user's perceived latency (the time difference between saying the wake word and perceiving the successful wake-up is defined as the perceived latency). This reduces the efficiency of intelligent interaction and affects the user experience.

[0025] Therefore, embodiments of this disclosure propose a method for waking up a device. In this method, a first device receives audio input from a user. Then, the first device determines a target evaluation result indicating that the received audio input from the user represents a part of a wake-up command for the first device. If the target evaluation result is greater than a threshold result, indicating that the received audio input is likely part of a wake-up command, the received audio input and subsequent audio input corresponding to another part of the wake-up command can be provided to a second device to determine whether to wake up the first device. If an instruction to wake up the first device is received from the second device, the second device is woken up. This method, by sending audio to the mobile device in advance by pre-wake-up time, effectively advances the start time of audio transmission to the mobile device, allowing the mobile device to receive all audio faster, reducing the user's waiting time to wake up the device, reducing perceived stuttering, and improving the user experience.

[0026] The embodiments of this disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 Example environments in which the devices and / or methods of embodiments of this disclosure may be implemented are shown. In environment 100, the first device 110 may provide the received audio input corresponding to the wake-up command to the second device 114 for verification as early as possible by detecting the pre-wake-up of audio.

[0027] The first device 110 can be a wearable device, such as a Bluetooth headset, smartwatch, smart bracelet, smart glasses, smart clothing, etc. Examples of the second device 114 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments that include any of the above systems or devices.

[0028] like Figure 1As described above, when the user needs to wake up the first device 110, generally, the wake-up command required by the corresponding device needs to be spoken out, such as wake-up words or sentences, that is, the corresponding user audio input 102. For example, when the user needs to wake up the first device, if the wake-up command is the four words "wake up the device", then during the process of the user speaking out these four words one by one, it can be divided into two parts, namely the part that has been spoken and the part to be spoken. The part that has been spoken is received by the first device 110 (such as a Bluetooth headset or smart glasses, etc.) in real time and serves as the received audio input 106. The received audio input 106 is evaluated in the first device 110. Therefore, in the first device 110, a target evaluation result 108 can be obtained by taking the received audio input 106 as a part of the pre-wake-up instruction of the first device 110. Additionally, the target evaluation result here is implemented by an audio processing model in the first device. For example, on the premise that the wake-up word is "wake up the device", a pre-wake-up event can be triggered in advance when the user has not finished speaking the wake-up word. The traditional solution is to transmit the audio after determining that the user has finished speaking the complete wake-up command. At this time, to determine that the audio input is a wake-up command, it needs to be compared with a relatively high wake-up threshold. For example, the wake-up threshold is 0.9. However, in this disclosure, the wake-up threshold for the wake-up command is set to a value smaller than the traditional wake-up threshold. For example, the wake-up threshold can be set to 0.5. Therefore, when detecting a partial voice corresponding to the wake-up command, it can be determined that the voice input is a wake-up command. For example, when the first device 110 receives the audio input of the word "wake", the score of the target evaluation result 108 calculated can be 0.2. At this time, the score of the target evaluation result has not reached the threshold score, so the first device will continue to receive the subsequent audio input from the user. When the first device 110 continues to receive the audio input of the word "up", assuming that the score of the target evaluation result 108 has accumulated to 0.5 at this time, reaching the threshold score, this means that the device confirms that the pre-wake-up condition is established and can perform the next action. If the second word received by the first device 110 is not "up", for example, it is an irrelevant other word, the device will determine that the current audio input does not meet the requirements of the pre-wake-up instruction. At this time, the target evaluation score directly returns to 0 from the original 0.2, and then continues to wait for and process new audio inputs, and perform target evaluation on the subsequent audio inputs until a voice meeting the wake-up condition is detected.

[0029] In some embodiments, target evaluation of received audio is performed using an audio processing model built into the first device 110. The core of this target evaluation is the analysis of audio frames. For example, when audio input is received by the first device, it divides the continuous audio signal into multiple small time segments, i.e., audio frames. Each audio frame typically contains audio data of a fixed length (e.g., 20 milliseconds or 50 milliseconds). This frame-segmentation process helps improve real-time performance and processing efficiency. The audio processing model built into the first device then extracts speech feature parameters from each frame, such as Mel-frequency cepstral coefficients, energy, and frequency characteristics. These feature parameters can be used to represent the core information of the speech. The first device processes the extracted audio features to obtain the target evaluation result. For example, the audio processing model can obtain an evaluation result for the current audio frame based on its audio features and the scores of previously processed frames; this result can also serve as the target evaluation result for the received audio input. When the target evaluation result is represented by a score, it can be a cumulative score for the received audio input. When this score reaches a set threshold, the first device confirms that a pre-wake-up event has been triggered. If the accumulated score does not reach the threshold, the first device will receive new audio input.

[0030] The first device 110 compares the target evaluation result 108 with the threshold result 112. If the target evaluation result 108 is greater than the threshold result 112, it indicates that the audio input now being received is an audio input for a wake-up command.

[0031] Therefore, both the received audio input 106 and the subsequent audio input 104, i.e., all the user audio inputs 102, can be provided to the second device 114 to determine whether to wake up the first device 110. Specifically, when the target evaluation result 108 is greater than the threshold result 112, the first device 110 begins sending the received audio input 106 to the second device 114. After all the received audio inputs 106 have been sent, it continues sending the received subsequent audio inputs 104 until all the user audio inputs 102 have been sent to the second device 114. Then, the second device 114 performs secondary verification processing on the received user audio inputs 102 to determine whether all the received user audio inputs 102 constitute the wake-up instruction 116. When the second device 114 receives subsequent audio inputs 104, the subsequent audio inputs 104 also need to be transmitted to the first device 110 first, and then the first device 110 transmits the subsequent audio inputs 104 to the second device 114.

[0032] Finally, if the first device 110 receives an instruction from the second device 114 to wake up the first device 110, the first device 110 is woken up to perform subsequent data processing. Therefore, after a successful wake-up, the first device 110 can respond to the user's audio.

[0033] This method utilizes a preset pre-wake event to begin transmitting audio to the second device before the user has fully uttered the wake word. This advances the start time of audio transmission and the end time of audio transmission to the second device, allowing the first device to respond to the user's wake word more quickly. This significantly shortens the response waiting time of the first device and improves the user experience.

[0034] The above combination Figure 1 The following is a schematic diagram illustrating an example environment in which some embodiments of this disclosure may be implemented, in conjunction with... Figure 2 A schematic diagram illustrating an example method for waking up a device according to some embodiments of the present disclosure. Figure 2 The method in can be derived from Figure 1 The first device 110 or any suitable computing device in the process shall execute.

[0035] like Figure 2 In example method 200, at block 202, received audio input from a user is determined at the first device. For example, the first device 110 receives user audio input from the user. At this time, the received audio input does not correspond to a complete wake-up command, but rather to a portion of the wake-up command. Figure 1 As shown, the user's audio input at this time is the received audio input 106, which is used to determine that the subsequent audio input 104 of the wake-up command has not yet been received. The received audio input 106 and the subsequent audio input 104 are combined in sequence to form all the audio inputs for the user's wake-up command.

[0036] At box 204, the target evaluation result of the received audio input representing a portion of the wake-up command of the first device is determined. For example, the first device 110 processes the audio frame of the received audio input 106, at which point the user's audio input is not yet fully completed, and is therefore considered incomplete audio input. The received audio portion is the voice content already spoken by the user, and has been successfully received by the first device. The first device 110 performs real-time analysis and processing on this portion of audio to evaluate whether it meets certain preset conditions (e.g., whether it meets some features of the wake-up word). The subsequent audio input 104 to be received is the voice content not yet spoken by the user. The first device 110 continues to listen to and receive this audio to complete the entire voice input processing flow. These two audio portions are combined in chronological order to ultimately form the user's complete audio input. This step-by-step reception and processing design enables dynamic response to user input during voice interaction and improves the device's real-time processing capabilities. Moreover, the device follows the user's input in real time, without reducing the smoothness of the interaction due to long waiting times, thus improving the user experience.

[0037] In some embodiments, when determining that the received audio input represents a target evaluation result that is part of a wake-up command for the first device, the first device 110 may process each received audio frame. For example, if the first device 110 receives the last audio frame in the received audio input, the first device 110 may use the last audio frame and historical evaluation results of previous audio frames in the received audio input to determine the target evaluation result for the received audio input at this time. For example, if the first device has an audio processing model, and the last audio frame in the received audio input is the currently received audio frame, it can calculate the target evaluation result for the received audio input after receiving the current audio frame based on the current audio frame and the evaluation results for previous frames.

[0038] Then, at box 206, in response to the target evaluation result being greater than the threshold result, the received audio input and the subsequent audio input corresponding to another part of the wake-up command are provided to the second device to determine whether to wake up the first device. To determine whether the received audio input 106 is part of the wake-up command, a threshold result 112 is set to judge the received audio input 106. The threshold result 112 is a preset value in the first device and can be adjusted and optimized through later software updates to the first device. The received audio input 106 and the subsequent audio input 104 together constitute the total audio input by the user. For example, since the first device is often a wearable device supporting voice interaction, such as a Bluetooth headset, its size is not large, and its computing resources are limited. Therefore, the first device 110 only has a small wake-up model. To ensure a high wake-up rate and a low false wake-up rate, the audio needs to be transmitted to the second device for secondary verification.

[0039] In some embodiments, if the target evaluation result is greater than the threshold result, it indicates that the received audio input has a high probability of being related to a wake-up command. Therefore, it is not necessary to wait until all audio input related to the wake-up command is received before transmitting the received audio input to the second device 114. Next, the first device 110 can continue to receive subsequent audio input from the user corresponding to another part of the wake-up command. Then, the first device 110 also transmits the received subsequent audio input to the second device 114. For ease of description, the subsequent audio input described above can also be referred to as the first subsequent audio input. If the target evaluation result is less than or equal to the threshold result, it indicates that the received audio input cannot be determined to be part of the wake-up command, and therefore it is necessary to continue receiving the second subsequent audio input. Then, the first device 110 can further determine the evaluation results for the received audio input and the second subsequent audio input based on the received audio input and the second subsequent audio input.

[0040] In some embodiments, the received audio input 106 and subsequent audio input 104 are also input to the second device in the order of the user's audio input. For example, the second device 114 can be a smartphone, in which case all audio is transmitted to the smartphone and will be processed by the corresponding application in the second device.

[0041] In some embodiments, when the second device 114 performs secondary verification on all audio sent by the first device 110, it compares and verifies the received audio and the wake word based on a machine learning model. This machine learning model can be a pre-trained neural network model, such as a convolutional neural network model or a recurrent neural network model. In some embodiments, the second device can determine the verification result of all audio based on a predetermined mapping relationship. The above examples are merely for describing this disclosure and are not intended to specifically limit this disclosure.

[0042] At block 208, the first device is woken up in response to receiving an instruction from the second device to wake up the first device. For example, if the first device 110 receives an instruction from the second device 114 to wake up the first device 110, it can perform the operation of waking up the first device 110.

[0043] When the second device 114 receives voice input from the user, it further detects whether the user intends to wake up the first device. For example, the user speaks a voice message containing a wake-up command while near the first device 110 (such as a Bluetooth headset or smart glasses). The first device forwards the received user audio data or command to the second device and records the status of the wake-up operation. This operation ensures that the first device can independently process the user's wake-up intention while retaining the flexibility of multi-device collaboration. After receiving the user's wake-up command, the second device further verifies the audio content to ensure the accuracy and legitimacy of the command. This secondary verification typically includes voice matching and background noise elimination. Voice matching analyzes whether keywords in the audio match the first device's preset wake-up word, and background noise elimination filters out interference noise in the environment, improving the accuracy of the verification. If the second device confirms that the audio command passes verification, it returns the verification result to the first device, along with relevant information (such as a wake-up success status indicator, user intent, etc.). This step ensures that the first device can respond based on accurate verification information. After receiving the verification result from the second device, the first device executes the wake-up operation, switches to interactive mode, and prepares to receive further commands from the user. It also sends a preset wake-up response to the user (such as a voice prompt "Device is awake" or an indicator light) to confirm a successful wake-up operation. In scenarios involving multiple devices working together, this mechanism ensures the efficiency and accuracy of the wake-up process while avoiding conflicts between devices.

[0044] This method allows the audio to be sent to the second device in advance, shortening the time it takes for the second device to receive all the audio. This enables the second device to return the results of the secondary verification to the first device more quickly, speeding up the first device's response to the user's wake-up call and improving the user experience.

[0045] The above combination Figure 2A schematic diagram illustrating an example method for waking up a device according to some disclosed embodiments is shown below. Figure 3 The illustration shows a schematic diagram of an example process of device and application interaction flow according to some embodiments of the present disclosure. Figure 3 In the example process, the headphones can be used as Figure 1 The first device in the process is the application (APP), which can be an application running on the second device.

[0046] Example 300 illustrates an example of the interaction process between the headphones and the app. t0, t1, t2, and t3, and T1, T2, and T3 represent the time points in the diagram where the headphones interact with the application. In box 302, the user says a wake-up word, and the headphones receive this word. This receiving phase is divided into two parts: the user saying part of the wake-up word and the user saying the entire wake-up word, corresponding to times t0 and t1 on the headphones, respectively. At t0, the headphones give a pre-wake-up event and begin sending audio for the wake-up word to the app. After a short delay in Bluetooth transmission, the app on the second device begins receiving the audio sent at t0 at time T1.

[0047] At point t1, after the user says the entire wake-up word, the earphones have already sent a portion of the audio to the app. The earphones will continue sending the audio from point t0 to t1 to the app. At point t2, the earphones have sent the user's entire audio message via Bluetooth. After a certain transmission time, the app on the second device receives all the audio at point t2. The time interval from t2 to t3 represents the app on the second device performing a secondary verification of the incoming audio. At point t3, the verification process is complete, and the result is sent to the earphones. After a certain transmission time, the earphones receive the wake-up result at point t3. If the verification confirms that the received voice corresponds to the wake-up word, it responds to the user's wake-up request, indicating successful wake-up. If the verification confirms that the received voice does not correspond to the wake-up word, the earphones are not activated.

[0048] Figure 4 The illustration shows a schematic block diagram of a wake-up device according to some embodiments of the present disclosure. Figure 4 As shown, the device 400 includes a candidate audio input receiving module 402 configured to determine received audio input from a user at a first device; a wake-up command evaluation module 404 configured to determine a target evaluation result that the received audio input represents a portion of a wake-up command for the first device; an audio input forwarding module 406 configured to provide the received audio input and subsequent audio input corresponding to another portion of the wake-up command to a second device in response to the target evaluation result being greater than a threshold result, for determining whether to wake up the first device; and a device wake-up module 408 configured to wake up the first device in response to receiving an instruction from the second device for waking up the first device.

[0049] In some embodiments, the first device is a wearable device and the second device is a terminal device.

[0050] In some embodiments, the wake-up command evaluation module includes: a last audio frame identification module configured to determine the last audio frame in the received audio input; and a target evaluation result calculation module configured to determine a target evaluation result based on the last audio frame and historical evaluation results for previous audio frames in the received audio input.

[0051] In some embodiments, the wake-up command evaluation module includes an audio processing and evaluation module configured to determine a target evaluation result by applying the received audio input to an audio processing model.

[0052] In some embodiments, the audio input forwarding module includes: an audio input transmission module configured to transmit received audio input to the second device in response to a target evaluation result being greater than a threshold result; a first subsequent audio receiving module configured to receive another portion of subsequent audio input from a user corresponding to a wake-up command; and a subsequent audio transmission module configured to transmit the subsequent audio input to the second device.

[0053] In some embodiments, the audio input forwarding module includes: a second subsequent audio receiving module configured to continue receiving a second subsequent audio input in response to a target evaluation result being less than or equal to a threshold result; and an audio evaluation result calculation module configured to determine an evaluation result for the received audio input and the second subsequent audio input based on the received audio input and the second subsequent audio input.

[0054] In some embodiments, the second device determines whether to wake up the first device based on received audio input and subsequent audio input.

[0055] In some embodiments, the audio input forwarding module further includes a Bluetooth transmission module configured to provide received audio input and subsequent audio input to a second device via Bluetooth communication.

[0056] In some embodiments, the second device utilizes a machine learning model to process received audio input and subsequent audio input.

[0057] In some embodiments, the wake-up command is a statement or a word.

[0058] Figure 5 A schematic block diagram of an example device 500 that can be used to implement embodiments of the present disclosure is shown. Figure 1The first device 110 and the second device 114 can be implemented using device 500. As shown, device 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 502 or loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 can also store various programs and data required for the operation of device 500. CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0059] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0060] The various processes and procedures described above, such as method 200 and process 300, can be executed by processing unit 501. For example, in some embodiments, method 200 and process 300 can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by CPU 501, one or more actions of the example methods 200 and process 300 described above can be performed.

[0061] This disclosure can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0062] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0063] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0064] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0065] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0066] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0067] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0069] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for waking up a device, comprising: The first device identifies the received audio input from the user; The received audio input represents a portion of the wake-up command of the first device; In response to the target evaluation result being greater than the threshold result, the received audio input and subsequent audio input corresponding to another part of the wake-up command are provided to the second device to determine whether to wake up the first device; as well as In response to receiving an instruction from the second device to wake up the first device, the first device is woken up.

2. The method according to claim 1, wherein the first device is a wearable device and the second device is a terminal device.

3. The method of claim 1, wherein determining the target evaluation result that represents a portion of the wake-up command of the first device by the received audio input comprises: Determine the last audio frame in the received audio input; as well as The target evaluation result is determined based on the last audio frame and the historical evaluation results for previous audio frames in the received audio input.

4. The method of claim 1, wherein determining the target evaluation result that the received audio input represents a portion of a wake-up command for the first device comprises: The target evaluation result is determined by applying the received audio input to the audio processing model.

5. The method of claim 1, wherein providing the received audio input and a subsequent audio input corresponding to another portion of the wake-up command to the second device comprises: In response to the target evaluation result being greater than the threshold result, the received audio input is transmitted to the second device; Receive the subsequent audio input from the user corresponding to another part of the wake-up command; as well as The subsequent audio input is transmitted to the second device.

6. The method of claim 1, wherein the subsequent audio input is a first subsequent audio input, and the method further comprises: In response to the target evaluation result being less than or equal to the threshold result, continue to receive a second subsequent audio input; as well as Based on the received audio input and the second subsequent audio input, an evaluation result is determined for the received audio input and the second subsequent audio input.

7. The method of claim 1, wherein the second device determines whether to wake up the first device based on the received audio input and the subsequent audio input.

8. The method of claim 7, wherein the second device processes the received audio input and the subsequent audio input using a machine learning model.

9. The method of claim 1, wherein providing the received audio input and a subsequent audio input corresponding to another portion of the wake-up command to the second device comprises: The received audio input and subsequent audio input are provided to the second device via Bluetooth communication.

10. The method of claim 1, wherein the wake-up command is a statement or a word.

11. An apparatus for waking up a device, comprising: An audio input receiving module is configured to determine, at a first device, the received audio input from the user; A wake-up command evaluation module is configured to determine a target evaluation result that represents a portion of the wake-up command of the first device from the received audio input. An audio input forwarding module is configured to, in response to the target evaluation result being greater than a threshold result, provide the received audio input and subsequent audio input corresponding to another part of the wake-up command to a second device for determining whether to wake up the first device; The device wake-up module is configured to wake up the first device in response to receiving an instruction from the second device to wake up the first device.

12. An electronic device, comprising: At least one processor; as well as A storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, the computer program implementing the method according to any one of claims 1-10 when executed by a processor.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.