Method, apparatus, device, and storage medium for waking up a device
By detecting the audio playback status of the device and comparing the response priority, the problem of inaccurate response of multiple devices wake-up instructions is solved, targeted device wake-up is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202210294677.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-03-22
AI Technical Summary
When there are multiple smart terminal devices that support voice wake-up function in the same space from the same manufacturer, the prior art cannot achieve targeted wake-up, resulting in a reduced accuracy of the wake-up device and inconvenience to users.
Determine the response priority by detecting the audio playback status of the device, and compare the energy values of the response priority and voice wake-up instructions with other devices, decide whether to respond to the wake-up instructions, and achieve targeted device wake-up.
Improves the accuracy of device wake-up and improves user experience.
Smart Images

Figure CN114627871B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of intelligent terminals, and in particular, to a method, apparatus, device, and storage medium for waking up a device. Background Art
[0002] With the continuous development of intelligent terminal technology and the wide application of speech recognition technology, currently many intelligent terminals (such as, intelligent TVs, intelligent speakers, and smart phones, etc.) have a voice wake-up function. In related technologies, intelligent terminals of the same manufacturer often support the same voice wake-up instruction, such as the same wake-up word, etc. When there are multiple terminal devices of the same manufacturer that support the voice wake-up function in the same space, if a user issues a voice wake-up instruction, then all devices in the current space will receive and respond to the wake-up instruction, while the user may only want to wake up the intelligent speaker at present. It can be seen that the method of waking up a device in related technologies cannot achieve targeted waking up, resulting in a decrease in the accuracy of waking up a device and bringing great inconvenience to users. Summary of the Invention
[0003] To overcome the problems existing in related technologies, embodiments of the present disclosure provide a method, apparatus, device, and storage medium for waking up a device to solve the defects in related technologies.
[0004] According to a first aspect of an embodiment of the present disclosure, a method for waking up a device is provided, the method including:
[0005] In response to receiving a voice wake-up instruction, detecting a first audio playback state of a first device;
[0006] Determining a first response priority corresponding to the first device based on the first audio playback state;
[0007] In response to receiving a second data packet sent by a second device, determining a comparison result between the first response priority and a second response priority carried in the second data packet, where the second response priority is a priority corresponding to a second audio playback state of the second device currently;
[0008] Determining whether to respond to the voice wake-up instruction based on the comparison result.
[0009] In an embodiment, the method further includes:
[0010] Sending a first data packet carrying the first response priority to the second device.
[0011] In an embodiment, the determining whether to respond to the voice wake-up instruction based on the comparison result includes:
[0012] In response to determining that the comparison result is that the first response priority is higher than the second response priority, respond to the voice wake-up instruction;
[0013] In response to determining that the comparison result is that the first response priority is lower than the second response priority, ignore the voice wake-up instruction.
[0014] In one embodiment, the method includes:
[0015] Detect a first energy value of the voice wake-up instruction;
[0016] Determining whether to respond to the voice wake-up instruction based on the comparison result includes:
[0017] In response to determining that the comparison result is that the first response priority is equal to the second response priority, determine whether to respond to the voice wake-up instruction based on a comparison result between the first energy value and a second energy value carried in the second data packet, where the second energy value is an energy value of the voice wake-up instruction detected by the second device.
[0018] In one embodiment, the method further includes:
[0019] Send a first data packet carrying the first response priority and the first energy value to the second device.
[0020] In one embodiment, detecting a current first audio playback state of the first device includes:
[0021] Dividing a reference channel audio signal extracted from the interleaved audio signal of the first device according to a set duration to obtain a plurality of the target audio signals; wherein, the length of each target audio signal is the set duration, the interleaved audio signal is a signal formed by interleaving multiple audio signals obtained from the system bottom layer, and the multiple audio signals include the reference channel audio signal received by the system bottom layer for playback and the ordinary channel audio signal collected by the system bottom layer based on a microphone;
[0022] Determine a count of playback state signals in the reference channel audio signal according to the value of each target audio signal;
[0023] Determine the current first audio playback state of the first device based on the count of the playback state signals.
[0024] In one embodiment, determining the count of the playback state signals in the reference channel audio signal according to the value of each target audio signal includes:
[0025] For each of the target audio signals, when it is detected that the value of any signal is greater than the set signal value threshold, increase the count of the playback status signal;
[0026] When it is detected that the value of no signal is greater than the set signal value threshold, decrease the count of the playback status signal.
[0027] In one embodiment, determining the current first audio playback status of the first device based on the count of the playback status signal includes:
[0028] In response to detecting that the count corresponding to a first set number of consecutive target audio signals is greater than a first set count threshold, determine that the current first audio playback status of the first device is the playing state.
[0029] In one embodiment, determining the current first audio playback status of the first device based on the count of the playback status signal includes:
[0030] In response to detecting that the count corresponding to a second set number of consecutive target audio signals is equal to a second set count threshold, determine that the current first audio playback status of the first device is the stopped playback state.
[0031] In one embodiment, the method further includes a reference channel audio signal extracted from the interleaved audio signal of the first device in the following manner:
[0032] Perform deinterleaving processing on the interleaved audio signal to obtain the multiplexed audio signals;
[0033] Extract the reference channel audio signal from the multiplexed audio signals.
[0034] In one embodiment, the method further includes:
[0035] Downsample the interleaved audio signal to obtain a downsampled interleaved audio signal;
[0036] The performing deinterleaving processing on the interleaved audio signal includes:
[0037] Perform deinterleaving processing on the downsampled interleaved audio signal.
[0038] According to a second aspect of the embodiments of the present disclosure, there is provided a device for waking up a device, the device includes:
[0039] A status detection module, configured to detect the current first audio playback status of the first device in response to receiving a voice wake-up instruction;
[0040] A level determination module, configured to determine a first response priority corresponding to the first device based on the first audio playback state;
[0041] A level comparison module, configured to determine a comparison result between the first response priority and a second response priority carried in a second data packet sent by a second device in response to receiving the second data packet, where the second response priority is a priority corresponding to a current second audio playback state of the second device;
[0042] A response determination module, configured to determine whether to respond to the voice wake-up instruction based on the comparison result.
[0043] In one embodiment, the apparatus further includes:
[0044] A first sending module, configured to send a first data packet carrying the first response priority to the second device.
[0045] In one embodiment, the response determination module includes:
[0046] A first determination unit, configured to respond to the voice wake-up instruction in response to determining that the comparison result is that the first response priority is higher than the second response priority;
[0047] A second determination unit, configured to ignore the voice wake-up instruction in response to determining that the comparison result is that the first response priority is lower than the second response priority.
[0048] In one embodiment, the apparatus includes:
[0049] An energy detection module, configured to detect a first energy value of the voice wake-up instruction;
[0050] The response determination module includes:
[0051] A third determination unit, configured to, in response to determining that the comparison result is that the first response priority is equal to the second response priority, determine whether to respond to the voice wake-up instruction based on a comparison result between the first energy value and a second energy value carried in the second data packet, where the second energy value is an energy value of the voice wake-up instruction detected by the second device.
[0052] In one embodiment, the apparatus further includes:
[0053] A second sending module, configured to send a first data packet carrying the first response priority and the first energy value to the second device.
[0054] In one embodiment, the state detection module includes:
[0055] A signal division unit, configured to divide a reference channel audio signal extracted from the interleaved audio signal of the first device according to a set duration, so as to obtain a plurality of the target audio signals; wherein the length of each target audio signal is the set duration, and the interleaved audio signal is a signal obtained from the system bottom layer and formed by interleaving multiple audio signals, and the multiple audio signals include the reference channel audio signal received by the system bottom layer for playing and the ordinary channel audio signal collected by the system bottom layer based on a microphone;
[0056] A signal counting unit, configured to determine the count of the playback status signal in the reference channel audio signal according to the value of each target audio signal;
[0057] A status determination unit, configured to determine the current first audio playback status of the first device based on the count of the playback status signal.
[0058] In one embodiment, the signal counting unit is further configured to:
[0059] For each target audio signal, when it is detected that the value of any signal is greater than a set signal value threshold, increase the count of the playback status signal;
[0060] When it is detected that the value of no signal is greater than the set signal value threshold, decrease the count of the playback status signal.
[0061] In one embodiment, the status determination unit is further configured to:
[0062] In response to detecting that the count corresponding to a continuous first set number of the target audio signals is greater than a first set count threshold, determine that the current first audio playback status of the first device is the playing status.
[0063] In one embodiment, the status determination unit is further configured to:
[0064] In response to detecting that the count corresponding to a continuous second set number of the target audio signals is equal to a second set count threshold, determine that the current first audio playback status of the first device is the stopped playback status.
[0065] In one embodiment, the device further includes a reference signal extraction module;
[0066] The reference signal extraction module includes:
[0067] A signal deinterleaving unit, configured to perform deinterleaving processing on the interleaved audio signal to obtain the multiple audio signals;
[0068] A reference signal extraction unit, configured to extract the reference channel audio signal from the multiple audio signals.
[0069] In one embodiment, the reference signal extraction module further includes:
[0070] A signal decimation unit, configured to decimate the interleaved audio signal to obtain a decimated interleaved audio signal;
[0071] The signal deinterleaving unit is further configured to perform deinterleaving processing on the decimated interleaved audio signal.
[0072] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, the device including:
[0073] A processor and a memory for storing a computer program;
[0074] Wherein, the processor is configured to, when executing the computer program, implement:
[0075] In response to receiving a voice wake-up instruction, detect a current first audio playback state of the electronic device;
[0076] Determine a first response priority corresponding to the electronic device based on the first audio playback state;
[0077] In response to receiving a second data packet sent by a second device, determine a comparison result between the first response priority and a second response priority carried in the second data packet, where the second response priority is a priority corresponding to a current second audio playback state of the second device;
[0078] Based on the comparison result, determine whether to respond to the voice wake-up instruction.
[0079] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements:
[0080] In response to receiving a voice wake-up instruction, detect a current first audio playback state of a first device;
[0081] Determine a first response priority corresponding to the first device based on the first audio playback state;
[0082] In response to receiving a second data packet sent by a second device, determine a comparison result between the first response priority and a second response priority carried in the second data packet, where the second response priority is a priority corresponding to a current second audio playback state of the second device;
[0083] Based on the comparison result, determine whether to respond to the voice wake-up instruction.
[0084] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0085] In response to receiving a voice wake-up instruction, the present disclosure detects the current first audio playback state of the first device, determines the first response priority corresponding to the first device based on the first audio playback state, and then, in response to receiving a second data packet sent by the second device, determines the comparison result between the first response priority and the second response priority carried in the second data packet, where the second response priority is the priority corresponding to the current second audio playback state of the second device. Furthermore, based on the comparison result, it is determined whether to respond to the voice wake-up instruction. Since the first response priority corresponding to the first device is determined based on the current first audio playback state of the first device, it is possible to determine whether to respond to the current voice wake-up instruction according to the comparison result between the first response priority corresponding to the first device and the second response priority corresponding to the second device, which can achieve targeted device wake-up, improve the accuracy of device wake-up, and thus enhance the user experience.
[0086] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0088] Figure 1 is a flowchart of a method for waking up a device according to an exemplary embodiment of the present disclosure;
[0089] Figure 2 is a flowchart of how to detect the current first audio playback state of the first device according to an exemplary embodiment of the present disclosure;
[0090] Figure 3A is a flowchart of how to extract a reference channel audio signal from the interleaved audio signal of the first device according to an exemplary embodiment of the present disclosure;
[0091] Figure 3B is a schematic diagram of an interleaved audio signal and a de-interleaved audio signal according to an exemplary embodiment of the present disclosure;
[0092] Figure 4 is a block diagram of a device for waking up a device according to an exemplary embodiment of the present disclosure;
[0093] Figure 5 is a block diagram of another device for waking up a device according to an exemplary embodiment of the present disclosure;
[0094] Figure 6 It is a block diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0095] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0096] Figure 1 It is a flowchart of a method for waking up a device shown according to an exemplary embodiment; the method of this embodiment can be applied to a first device supporting a voice wake-up function (such as a smart phone, a tablet computer, a smart speaker, a smart home, etc.), or can be applied to a control device of the first device, and the control device can be used to determine whether the first device responds to a voice wake-up instruction. Hereinafter, an exemplary description will be given by taking the application to the first device as an example.
[0097] As Figure 1 shown, the method includes the following steps S101-S104:
[0098] In step S101, in response to receiving a voice wake-up instruction, detect the current first audio playback state of the first device.
[0099] In this embodiment, the first device can support the voice wake-up function by installing a voice assistant program. The wake-up instruction can be a pre-set wake-up word or wake-up statement. For example, the wake-up instruction for devices of Xiaomi manufacturer is usually "Xiaomi AI Assistant". The content of the voice wake-up instruction is not limited in this embodiment.
[0100] For example, after the user issues a voice wake-up instruction, the first device can, in response to receiving the voice wake-up instruction, detect the current first audio playback state of the first device. Among them, the types of the first audio playback state can include a playing state and a stopped playback state.
[0101] In one embodiment, the first device can detect the current first audio playback state of the first device based on the audio signal collected by its own microphone, and this embodiment does not limit this.
[0102] In another embodiment, the above-mentioned manner of detecting the current first audio playback state of the first device can refer to the embodiment shown below Figure 2 shown, and details will not be described here first.
[0103] In step S102, determine the first response priority corresponding to the first device based on the first audio playback state.
[0104] In this embodiment, when responding to receiving a voice wake-up instruction and detecting the current first audio playback state of the first device, the first response priority corresponding to the first device can be determined based on the first audio playback state.
[0105] In one embodiment, based on the pre-determined correspondence between various audio playback states and response priorities, query the response priority corresponding to the above first audio playback state, and then determine it as the first response priority corresponding to the first device.
[0106] Among them, the above response priority can include preset levels such as level I and level II, where level I > level II, and level I can correspond to the "playing state", and level II can correspond to the "stopped playback state"; or, it can include preset levels of scores such as 5 points and 10 points, where 10 points > 5 points, and 10 points can correspond to the "playing state", and 5 points can correspond to the "stopped playback state".
[0107] In step S103, in response to receiving a second data packet sent by the second device, determine the comparison result between the first response priority and the second response priority carried in the second data packet, where the second response priority is the priority corresponding to the current second audio playback state of the second device.
[0108] In this embodiment, after determining the first response priority corresponding to the first device based on the first audio playback state, in response to receiving a second data packet sent by the second device, determine the comparison result between the first response priority and the second response priority carried in the second data packet, where the second response priority is the priority corresponding to the current second audio playback state of the second device.
[0109] Among them, the first device and the second device can be products of the same manufacturer, that is, both support the same wake-up instruction.
[0110] For example, when the user issues a voice wake-up command, both the first device and the second device can respond to the received voice wake-up command and detect their current audio playback states. Among them, the first audio playback state corresponds to the first device, and the second audio playback state corresponds to the second device. On this basis, the first device and the second device can send packets carrying their respective priorities to each other, that is, the first device sends a first packet carrying the first response priority to the second device, and the second device sends a second packet carrying the second response priority to the first device. Furthermore, both can determine the comparison result of the first response priority and the second response priority carried in the second packet in response to receiving the second packet sent by the other party. The second response priority is the priority corresponding to the current second audio playback state of the second device, such as: the first response priority is higher than the second response priority, the first response priority is lower than the second response priority, or the first response priority is equal to the second response priority, etc.
[0111] In step S104, based on the comparison result, it is determined whether to respond to the voice wake-up command.
[0112] In this embodiment, when, in response to receiving the second packet sent by the second device, the comparison result of the first response priority and the second response priority carried in the second packet is determined, it can be determined whether to respond to the voice wake-up command based on the comparison result.
[0113] In one embodiment, it can respond to the voice wake-up command in response to determining that the comparison result is that the first response priority is higher than the second response priority;
[0114] In response to determining that the comparison result is that the first response priority is lower than the second response priority, the voice wake-up command is ignored.
[0115] For the case where the first response priority is equal to the second response priority, this embodiment may further include: detecting a first energy value of the voice wake-up instruction. Then, after determining the first energy value, the first device may send a first data packet carrying the first response priority and the first energy value to the second device. Similarly, the second device may also send a second data packet carrying its own second response priority and a second energy value to the first device. For example, when the first device and the second device receive a voice wake-up instruction, they may respectively detect the energy value of the voice wake-up instruction. On this basis, when the first device determines that the comparison result is that the first response priority is equal to the second response priority, it may determine whether to respond to the voice wake-up instruction based on the comparison result between the first energy value and the second energy value carried in the second data packet. Exemplarily, when the first energy value is higher than the second energy value, it is determined that the first device responds to the voice wake-up instruction; on the contrary, when the first energy value is lower than the second energy value, it is determined that the second device responds to the voice wake-up instruction.
[0116] In one embodiment, the energy value of the voice wake-up instruction may be related to the distance from which the voice wake-up instruction is currently detected, that is, the closer the distance, the greater the energy value; and the farther the distance, the smaller the energy value. Exemplarily, the first terminal device may detect the energy value of the voice wake-up instruction based on the sound pickup microphones in the microphone array. The specific method may refer to the explanations and descriptions in the related art, and this embodiment does not limit it.
[0117] As can be seen from the above description, the method of this embodiment responds to receiving a voice wake-up instruction, detects the current first audio playback state of the first device, determines the first response priority corresponding to the first device based on the first audio playback state, and then responds to receiving the second data packet sent by the second device to determine the comparison result between the first response priority and the second response priority carried in the second data packet. The second response priority is the priority corresponding to the current second audio playback state of the second device. Furthermore, based on the comparison result, it is determined whether to respond to the voice wake-up instruction. Since the first response priority corresponding to the first device is determined based on the current first audio playback state of the first device, it is possible to determine whether to respond to the current voice wake-up instruction according to the comparison result between the first response priority corresponding to the first device and the second response priority corresponding to the second device, which can achieve targeted device wake-up, improve the accuracy of device wake-up, and thus improve the user experience.
[0118] Figure 2 is a flowchart showing how to detect the current first audio playback state of the first device according to an exemplary embodiment of the present disclosure; this embodiment takes how to detect the current first audio playback state of the first device as an example for exemplary illustration on the basis of the above embodiment. As Figure 2As shown, the detection of the current first audio playback state of the first device in the above step S101 may include the following steps S201-S203:
[0119] In step S201, a reference channel audio signal extracted from an interleaved audio signal of the first device is divided according to a set time length to obtain a plurality of target audio signals.
[0120] In this embodiment, in order to detect the current first audio playback state of the first device, the interleaved audio signal of the first device can be obtained first, wherein the interleaved audio signal is a signal interleaved from multiple audio signals obtained from the system bottom layer (such as the system bottom layer hardware driver such as the Advanced Linux Sound Architecture ALSA), and the multiple audio signals include the reference channel audio signal for playback received by the system bottom layer, and the ordinary channel audio signal collected by the system bottom layer based on the microphone. Then, the reference channel audio signal is extracted from the interleaved audio signal of the first device, and then the reference channel audio signal is divided according to the set duration, so as to obtain a plurality of the target audio signals, wherein the length of each target audio signal is the set duration, and the set duration can be set based on actual needs, such as being set to 8ms, etc., which is not limited in this embodiment.
[0121] For example, when the first device plays music through a player or other program, the audio data to be played can be downloaded from the local device of the first device or from the server, and then the audio signal corresponding to the audio data is sent to the system bottom hardware for playback through the system bottom hardware driver. Among them, the audio signal sent by the system bottom hardware driver to the system bottom hardware is the above-mentioned "reference channel audio signal for playback received by the system bottom layer"; and the sound actually played by the system bottom hardware is collected by the microphone of the first device (such as each microphone in the microphone array) and sent to the system bottom hardware, that is, the above-mentioned "normal channel audio signal collected by the system bottom layer based on the microphone". On this basis, the above-mentioned reference channel audio signal and normal channel audio signal will be collected and interwoven into the above-mentioned interleaved audio signal by the system bottom hardware driver.
[0122] In another embodiment, the method for extracting the reference channel audio signal from the interleaved audio signal of the first device can refer to the following Figure 3A The illustrated embodiment will not be described in detail here.
[0123] In step S202, the count of the playing status signal in the reference channel audio signal is determined according to the value of each target audio signal.
[0124] For example, for each of the target audio signals, when it is detected that the value of any signal is greater than the set signal value threshold, the count of the above playback status signal can be increased. For example, the count of the playback status signal can be incremented by one.
[0125] When it is detected that the value of no signal is greater than the set signal value threshold, the count of the above playback status signal can be decreased. For example, the count of the playback status signal can be decremented by one.
[0126] It should be noted that the signal of the above playback status can be understood as the target audio signal in which the value of any signal is greater than the set signal value threshold. When it is detected that the value of no signal is greater than the set signal value threshold in the above target audio signal, decrementing the count of the playback status signal is to avoid incrementing the count due to noise. In other words, the purpose of doing this is to avoid counting due to the presence of noise in the above target audio signal. It can be understood that usually the duration of noise is relatively short. If the count is incremented for such noise, it may lead to misjudgment. Therefore, in this embodiment, when a short-term noise is detected, although the count is incremented, when it is detected that the value of no signal is greater than the set signal value threshold in the next adjacent target audio signal, the count of the playback status signal will be decremented, thereby eliminating the influence of incrementing the count for the previous noise and improving the accuracy of the count.
[0127] In step S203, determine the current first audio playback status of the first device based on the count of the playback status signal.
[0128] For example, in response to detecting that the count corresponding to a continuous first set number of the target audio signals is greater than the first set count threshold, it can be determined that the current first audio playback status of the first device is the playing state.
[0129] Among them, the above continuous first set number can be set according to actual needs, such as set to 64, etc., and this embodiment does not limit this. For example, when continuously judging 64 audio signals of 8 ms (i.e., 512 ms), if the obtained count is greater than the first set count threshold (such as 128), it can be determined that the current first audio playback status of the first device is the playing state.
[0130] In another embodiment, when it is detected that the count corresponding to a continuous second set number of the target audio signals is equal to the second set count threshold, it can be determined that the current first audio playback status of the first device is the stopped playback state.
[0131] Among them, the above-mentioned second set number in succession can be set according to actual needs, such as set to 1000, etc., and this embodiment does not limit this. For example, after continuously judging 1000 audio signals of 8 ms, if the obtained count is equal to the second set count threshold (for example, 0), it can be determined that the current first audio playback state of the first device is the stop playback state. It can be understood that when the stop playback state appears, each of the 8 ms audio signals should be less than the set signal value threshold, and at this time the count is decremented by one; and after continuously decrementing the count multiple times, the count value will return to 0, so at this time it can be determined that the audio playback state of the first device is the stop playback state. The inventor found during the implementation of this embodiment of the present disclosure that the playback state of the first device can be determined through inter-process communication. For example, when the player of the first device plays audio, the playback process of the player can notify the process for controlling whether the device responds to the wake-up instruction (hereinafter simply referred to as the "wake-up process") of the current playback state, so that the wake-up process can know the current audio playback state of the first device. However, this solution requires communication between the playback process of the player and the wake-up process, and if communication fails, this solution cannot be implemented; in addition, for the purpose of protecting user privacy, some third-party players cannot notify the playback state of the player, so this solution cannot be implemented either. From the above description, it can be seen that in this embodiment, the reference channel audio signal extracted from the interleaved audio signal of the first device is divided to obtain a plurality of target audio signals, and the count of the playback state signal in the reference channel audio signal is determined according to the value of each target audio signal, and then the current first audio playback state of the first device is determined based on the count of the playback state signal, which can accurately determine the audio playback state of the first device, and does not rely on inter-process communication, nor on the support of external third-party players, which can improve the feasibility of the solution.
[0132] Figure 3A is a flowchart of how to extract the reference channel audio signal from the interleaved audio signal of the first device shown according to an exemplary embodiment of the present disclosure; this embodiment takes how to extract the reference channel audio signal from the interleaved audio signal of the first device as an example for exemplary illustration. As Figure 3A shown, this embodiment further includes extracting the reference channel audio signal from the interleaved audio signal of the first device based on the following steps S301 - S302:
[0133] In step S301, the interleaved audio signal is deinterleaved to obtain the multiplexed audio signals.
[0134] In this embodiment, when detecting the current first audio playback state of the first device, the interleaved audio signal of the first device may be obtained first, and then the interleaved audio signal is deinterleaved to obtain the multiplexed audio signals.
[0135] For example, if the interleaved audio signal is a signal interleaved by 6 audio signals, namely ch1, ch2, ch3, ch4, ch5, and ch6, the form of the interleaved audio signal is Figure 3B the first row signal in. From Figure 3B it can be seen that in the interleaved audio signal, the 6 audio signals of ch1, ch2, ch3, ch4, ch5, and ch6 are interleaved together, and each signal appears only once in a period. In this embodiment, the above interleaved audio signal can be deinterleaved based on the signal deinterleaving method in the related art to obtain the 6 audio signals of ch1, ch2, ch3, ch4, ch5, and ch6, as Figure 3B the signals with a length of 128 in the second row shown in. Among them, the above interleaved signal is 16Kib 16bit data.
[0136] In step S302, the reference channel audio signal is extracted from the multiplexed audio signals.
[0137] In this embodiment, after the interleaved audio signal is deinterleaved to obtain the multiplexed audio signals, the reference channel audio signal can be extracted from the multiplexed audio signals.
[0138] Still taking Figure 3B as an example, it is known that ch5 and ch6 are the reference channel audio signals. Therefore, after obtaining the 6 deinterleaved audio signals of ch1, ch2, ch3, ch4, ch5, and ch6, the audio signals of ch5 and ch6 can be extracted from these six signals.
[0139] It should be noted that which signal is the reference channel audio signal is determined according to the device information. For example, if there are 4 ordinary microphones and 2 reference channel microphones in the microphone array of the device, then the previous ch1 to ch4 are the ordinary channel audio signals collected by the microphone based on the system bottom layer, and the subsequent ch5 and ch6 are the reference channel audio signals received by the system bottom layer for playback.
[0140] It can be understood that in practical applications, there may also be only one reference signal audio signal, such as ch6, etc. The above ch5 and ch6 as the reference channel audio signals are only for illustrative purposes, and this embodiment does not limit this.
[0141] In this embodiment, before performing de-interleaving processing on the interleaved audio signal, it is possible to first determine whether the interleaved audio signal is 16 KiB 16-bit data. If not, the interleaved audio signal can be downsampled first to obtain a downsampled interleaved audio signal, that is, a 16 KiB 16-bit interleaved audio signal, and then de-interleaving processing can be performed on the downsampled interleaved audio signal. This can ensure the unity of the audio signal and thus lay a foundation for subsequent signal division and judgment.
[0142] Figure 4 It is a block diagram of a device for a wake-up device shown according to an exemplary embodiment; the device of this embodiment can be applied to a first device supporting a voice wake-up function (such as a smart phone, a tablet computer, a smart speaker, a smart home, etc.), or can be applied to a control device of the first device, and the control device can be used to determine whether the first device responds to a voice wake-up instruction.
[0143] Such as Figure 4 As shown, the device includes: a status detection module 110, a level determination module 120, a level comparison module 130, and a response judgment module 140, where:
[0144] The status detection module 110 is configured to detect the current first audio playback status of the first device in response to receiving a voice wake-up instruction;
[0145] The level determination module 120 is configured to determine the first response priority corresponding to the first device based on the first audio playback status;
[0146] The level comparison module 130 is configured to determine a comparison result between the first response priority and a second response priority carried in the second data packet in response to receiving the second data packet sent by the second device, where the second response priority is the priority corresponding to the current second audio playback status of the second device;
[0147] The response judgment module 140 is configured to determine whether to respond to the voice wake-up instruction based on the comparison result.
[0148] As described above, the device of this embodiment detects the current first audio playback state of the first device in response to receiving a voice wake-up instruction, determines the first response priority corresponding to the first device based on the first audio playback state, and then determines the comparison result between the first response priority and the second response priority carried in the second data packet in response to receiving the second data packet sent by the second device. The second response priority is the priority corresponding to the current second audio playback state of the second device. Furthermore, based on the comparison result, it is determined whether to respond to the voice wake-up instruction. Since the first response priority corresponding to the first device is determined based on the current first audio playback state of the first device, it is possible to determine whether to respond to the current voice wake-up instruction according to the comparison result between the first response priority corresponding to the first device and the second response priority corresponding to the second device, which can achieve targeted device wake-up, improve the accuracy of device wake-up, and further enhance the user experience.
[0149] Figure 5 It is a block diagram of a device for waking up a device shown according to another exemplary embodiment; the device of this embodiment can be applied to a first device that supports the voice wake-up function (such as a smart phone, a tablet computer, a smart speaker, a smart home, etc.), or can be applied to a control device of the first device, and the control device can be used to determine whether the first device responds to a voice wake-up instruction. Among them, the state detection module 210, the level determination module 220, the level comparison module 230, and the response determination module 240 have the same functions as the state detection module 110, the level determination module 120, the level comparison module 130, and the response determination module 140 in the foregoing Figure 4 illustrated embodiment, and will not be elaborated here. As Figure 5 shown, the device may further include:
[0150] A first sending module 250, configured to send a first data packet carrying the first response priority to the second device.
[0151] In one embodiment, the response determination module 240 may include:
[0152] A first determination unit 241, configured to respond to the voice wake-up instruction in response to determining that the comparison result is that the first response priority is higher than the second response priority;
[0153] A second determination unit 242, configured to ignore the voice wake-up instruction in response to determining that the comparison result is that the first response priority is lower than the second response priority.
[0154] In one embodiment, the above device may include:
[0155] An energy detection module 260, configured to detect a first energy value of the voice wake-up instruction;
[0156] A response determination module 240 may include:
[0157] A third determination unit 243, configured to, in response to determining that the comparison result is that the first response priority is equal to the second response priority, determine whether to respond to the voice wake-up instruction based on a comparison result between the first energy value and a second energy value carried in the second data packet, where the second energy value is an energy value of the voice wake-up instruction detected by the second device.
[0158] In one embodiment, the above device may further include:
[0159] A second sending module 270, configured to send a first data packet carrying the first response priority and the first energy value to the second device.
[0160] In one embodiment, a status detection module 210 may include:
[0161] A signal partitioning unit 211, configured to partition a reference channel audio signal extracted from an interleaved audio signal of the first device according to a set duration to obtain a plurality of target audio signals; where the length of each target audio signal is the set duration, the interleaved audio signal is a signal obtained from the system bottom layer and formed by interleaving multiple audio signals, and the multiple audio signals include a reference channel audio signal received by the system bottom layer for playing and a common channel audio signal collected by the system bottom layer based on a microphone;
[0162] A signal counting unit 212, configured to determine a count of a playback status signal in the reference channel audio signal according to a value of each target audio signal;
[0163] A status determination unit 213, configured to determine a current first audio playback status of the first device based on the count of the playback status signal.
[0164] In one embodiment, the signal counting unit 212 may further be configured to:
[0165] For each target audio signal, when it is detected that there is any signal value greater than a set signal value threshold, increase the count of the playback status signal;
[0166] When it is detected that there is no signal value greater than the set signal value threshold, decrease the count of the playback status signal.
[0167] In one embodiment, the status determination unit 213 may further be configured to:
[0168] In response to detecting that the count corresponding to the consecutive first set number of the target audio signals is greater than the first set count threshold, it is determined that the current first audio playback state of the first device is the playing state.
[0169] In one embodiment, the state determination unit 213 can also be used for:
[0170] In response to detecting that the count corresponding to the consecutive second set number of the target audio signals is equal to the second set count threshold, it is determined that the current first audio playback state of the first device is the stopped playback state.
[0171] In one embodiment, the above device may further include a reference signal extraction module 270;
[0172] The reference signal extraction module 270 may include:
[0173] A signal deinterleaving unit 271, configured to perform deinterleaving processing on the interleaved audio signal to obtain the multiplexed audio signals;
[0174] A reference signal extraction unit 272, configured to extract the reference channel audio signal from the multiplexed audio signals.
[0175] In one embodiment, the above reference signal extraction module 270 may further include:
[0176] A signal downsampling unit 273, configured to perform downsampling on the interleaved audio signal to obtain the downsampled interleaved audio signal;
[0177] The signal deinterleaving unit 271 may also be used to perform deinterleaving processing on the downsampled interleaved audio signal.
[0178] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0179] Figure 6 It is a block diagram of an electronic device shown according to an exemplary embodiment. For example, the device 900 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0180] Referring to Figure 6 , the device 900 may include one or more of the following components: a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.
[0181] The processing component 902 generally controls the overall operation of the device 900, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above-described methods. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.
[0182] The memory 904 is configured to store various types of data to support the operation of the device 900. Examples of such data include instructions for any application or method operating on the device 900, contact data, phone book data, messages, pictures, videos, and the like. The memory 904 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0183] The power component 906 provides power to various components of the device 900. The power component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 900.
[0184] The multimedia component 908 includes a screen that provides an output interface between the device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0185] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC) that is configured to receive external audio signals when the device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.
[0186] The I / O interface 912 provides an interface between the processing component 902 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0187] The sensor component 914 includes one or more sensors for providing status assessments of various aspects of the device 900. For example, the sensor component 914 can detect the on / off state of the device 900, the relative positioning of components, such as the display and keypad of the device 900. The sensor component 914 can also detect a change in the position of the device 900 or a component of the device 900, the presence or absence of user contact with the device 900, the orientation or acceleration / deceleration of the device 900, and the temperature change of the device 900. The sensor component 914 can also include a proximity sensor that is configured to detect the presence of nearby objects without any physical contact. The sensor component 914 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 914 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0188] The communication component 916 is configured to facilitate communication between the device 900 and other devices in a wired or wireless manner. The device 900 can access a wireless network based on communication standards, such as WiFi, 2G or 3G, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0189] In an exemplary embodiment, the device 900 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0190] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 904 including instructions, may be provided, and the above instructions may be executed by the processor 920 of the device 900 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0191] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the disclosure herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed herein. The specification and embodiments are to be considered as exemplary only, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0192] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for waking up a device, characterized in that, The method includes: In response to receiving a voice wake-up instruction, detecting a first audio playback state of a first device; Determining a first response priority corresponding to the first device based on the first audio playback state; In response to receiving a second data packet sent by a second device, determining a comparison result between the first response priority and a second response priority carried in the second data packet, where the second response priority is a priority corresponding to a second audio playback state of the second device currently; Determining whether to respond to the voice wake-up instruction based on the comparison result; The detecting the first audio playback state of the first device currently includes: Dividing a reference channel audio signal extracted from an interleaved audio signal of the first device according to a set duration to obtain a plurality of target audio signals; where the length of each target audio signal is the set duration, and the interleaved audio signal is a signal formed by interleaving multiple audio signals obtained from the system bottom layer, and the multiple audio signals include the reference channel audio signal received by the system bottom layer for playback and the ordinary channel audio signal collected by the system bottom layer based on a microphone; Determining a count of playback state signals in the reference channel audio signal according to the value of each target audio signal; Determining the first audio playback state of the first device currently based on the count of the playback state signals.
2. The method according to claim 1, characterized in that, The method further includes: Sending a first data packet carrying the first response priority to the second device.
3. The method according to claim 1, wherein The determining whether to respond to the voice wake-up instruction based on the comparison result includes: In response to determining that the comparison result is that the first response priority is higher than the second response priority, responding to the voice wake-up instruction; In response to determining that the comparison result is that the first response priority is lower than the second response priority, ignoring the voice wake-up instruction.
4. The method according to claim 1, wherein The method includes: Detecting a first energy value of the voice wake-up instruction; The determining whether to respond to the voice wake-up instruction based on the comparison result includes: In response to determining that the comparison result is that the first response priority is equal to the second response priority, determining whether to respond to the voice wake-up instruction based on a comparison result between the first energy value and a second energy value carried in the second data packet, where the second energy value is an energy value of the voice wake-up instruction detected by the second device.
5. The method according to claim 4, wherein The method further includes: Sending a first data packet carrying the first response priority and the first energy value to the second device.
6. The method according to claim 1, wherein The determining the count of playback state signals in the reference channel audio signal according to the value of each target audio signal includes: For each target audio signal, when it is detected that there is any signal value greater than a set signal value threshold, increasing the count of the playback state signals; When it is detected that there is no signal value greater than the set signal value threshold, decreasing the count of the playback state signals.
7. The method according to claim 1, wherein The determining the first audio playback state of the first device currently based on the count of the playback state signals includes: In response to detecting that the count corresponding to the first set number of consecutive target audio signals is greater than the first set count threshold, it is determined that the current first audio playback state of the first device is the playing state.
8. The method according to claim 1, characterized in that, The determining of the current first audio playback state of the first device based on the count of the playback state signal includes: In response to detecting that the count corresponding to the second set number of consecutive target audio signals is equal to the second set count threshold, it is determined that the current first audio playback state of the first device is the stopped playback state.
9. The method according to claim 1, characterized in that, The method further includes a reference channel audio signal extracted from the interleaved audio signal of the first device in the following manner: Perform deinterleaving processing on the interleaved audio signal to obtain the multiplex audio signals; Extract the reference channel audio signal from the multiplex audio signals.
10. The method according to claim 9, wherein The method further includes: Perform downsampling on the interleaved audio signal to obtain the downsampled interleaved audio signal; The performing of the deinterleaving processing on the interleaved audio signal includes: Perform deinterleaving processing on the downsampled interleaved audio signal.
11. A device for waking up a device, characterized in that, The apparatus includes: A state detection module, configured to detect the current first audio playback state of the first device in response to receiving a voice wake-up instruction; A level determination module, configured to determine the first response priority corresponding to the first device based on the first audio playback state; A level comparison module, configured to determine the comparison result between the first response priority and the second response priority carried in the second data packet in response to receiving the second data packet sent by the second device, where the second response priority is the priority corresponding to the current second audio playback state of the second device; A response judgment module, configured to determine whether to respond to the voice wake-up instruction based on the comparison result; The state detection module includes: A signal division unit, configured to divide the reference channel audio signal extracted from the interleaved audio signal of the first device according to a set duration to obtain a plurality of target audio signals; wherein, the length of each target audio signal is the set duration, the interleaved audio signal is a signal obtained from the system bottom layer and formed by interleaving multiplex audio signals, and the multiplex audio signals include the reference channel audio signal received by the system bottom layer for playback and the ordinary channel audio signal collected by the system bottom layer based on a microphone; A signal counting unit, configured to determine the count of the playback state signal in the reference channel audio signal according to the value of each target audio signal; A state determination unit, configured to determine the current first audio playback state of the first device based on the count of the playback state signal.
12. An electronic device, characterized in that, The device includes: A processor and a memory for storing a computer program; Wherein, the processor is configured to, when executing the computer program, implement: Detect the current first audio playback state of the electronic device in response to receiving a voice wake-up instruction; Determine the first response priority corresponding to the electronic device based on the first audio playback state; In response to receiving a second data packet sent by a second device, determine the comparison result between the first response priority and the second response priority carried in the second data packet, where the second response priority is the priority corresponding to the current second audio playback state of the second device; Determine whether to respond to the voice wake-up instruction based on the comparison result; The detecting the current first audio playback state of the first device includes: Dividing the reference channel audio signal extracted from the interleaved audio signal of the first device according to a set duration to obtain a plurality of target audio signals; wherein the length of each target audio signal is the set duration, and the interleaved audio signal is a signal formed by interleaving multiple audio signals obtained from the system bottom layer, and the multiple audio signals include the reference channel audio signal received by the system bottom layer for playback and the common channel audio signal collected by the system bottom layer based on the microphone; Determine the count of the playback state signals in the reference channel audio signal according to the value of each target audio signal; Determine the current first audio playback state of the first device based on the count of the playback state signals.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it realizes: In response to receiving a voice wake-up instruction, detect the current first audio playback state of the first device; Determine the first response priority corresponding to the first device based on the first audio playback state; In response to receiving a second data packet sent by a second device, determine the comparison result between the first response priority and the second response priority carried in the second data packet, where the second response priority is the priority corresponding to the current second audio playback state of the second device; Determine whether to respond to the voice wake-up instruction based on the comparison result; The detecting the current first audio playback state of the first device includes: Dividing the reference channel audio signal extracted from the interleaved audio signal of the first device according to a set duration to obtain a plurality of target audio signals; wherein the length of each target audio signal is the set duration, and the interleaved audio signal is a signal formed by interleaving multiple audio signals obtained from the system bottom layer, and the multiple audio signals include the reference channel audio signal received by the system bottom layer for playback and the common channel audio signal collected by the system bottom layer based on the microphone; Determine the count of the playback state signals in the reference channel audio signal according to the value of each target audio signal; Determine the current first audio playback state of the first device based on the count of the playback state signals.
Citation Information
Patent Citations
Voice awakening method and device
CN111276139A
Method and device for waking up equipment, electronic equipment and storage medium
CN112201242A