Interaction method of intelligent sound box and intelligent sound box
By using the difference in microphone signal strength of smart speakers to determine wake-up and control noise in linkage devices, the problem of low wake-up success rate in high-noise environments is solved, and more efficient voice control is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU ROBAM APPLIANCES CO LTD
- Filing Date
- 2026-03-26
- Publication Date
- 2026-04-21
AI Technical Summary
In high-noise environments, smart speakers have a low wake-up success rate and cannot effectively receive user voice commands, leading to the failure of voice-controlled linked devices.
The smart speaker is equipped with at least two microphones. By detecting the difference in microphone signal strength, it can determine that one of the microphones is blocked, wake up the smart speaker, and control linked devices to reduce noise and improve the success rate of voice signal reception in high-noise environments via wireless communication.
It improves the wake-up success rate of smart speakers in high-noise environments and the success rate of voice-controlled linkage devices, reduces noise interference, and enhances the reliability of voice interaction.
Smart Images

Figure CN121905178A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart home technology, and in particular to an interaction method for a smart speaker and a smart speaker. Background Technology
[0002] Currently, smart speakers can be used in kitchen settings, allowing users to control connected devices such as range hoods and steam ovens by speaking specific commands to the smart speaker.
[0003] The recognition of voice control commands by smart speakers mainly consists of three parts: wake-up, local recognition, and cloud recognition. Among them, the user's wake-up of the smart speaker is the first step to realize voice control of linked devices. By saying a specific wake-up word to the smart speaker, the user can wake up the smart speaker from the standby state, and then conduct subsequent voice dialogues, such as turning off the range hood, turning on the steam oven, setting the temperature of the steam oven, and so on.
[0004] However, with current voice signal processing technology, when the range hood is operating at the high-noise setting, the smart speaker is in a high-noise environment, which often results in the user's voice input not being effectively received. The smart speaker has a low wake-up success rate, or may not even be able to be woken up properly, and cannot achieve voice control of linked devices. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide an interaction method for a smart speaker and a smart speaker, so as to improve the success rate of waking up a smart speaker in a high-noise environment.
[0006] In a first aspect, embodiments of the present invention provide a method for waking up a smart speaker, the smart speaker being provided with at least two microphones, the method comprising: waking up the smart speaker in response to one of the microphones being blocked.
[0007] In an optional embodiment of this application, the step of waking up the smart speaker in response to one of the microphones being blocked includes: determining that one of the microphones is blocked based on the difference in signal strength between the audio signals picked up by each microphone.
[0008] In an optional embodiment of this application, the step of determining that one of the microphones is blocked based on the difference in signal strength between the audio signals picked up by each microphone includes: acquiring the audio signals picked up by each microphone and determining the signal strength of each audio signal; selecting the audio signal with the weakest signal strength as the target audio signal and determining the target audio signal strength value based on the signal strength of the target audio signal; determining a reference audio signal strength value based on the signal strength of other audio signals besides the target audio signal; and determining that the microphone corresponding to the target audio signal is blocked if the difference between the reference audio signal strength value and the target audio signal strength value exceeds a corresponding difference threshold.
[0009] In an optional embodiment of this application, the step of determining the reference audio signal strength value based on the signal strength of other audio signals besides the target audio signal includes: performing a square root operation on the signal strength values of other audio signals picked up by microphones other than the microphone corresponding to the target audio signal to obtain the reference audio signal strength value.
[0010] In an optional embodiment of this application, the smart speaker is further provided with a wireless communication unit. The smart speaker establishes a communication connection with a corresponding linkage device through the wireless communication unit, and the linkage device is a noise source. After the step of waking up the smart speaker in response to one of the microphones being blocked, the method further includes: in response to the noise signal picked up by the microphone exceeding the corresponding noise threshold, controlling the linkage device to reduce its working level in order to pick up the corresponding voice signal for dialogue.
[0011] In an optional embodiment of this application, after the step of controlling the linkage device to reduce its operating level in response to the noise signal picked up by the microphone exceeding the corresponding noise threshold in order to pick up the corresponding voice signal for dialogue, the method further includes: in response to the end of the dialogue, controlling the linkage device to restore the operating level before the reduced operating level.
[0012] In an optional embodiment of this application, the microphone is disposed on the top or side of the smart speaker.
[0013] In an optional embodiment of this application, the microphone corresponding to the target audio signal is blocked based on the corresponding pressing operation.
[0014] In an optional embodiment of this application, a touch sensor is provided on the smart speaker at the position corresponding to the microphone; the step of waking up the smart speaker in response to one of the microphones being blocked includes: in response to the touch sensor being triggered, determining that the microphone corresponding to the touch sensor is blocked, and waking up the smart speaker.
[0015] Secondly, embodiments of the present invention also provide a smart speaker, characterized in that the smart speaker is used to execute the above-described smart speaker interaction method.
[0016] The embodiments of the present invention bring the following beneficial effects: This invention provides an interaction method and a smart speaker that wakes up in response to one of its microphones being blocked. In this method, the smart speaker's microphones are used as buttons, and the blocking of one microphone determines the wake-up of the smart speaker, thus improving the success rate of waking up the smart speaker.
[0017] Other features and advantages of this disclosure will be set forth in the following description, or some features and advantages may be inferred from the description or determined without doubt, or may be learned by practicing the techniques described above.
[0018] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of the structure of a smart speaker provided in an embodiment of the present invention; Figure 2 A flowchart illustrating an interaction method for a smart speaker provided in an embodiment of the present invention; Figure 3 A flowchart illustrating another interaction method for a smart speaker provided in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the communication connection between a smart speaker and a linked device, provided as an embodiment of the present invention. Figure 5 This is a structural schematic diagram of a smart speaker provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Currently, when the range hood is operating on the high-noise setting, the smart speaker is in a high-noise environment, which often results in the user's voice input not being effectively received. The smart speaker has a low wake-up success rate, or may not even be able to be woken up properly, and cannot achieve voice control of linked devices.
[0023] Based on this, embodiments of the present invention provide an interaction method for a smart speaker and a smart speaker, specifically providing a method to improve the wake-up efficiency of a smart speaker, which can increase the success rate of waking up a smart speaker in a high-noise environment.
[0024] To facilitate understanding of this embodiment, a detailed description of an interaction method for a smart speaker disclosed in this embodiment of the invention will be provided first.
[0025] Example 1: This invention provides an interaction method for a smart speaker, wherein the smart speaker is equipped with at least two microphones, as described above. Figure 1 The diagram shown is a structural schematic of a smart speaker. Figure 1 The smart speaker shown has four microphones, which are located in the four corners of the smart speaker (i.e., the top left, bottom left, top right, and bottom right corners).
[0026] In some embodiments, the microphone can be located on the top or side of the smart speaker. Positioning the microphone on the top or side of the smart speaker makes it convenient for the user to press it.
[0027] Based on the above description, see Figure 2 The flowchart shown illustrates an interaction method for a smart speaker, which includes the following steps: Step S202: In response to one of the microphones being blocked, the smart speaker is woken up.
[0028] In this embodiment, the smart speaker can be woken up after one of its microphones is blocked.
[0029] In some embodiments, one of the microphones may be determined to be blocked based on the difference in signal strength between the audio signals picked up by each microphone.
[0030] In this embodiment, the user can use the microphone as a button. When the user presses one of the microphones, the microphone is blocked, and the smart speaker can be woken up in response to the blocked microphone.
[0031] In this embodiment, audio signals picked up by multiple microphones can be acquired, and one microphone can be determined to be blocked based on the difference in signal strength between the multiple audio signals.
[0032] like Figure 1 As shown, the four microphones can be connected to the processor of the smart speaker via the PDM (Pulse Density Modulation) bus. The smart speaker can read the audio signals output by the four microphones in sequence via the PDM bus and then calculate the signal strength of multiple audio signals.
[0033] In some embodiments, the audio signal picked up by each microphone can be acquired, and the signal strength of each audio signal can be determined; the audio signal with the weakest signal strength can be selected as the target audio signal, and the target audio signal strength value can be determined based on the signal strength of the target audio signal; a reference audio signal strength value can be determined based on the signal strength of other audio signals besides the target audio signal; if the difference between the reference audio signal strength value and the target audio signal strength value exceeds the corresponding difference threshold, it is determined that the microphone corresponding to the target audio signal is blocked.
[0034] In some embodiments, the microphone corresponding to the target audio signal is blocked based on the corresponding pressing operation.
[0035] In this embodiment, the user can press the microphone corresponding to the target audio signal, and at this time it can be determined that the microphone corresponding to the target audio signal is blocked.
[0036] like Figure 1 As shown, in this embodiment, the weakest audio signal can be identified and used as the target audio signal, and its strength value can be determined. Alternatively, a reference audio signal strength value can be determined based on the signal strength of other audio signals besides the target audio signal. Then, it can be determined whether the difference between the reference audio signal strength value and the target audio signal strength value exceeds a corresponding difference threshold. If the difference threshold is exceeded, it can be assumed that the user is pressing the microphone corresponding to the target audio signal with their finger. In this case, the microphone is used as a button by the user, which can wake up the smart speaker, thereby increasing the success rate of waking up the smart speaker.
[0037] In this embodiment, it can also be determined whether the smart speaker is in a high-noise environment by determining whether the average signal strength value of other audio signals besides the standard audio signal is greater than a preset threshold. If it is greater, the smart speaker can be considered to be in a high-noise environment.
[0038] If the smart speaker is in a high-noise environment, and since it has been previously determined that the user can wake up the smart speaker by pressing the microphone corresponding to the target audio signal with their finger, the success rate of waking up the smart speaker in a high-noise environment is improved.
[0039] This invention provides an interaction method for a smart speaker, which wakes up the smart speaker in response to one of its microphones being blocked. In this method, the smart speaker's microphones are used as buttons, and the blocking of one microphone determines the wake-up of the smart speaker, thus improving the success rate of waking up the smart speaker.
[0040] Example 2: This embodiment provides another interaction method for a smart speaker, which is implemented based on the above embodiment. It focuses on describing the control method of the smart speaker with a range hood in a high-noise environment. (See also...) Figure 3 The flowchart shown represents another interaction method for a smart speaker, which includes the following steps: Step S302: In response to one of the microphones being blocked, the smart speaker is woken up.
[0041] In some embodiments, the signal strength values of other audio signals picked up by microphones other than the microphone corresponding to the target audio signal can be squared to obtain the reference audio signal strength value.
[0042] In this embodiment, the root mean square (RMS) calculation can be performed on the signal strength values of other audio signals picked up by microphones other than the microphone corresponding to the target audio signal to obtain the reference audio signal strength value.
[0043] In this embodiment, the code for the square root operation function can be: float calculate_rms(int16_t buffer, int samples) { long sum = 0; for(int i = 0; i <samples; i++) { sum += (long)buffer[i] (long)buffer[i]; } return sqrtf(sum / (float)samples); } in, The buffer refers to the starting address where the audio signal is stored, and the samples refer to the length of the data to be calculated.
[0044] For example, if the audio signal is sampled for 100 milliseconds, to obtain the root mean square (RMS) of the 100-millisecond audio signal, the audio signal sampling rate is 8kHz, which means 8000 samplings per second. For 100 milliseconds, this means 800 samplings, resulting in 800 audio signals. These 800 audio signals are then substituted into the function mentioned above, squared, and summed to obtain the sum of squares. Finally, the square root of the sum of squares is calculated to obtain the RMS of the 100-millisecond audio signal, which is used as the reference audio signal strength value.
[0045] In this embodiment, the reference audio signal strength value can also be converted into a decibel value: dB=20×log 10 (reference / linear); where dB is the decibel value, reference is the preset reference value, and linear is the reference audio signal strength value.
[0046] Where `linear` can be the linear amplitude of the input, i.e., the reference audio signal strength value calculated above. `reference` can be a reference value, for example: 32768.0f, corresponding to the maximum value of a 16-bit signed integer ±32768. `log10` is a logarithmic function with base 10, which can be a single-precision floating-point number. `20` is a preset coefficient used for converting voltage / amplitude to decibels. The code for this function can be: float linear_to_db(float linear) { return 20.0f log10f(linear / 32768.0f);} by Figure 1 Taking the four microphones shown as an example, the signal strength of the audio signals corresponding to the four microphones can be calculated. If one of these four signal strengths represents the target audio signal strength value, the average of the signal strength values of the other three audio signals is used as the reference audio signal strength value. If the difference between the reference audio signal strength value and the target audio signal strength value is greater than a preset difference threshold, it indicates that the user has pressed the microphone corresponding to the target audio signal.
[0047] This embodiment can also perform noise detection. Continuing with... Figure 1 Taking the four microphones shown as an example, we can determine the three other audio signals collected by the three microphones other than the microphone corresponding to the target audio signal, and then determine whether the average signal strength of the three other audio signals is greater than a preset threshold (which can be 70dB). Thus, we can determine whether it is a high-noise environment through the three other microphones.
[0048] If the noise level exceeds the sound pressure threshold, it is identified as a high-noise environment. The smart speaker can be activated when the user presses the microphone in a high-noise environment.
[0049] Therefore, this embodiment of the invention provides an interactive method for waking up a smart speaker in a high-noise environment. The method can use a program to detect whether a microphone of the smart speaker is pressed by a finger and the ambient noise is greater than a certain threshold. If so, it is determined that the smart speaker is manually woken up, so that the smart speaker can perform subsequent voice command recognition.
[0050] In step S304, in response to the noise signal picked up by the microphone exceeding the corresponding noise threshold, the linkage device is controlled to reduce its working level in order to pick up the corresponding voice signal for dialogue.
[0051] In some embodiments, the smart speaker is further provided with a wireless communication unit, through which the smart speaker establishes a communication connection with a corresponding linked device, which is a noise source.
[0052] In this embodiment, the smart speaker communicates with the linked device via a wireless communication unit. See also... Figure 4 The diagram shows the connection between a smart speaker and a linked device. The wireless communication unit can be WiFi (mobile hotspot), and the smart speaker and the linked device can communicate with the server via WiFi (mobile hotspot).
[0053] If the noise signal picked up by the microphone exceeds the corresponding noise threshold, it can be determined that the smart speaker is in a high-noise environment, and the noise is very likely generated by the linked devices, meaning the linked devices are likely the noise source. Therefore, the smart speaker can control the linked devices to lower their operating level to pick up the appropriate voice signals for dialogue, thereby reducing the noise generated by the linked devices.
[0054] Among them, the linkage device can be a range hood; it can determine the working level of the range hood; if the working level is high or stir-frying, it controls the range hood to lower the working level to medium.
[0055] In this embodiment, a range hood can be used as an example. The smart speaker can query the operating setting of the range hood via a wireless communication unit (such as a WiFi network). When the queried operating setting is high or stir-fry setting, the noise can be considered to be generated by the range hood at high or stir-fry setting.
[0056] In this embodiment, it can be determined that the user manually wakes up the smart speaker, the smart speaker enters the dialogue state, and can start a countdown for the dialogue timeout (the countdown value is the time threshold). At the same time, it sends a command via WiFi to adjust the range hood to the medium setting.
[0057] In some embodiments, in response to the end of the dialogue, the control linkage device is restored to the working level before the working level was lowered.
[0058] When the countdown reaches 0, it can be considered that the duration of the dialogue state has exceeded the time threshold. The smart speaker can exit the dialogue state and return to the standby state. At the same time, it can send a command via WiFi to adjust the range hood back to its original working level (such as high or stir-fry).
[0059] Therefore, this invention provides a method to improve the success rate of voice recognition when the range hood is operating at high speed or stir-fry speed. When the smart speaker is manually woken up by the user in a high-noise environment, the operating speed of the range hood is automatically lowered to reduce the noise generated by the range hood. After the user's voice interaction ends, the operating speed of the range hood is automatically restored, thereby reducing the noise emitted by the range hood during the user's use of the smart speaker's voice function, and thus improving the success rate of voice control of linked devices during this period.
[0060] In some embodiments, a touch sensor is provided on the smart speaker at the location corresponding to the microphone; the smart speaker can be woken up in response to the touch sensor being triggered, which determines that the microphone corresponding to the touch sensor is blocked.
[0061] In this embodiment, in addition to determining whether the user has pressed the microphone by the difference in the signal strength of the microphones, the user can also directly determine whether the user has pressed the microphone by the touch sensor set at the microphone location. Although an additional touch sensor needs to be set in the smart speaker, it can reduce the number of algorithms and thus improve the efficiency of determining whether the user has pressed the microphone.
[0062] In summary, the method provided by the embodiments of the present invention can use the microphone of the smart speaker as a button in high-noise environments caused by the high setting or stir-fry setting of the range hood. The signal strength collected by the microphone determines whether to wake up the smart speaker, which can improve the success rate of waking up the smart speaker in high-noise environments.
[0063] The method provided in this embodiment of the invention can utilize the existing microphone of the smart speaker to detect button presses. When a button press is detected and the background noise exceeds a certain threshold, it is determined to be a wake-up event for the smart speaker. This increases the wake-up success rate in high-noise scenarios without adding any additional devices.
[0064] The method provided in this embodiment of the invention can also be used for the coordinated control of a smart speaker and a range hood. When the range hood is on a high setting or stir-fry setting and there is a need for voice interaction, the range hood's setting is automatically lowered to reduce noise and improve the success rate of voice interaction. The method can automatically adjust the range hood's setting to reduce noise, facilitating voice interaction with the smart speaker. After the voice interaction ends, the range hood's setting can be automatically restored.
[0065] Example 3: This invention also provides a smart speaker for running the above-described smart speaker interaction method; see [link to related documentation]. Figure 5The diagram shows the structure of a smart speaker, which includes a memory 100 and a processor 101. The memory 100 stores one or more computer instructions, which are executed by the processor 101 to implement the interaction method of the smart speaker.
[0066] Furthermore, Figure 5 The smart speaker shown also includes a bus 102 and a communication interface 103. The processor 101, the communication interface 103, and the memory 100 are connected via the bus 102.
[0067] The memory 100 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0068] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. Processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 100, and processor 101 reads information from memory 100 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0069] This invention also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are called and executed by a processor, they cause the processor to implement the above-described smart speaker interaction method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0070] The smart speaker interaction method and smart speaker computer program product provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0071] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and / or device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0072] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0073] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0074] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0075] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for waking up a smart speaker, characterized in that, The smart speaker is equipped with at least two microphones, and the method includes: The smart speaker is activated in response to one of its microphones being blocked.
2. The method according to claim 1, characterized in that, The steps for waking up the smart speaker in response to one of the microphones being blocked include: Based on the difference in signal strength between the audio signals picked up by each of the microphones, it is determined that one of the microphones is blocked.
3. The method according to claim 2, characterized in that, The step of determining which microphone is blocked based on the difference in signal strength between the audio signals picked up by each microphone includes: Acquire the audio signal picked up by each microphone and determine the signal strength of each audio signal; The audio signal with the weakest signal strength is selected as the target audio signal, and the target audio signal strength value is determined based on the signal strength of the target audio signal. A reference audio signal strength value is determined based on the signal strength of other audio signals besides the target audio signal; If the difference between the reference audio signal strength value and the target audio signal strength value exceeds the corresponding difference threshold, it is determined that the microphone corresponding to the target audio signal is blocked.
4. The method according to claim 3, characterized in that, The step of determining a reference audio signal strength value based on the signal strength of other audio signals besides the target audio signal includes: The reference audio signal strength value is obtained by taking the square root of the signal strength values of other audio signals picked up by microphones other than the microphone corresponding to the target audio signal.
5. The method according to claim 1, characterized in that, The smart speaker is also equipped with a wireless communication unit, through which the smart speaker establishes a communication connection with a corresponding linked device, which is a noise source; Following the step of waking up the smart speaker in response to one of the microphones being blocked, the method further includes: In response to the noise signal picked up by the microphone exceeding the corresponding noise threshold, the linkage device is controlled to reduce its operating level in order to pick up the corresponding voice signal for dialogue.
6. The method according to claim 5, characterized in that, After the step of controlling the linkage device to reduce its operating level to pick up the corresponding voice signal for dialogue in response to the noise signal picked up by the microphone exceeding a corresponding noise threshold, the method further includes: In response to the end of the dialogue, the control of the linkage device is restored to the working level before the working level was lowered.
7. The method according to any one of claims 1-6, characterized in that, The microphone is located on the top or side of the smart speaker.
8. The method according to claim 3, characterized in that, The microphone corresponding to the target audio signal is blocked based on the corresponding pressing operation.
9. The method according to claim 1, characterized in that, A touch sensor is provided on the smart speaker at the position corresponding to the microphone; the step of waking up the smart speaker in response to one of the microphones being blocked includes: In response to the touch sensor being triggered, it is determined that the microphone corresponding to the touch sensor is blocked, and the smart speaker is woken up.
10. A smart speaker, characterized in that, The smart speaker is used to execute the interaction method of the smart speaker according to any one of claims 1-9.
Citation Information
Patent Citations
Voice receiving quality detection method and apparatus, and terminal device
CN106453970A
Voice wake-up method and equipment
CN113571053A
Pickup method, pickup device and storage medium
CN118870239A
Dehumidifier and control method and device thereof, storage medium and computer program product
CN120799643A
Method for waking up intelligent assistant and electronic equipment
CN120929144A
Cited By
Sound detection method
CN122116940A