Equipment awakening method and device, storage medium and electronic device
By calculating the audio energy, capability, and inertia value of voice interaction devices in a smart home environment, the wake-up weight is determined, the device wake-up logic is optimized, the false wake-up problem is solved, and the accuracy of device wake-up and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202411161271.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-03
AI Technical Summary
In smart home environments with multiple devices, voice interaction devices frequently experience false wake-ups, leading to a decline in user experience and a waste of resources.
By receiving audio energy values, voice capability values, and wake-up inertia values from multiple voice interaction devices, the wake-up weight is calculated, and the device with the highest weight is woken up first, and a wake-up command is sent.
It improves the accuracy of device wake-up, avoids false wake-ups and resource waste from simultaneous wake-ups of multiple devices, ensures coordinated operation of multiple voice interaction devices, and optimizes device wake-up logic.
Smart Images

Figure CN121600919A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home / intelligent home technology, and more specifically, to a device wake-up method, apparatus, storage medium, and electronic device. Background Technology
[0002] With the continuous advancement of technologies such as the Internet of Things and artificial intelligence, smart home systems are becoming increasingly complex, making collaborative work between different voice interaction devices particularly important. Distributed wake-up technology, through sophisticated intelligent algorithms, can effectively manage voice commands in multi-device environments, ensuring that only the device closest to the user or the preset target responds when a wake-up signal is received.
[0003] However, since multiple devices coexist in a smart home system and all of them have voice interaction capabilities, false wake-up problems often occur when users interact with the voice interaction devices in the smart home system, affecting the user experience. Summary of the Invention
[0004] This application provides a device wake-up method, apparatus, storage medium, and electronic device to solve the technical problem of frequent false wake-ups when multiple voice interaction devices coexist.
[0005] In a first aspect, this application provides a device wake-up method, comprising:
[0006] Receive the first audio energy value of the first wake-up audio received by the voice interaction device from each of the multiple voice interaction devices;
[0007] Based on the first audio energy value, voice capability value and wake-up inertia value corresponding to each of the voice interaction devices, the wake-up weight of each of the voice interaction devices is determined, wherein the voice capability value is used to characterize the range of voice capabilities supported by the voice interaction device, and the wake-up inertia value is used to characterize the record of the voice interaction being woken up and the record of the user interacting with the voice interaction after being woken up.
[0008] Based on the wake-up weights of the multiple voice interaction devices, the first target device among the multiple voice interaction devices is woken up.
[0009] Optionally, the method further includes:
[0010] The voice capability value of the voice interaction device is determined based on the number and / or types of voice capabilities supported by the voice interaction device.
[0011] Optionally, the method further includes:
[0012] Based on the records of the voice interaction device being woken up and the records of the user's voice interaction with the voice interaction device after being woken up, the historical wake-up inertia value of the voice interaction device is adjusted to obtain the wake-up inertia value of the corresponding device, wherein the initial wake-up inertia value of the voice interaction device is a preset wake-up inertia value.
[0013] Optionally, adjusting the historical wake-up inertia value of the voice interaction device to obtain the corresponding device's wake-up inertia value based on the records of the voice interaction device being woken up and the records of the user's voice interaction with the voice interaction device after being woken up includes:
[0014] If the first device is woken up and the user interacts with the first device via voice, and the user's voice intent is the voice capability supported by the first device, then the historical wake-up inertia value of the first device is increased by a preset value to obtain the wake-up inertia value of the first device.
[0015] If the second device is woken up and the user interacts with the second device via voice, but the user's voice intent is the voice capability supported by the third device, then the historical wake-up inertia value of the third device is increased by the preset value to obtain the wake-up inertia value of the third device.
[0016] If the fourth device is woken up, but the user does not interact with the fourth device via voice, and the fifth device is woken up again, and the user interacts with the fifth device via voice, and the user's voice intent is the voice capability supported by the fifth device, then the historical wake-up inertia value of the fifth device is increased by the preset value to obtain the wake-up inertia value of the fifth device, and the historical wake-up inertia value of the fourth device is decreased by the preset value to obtain the wake-up inertia value of the fourth device.
[0017] Optionally, determining the wake-up weight of each voice interaction device based on the first audio energy value, voice capability value, and wake-up inertia value corresponding to each voice interaction device includes:
[0018] The wake-up weight is obtained by weighting and summing the first audio energy value, the voice capability value, and the wake-up inertia value, wherein the weighting coefficient of the first audio energy value is greater than the weighting coefficient of the voice capability value, and the weighting coefficient of the voice capability value is greater than the weighting coefficient of the wake-up inertia value.
[0019] Optionally, waking up the first target device among the plurality of voice interaction devices based on their respective wake-up weights includes:
[0020] If the first target device is the device with the highest wake-up weight among the plurality of voice interaction devices, then a wake-up command is sent to the first target device, and the wake-up command is used to wake up the first target device.
[0021] Optionally, the method further includes:
[0022] After the first target device is woken up, if the user does not interact with the first target device by voice, and other devices are woken up again, and the user interacts with the other devices by voice, and the user's voice intent is the voice capability supported by the other devices, then in the next instance of receiving the second audio energy value of the second wake-up audio received by each of the multiple devices, the second target device among the multiple voice interaction devices is woken up based on the second audio energy value.
[0023] Secondly, this application provides a device wake-up device, comprising:
[0024] The processing module is used to receive the first audio energy value of the first wake-up audio received by the first voice interaction device from each of the multiple voice interaction devices;
[0025] The determining module is used to determine the wake-up weight of each of the voice interaction devices based on the first audio energy value, voice capability value and wake-up inertia value corresponding to each of the voice interaction devices, wherein the voice capability value is used to characterize the range of voice capabilities supported by the voice interaction device, and the wake-up inertia value is used to characterize the record of the voice interaction being woken up and the record of the user interacting with the voice interaction after being woken up.
[0026] The processing module is further configured to wake up the first target device among the plurality of voice interaction devices based on the wake-up weights of the respective voice interaction devices.
[0027] Optionally, the determining module is further configured to determine the voice capability value of the voice interaction device based on the number and / or types of voice capabilities supported by the voice interaction device.
[0028] Optionally, the processing module is further configured to adjust the historical wake-up inertia value of the voice interaction device to obtain the wake-up inertia value of the corresponding device based on the records of the voice interaction device being woken up and the records of the user interacting with the voice interaction device after being woken up, wherein the initial wake-up inertia value of the voice interaction device is a preset wake-up inertia value.
[0029] Optionally, if the first device is woken up and the user interacts with the first device via voice, and the user's voice intent is the voice capability supported by the first device, then the processing module is further configured to increase the historical wake-up inertia value of the first device by a preset value to obtain the wake-up inertia value of the first device.
[0030] If the second device is woken up and the user interacts with the second device via voice, but the user's voice intent is for the voice capability supported by the third device, then the processing module is further configured to add the preset value to the historical wake-up inertia value of the third device to obtain the wake-up inertia value of the third device.
[0031] If the fourth device is woken up, but the user does not interact with the fourth device via voice, and the fifth device is woken up again, and the user interacts with the fifth device via voice, and the user's voice intent is the voice capability supported by the fifth device, then the processing module is further configured to increase the historical wake-up inertia value of the fifth device by the preset value to obtain the wake-up inertia value of the fifth device, and decrease the historical wake-up inertia value of the fourth device by the preset value to obtain the wake-up inertia value of the fourth device.
[0032] Optionally, the processing module is further configured to perform a weighted summation of the first audio energy value, the voice capability value, and the wake-up inertia value to obtain the wake-up weight, wherein the weighting coefficient of the first audio energy value is greater than the weighting coefficient of the voice capability value, and the weighting coefficient of the voice capability value is greater than the weighting coefficient of the wake-up inertia value.
[0033] Optionally, if the first target device is the device with the highest wake-up weight among the plurality of voice interaction devices, the processing module is further configured to send a wake-up command to the first target device, the wake-up command being used to wake up the first target device.
[0034] Optionally, after the first target device is woken up, if the user does not interact with the first target device via voice, and other devices are continuously woken up, and the user interacts with the other devices via voice, and the user's voice intent is the voice capability supported by the other devices, then the processing module is further configured to wake up the second target device among the multiple voice interaction devices based on the second audio energy value when receiving the second audio energy value of the second wake-up audio received by the device from each of the multiple devices in the next instance.
[0035] Thirdly, this application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0036] The memory stores computer-executed instructions;
[0037] The processor executes computer execution instructions stored in the memory to implement the device wake-up method as described in the first aspect and various possible implementations of the first aspect above.
[0038] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions thereon, which, when executed by a processor, are used to implement the device wake-up method as described in the first aspect and various possible implementations of the first aspect.
[0039] Fifthly, this application provides a program product including a computer program that, when executed by a processor, implements the device wake-up method described above.
[0040] The device wake-up method, apparatus, storage medium, and electronic device provided in this application involve receiving the first audio energy value of a first wake-up audio signal received by each of multiple voice interaction devices; determining the voice capability value of a corresponding device based on the number and / or types of voice capabilities supported by different voice interaction devices; adjusting the historical wake-up inertia value of the corresponding voice interaction device based on records of different voice interaction devices being woken up and records of user voice interaction with the corresponding device after being woken up to obtain the wake-up inertia value of multiple voice interaction devices; weighted summing of the first audio energy value, voice capability value, and wake-up inertia value to obtain the wake-up weight; and if the first target device is the device with the highest wake-up weight among the multiple voice interaction devices, sending a wake-up command to the first target device, which then provides services to the user. This method optimizes the device wake-up logic in a smart home environment, improves the accuracy of device wake-up, avoids false wake-ups and resource waste from multiple voice interaction devices being woken up simultaneously, ensures coordinated operation of multiple voice interaction devices, and can continuously optimize the device wake-up logic by dynamically adjusting the wake-up inertia value of the device, thereby improving the accuracy of distributed wake-up. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0042] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 A schematic diagram illustrating a scenario for the device wake-up method provided in this application;
[0044] Figure 2 Flowchart of the device wake-up method provided in this application Figure 1 ;
[0045] Figure 3Flowchart of the device wake-up method provided in this application Figure 2 ;
[0046] Figure 4 A schematic diagram of the device wake-up device provided in this application;
[0047] Figure 5 A schematic diagram of the device wake-up device provided in this application. Detailed Implementation
[0048] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0049] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0050] According to one aspect of the embodiments of this application, an interaction method for smart home devices is provided. This interaction method for smart home devices is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned interaction method for smart home devices can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 101 and server 102. For example... Figure 1 As shown, server 102 is connected to terminal device 101 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 102. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 102.
[0051] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. Terminal device 101 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0052] With the continuous advancement of technologies such as the Internet of Things and artificial intelligence, smart home systems are becoming increasingly complex, making collaborative work between different voice interaction devices particularly important. Distributed wake-up technology, through sophisticated intelligent algorithms, can effectively manage voice commands in multi-device environments, ensuring that only the device closest to the user or the preset target responds when a wake-up signal is received.
[0053] However, in a smart home environment with multiple devices, the following technical issues arise when users attempt to wake up and control specific devices (such as televisions) via voice:
[0054] (1) Voice recognition conflict: Since multiple devices (such as TV, central control screen, and standing air conditioner) are close to the user and all have voice recognition functions, the user's wake word or command may be captured by multiple devices at the same time, making it impossible to accurately determine which device the user intends to target.
[0055] (2) Limitations of the wake-up mechanism: The existing device wake-up mechanism is easily affected by environmental factors (such as background noise, similar wake words of other devices, etc.) and triggers false wake-ups, thus affecting the user experience.
[0056] The device wake-up method provided in this application aims to solve the above-mentioned technical problems of the prior art.
[0057] First, the implementation scenarios involved in this application will be explained.
[0058] Figure 1 This is a schematic diagram illustrating a scenario of the device wake-up method provided in this application. For example... Figure 1As shown, the smart terminal device group 101 is communicatively connected to the server 102, and multiple devices in the smart terminal device group 101 are in the same smart home system. Each device can receive energy information from voice commands issued by the user and send this energy information to the server 102. The server 102 stores a voice wake-up module. Using this voice wake-up module and the energy information obtained from the corresponding device, it determines the device interacting with the user and sends a wake-up command to that device to initiate voice interaction with the user. The smart terminal device group 101 can be, for example, multiple smart devices in the same smart home environment, including a smart TV, a smart air conditioner, and a smart washing machine. The server 102 can be, for example, a smart home central server equipped with a voice wake-up module.
[0059] This application provides a device wake-up method. When a user is in a smart home environment and interacts with multiple smart devices in that environment, the method acquires the audio energy values received by different smart devices and determines the voice capabilities and wake-up inertia values of each device. Based on the audio energy values, voice capabilities, and wake-up inertia values of different devices, the wake-up weight of each device is calculated to determine the target device to be woken up. This method optimizes the device wake-up logic in a smart home environment, improves the accuracy of device wake-up, avoids false wake-ups and resource waste caused by multiple devices being woken up simultaneously, ensures the coordinated operation of multiple voice interaction devices, and can continuously optimize the device wake-up logic by dynamically adjusting the wake-up inertia values of the devices, thereby improving the accuracy of distributed wake-up.
[0060] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0061] Figure 2 Flowchart of the device wake-up method provided in the embodiments of this application Figure 1 The implementing entity in this embodiment can be, for example, a smart home system equipped with a voice wake-up module. Figure 2 As shown, the device wake-up method provided in this embodiment includes:
[0062] S201: Receive the first audio energy value of the first wake-up audio received by the voice interaction device from each of the multiple voice interaction devices.
[0063] The audio energy value is used to indicate the strength or energy level of the wake-up audio signal received by the voice interaction device.
[0064] Understandably, multiple smart devices in a smart home system are typically equipped with audio receiving units and audio processing units, enabling them to receive and process audio signals emitted by the user. These devices operate on the same network, facilitating communication and collaboration between different voice interaction devices or between the voice interaction devices and the smart home system. When a user issues a voice command, the audio receiving units on different smart devices receive the corresponding audio information, and the audio processing unit on the corresponding voice interaction device calculates the energy level of the audio information, thus obtaining the audio energy value for that device. Furthermore, after calculating the audio energy values of the audio information received by each smart device, the multiple smart devices communicate with the smart home system, and the smart home system can receive the audio energy values sent by the multiple smart devices.
[0065] For example, the audio receiving unit in a smart home system can be a microphone, and the audio processing unit can be a digital signal processor. When a user is in a smart home environment and issues a voice command, the microphones in smart devices A, B, and C respectively receive the changes in the sound waveform within a time period of the user's voice command, i.e., the first wake-up audio. Then, the digital signal processing device in the corresponding device converts the changes in the corresponding sound waveform into sound intensity, which can be measured in decibels. The processed decibel value is then sent to the smart home system. The first audio energy values obtained by the smart home system for different devices at that time are: smart device A - decibel value A, smart device B - decibel value B, and smart device C - decibel value C, respectively. This application does not impose any special restrictions on the type of audio energy value.
[0066] S202: Determine the wake-up weight of each voice interaction device based on the first audio energy value, voice capability value, and wake-up inertia value corresponding to each voice interaction device.
[0067] Among them, the voice capability value is used to characterize the range of voice capabilities supported by the voice interaction device, the wake-up inertia value is used to characterize the record of the voice interaction device being woken up and the record of the user's voice interaction with the voice interaction device after being woken up, and the wake-up weight is used to indicate the priority of multiple voice interaction devices being woken up.
[0068] Understandably, the voice capability values of different voice interaction devices are fixed, and these values reflect the richness of the device's functions. Devices that support more voice functions usually have higher voice capability values. Wake-up inertia values reflect the device's performance in historical usage and the frequency with which users use different voice interaction devices. For example, in a smart home system, if any voice interaction device is frequently woken up by the user and the user engages in frequent voice interactions with it, then that device has a high wake-up inertia value.
[0069] For any voice interaction device, the first audio energy value, voice capability value and wake-up inertia value of the corresponding device are standardized so that the three types of data can be directly calculated, and the calculation result is determined as the wake-up weight of the corresponding device.
[0070] For example, for smart devices A, B, and C, the first audio energy value, voice capability value, and wake-up inertia value of each device are converted to the same scale. Specifically, the decibel values A, B, and C are scaled, resulting in first audio energy values of 80, 85, and 75, respectively. If the voice capability values of smart devices A, B, and C after scale conversion are 70, 65, and 80, and the wake-up inertia values are 60, 55, and 70, then the corresponding first audio energy value, voice capability value, and wake-up inertia value are directly calculated for each smart device. The calculated results are the wake-up weights for each device, and the resulting wake-up weights are: Smart device A - 70, Smart device B - 68, and Smart device C - 75. This application does not impose special restrictions on the standardization of the first audio energy value, the voice capability value, and the wake-up inertia value.
[0071] S203: Based on the wake-up weights of the multiple voice interaction devices, wake up the first target device among the multiple voice interaction devices.
[0072] The first target device is used to demonstrate the device most suitable for being woken up in a multi-device environment.
[0073] Understandably, different voice interaction devices have different wake-up weights, and the wake-up weight represents the priority of multiple voice interaction devices being woken up. According to this priority, the most suitable voice interaction device to be woken up in the smart home environment after the user issues a voice command can be determined, that is, the first target device. In the process of determining the first target device according to the wake-up weight, the device with the highest wake-up weight can be determined as the first target device, or the device with a wake-up weight less than the highest wake-up weight but greater than other wake-up weights can be determined as the first target device.
[0074] For example, in a smart home system, the wake-up weights for different smart devices are as follows: smart device A-70, smart device B-68, and smart device C-75. By comparing the multiple wake-up weights, the priority of waking up different smart devices is determined according to their size. That is, smart device C has a higher priority than smart device A, and smart device A has a higher priority than smart device B. Therefore, smart device C is the most suitable device to be woken up at this time. At this point, smart device C can be identified as the first target device and woke up.
[0075] The device wake-up method provided in this embodiment receives the first audio energy value of the first wake-up audio received by each of multiple voice interaction devices. Based on the first audio energy value, voice capability value, and wake-up inertia value corresponding to each voice interaction device, the wake-up weight of each voice interaction device is determined. Then, based on the wake-up weights of each of the multiple voice interaction devices, a first target device among the multiple voice interaction devices is woken up. This method optimizes the device wake-up logic in a smart home environment, improves the accuracy of device wake-up, avoids false wake-ups and resource waste from multiple devices being woken up simultaneously, and ensures coordinated operation of multiple voice interaction devices.
[0076] Figure 3 Flowchart of the device wake-up method provided in the embodiments of this application Figure 2 .like Figure 3 As shown, in this embodiment... Figure 2 Based on the embodiments, the device wake-up method is described in detail below. The device wake-up method shown in this embodiment includes:
[0077] S301: Receive the first audio energy value of the first wake-up audio received by the voice interaction device from each of the multiple voice interaction devices.
[0078] Step S301 is similar to step S201 above, and will not be repeated here.
[0079] S302: Determine the voice capability value of the voice interaction device based on the number and / or types of voice capabilities supported by the voice interaction device.
[0080] Voice capabilities include, for example, providing users with media assets, videos, casual conversation, storytelling, monitoring queries, and control of other home appliances (refrigerators, washing machines, air conditioners, water heaters), among other functions.
[0081] Understandably, smart home systems can acquire the voice capabilities supported by different voice interaction devices and determine the voice capability value of each device based on the number and / or types of voice capabilities supported. The more voice capabilities a device supports, the higher its voice capability value. Furthermore, when determining the voice capability value of a voice interaction device, there are upper and lower limits. For example, if a smart device only supports one voice skill, such as querying monitoring, its voice capability value can be set to 1 point. This application does not impose special restrictions on the upper and lower limits of the voice capability value.
[0082] For example, a smart home system might include a voice capability scoring module. This module quantifies the voice capabilities of different devices, making comparisons between them more intuitive and simple. Smart device A, whose voice skills include media assets, casual conversation, storytelling, monitoring queries, and control of other home appliances (but not the TV), would have a voice capability score of 9. Smart device B, also supporting these skills, would have a voice capability score of 8. Smart device C, supporting media assets, video, casual conversation, storytelling, monitoring queries, and control of other home appliances (refrigerator, washing machine, air conditioner, water heater, etc.), would have a voice capability score of 10. Therefore, smart device C has a higher voice capability score than smart device A, and smart device A has a higher voice capability score than smart device B.
[0083] S303: Based on the records of the voice interaction device being woken up and the records of the user's voice interaction with the voice interaction device after being woken up, adjust the historical wake-up inertia value of the voice interaction device to obtain the wake-up inertia value of the corresponding device.
[0084] The initial wake-up inertia value of the voice interaction device is a preset wake-up inertia value, which can be, for example, 0 or other values.
[0085] Based on the records of different voice interaction devices being woken up and whether the user interacts with the device after being woken up, the historical wake-up inertia value of multiple voice interaction devices is calculated. After the calculation of the historical wake-up inertia value is completed, the updated historical wake-up inertia value is used as the wake-up inertia value for the next interaction with the user. This achieves dynamic updating of the wake-up inertia value and improves the accuracy of locating the wake-up device.
[0086] Specifically, the process of dynamically updating the wake-up inertia value based on the records of the voice interaction device being woken up and the records of the user's voice interaction with the voice interaction device after being woken up includes:
[0087] If the first device is woken up and the user interacts with it via voice, and the user's voice intent is a voice capability supported by the first device, then the historical wake-up inertia value of the first device is increased by a preset value to obtain the wake-up inertia value of the first device. If the second device is woken up and the user interacts with it via voice, but the user's voice intent is a voice capability supported by the third device, then the historical wake-up inertia value of the third device is increased by a preset value to obtain the wake-up inertia value of the third device. If the fourth device is woken up but the user does not interact with it via voice, and the fifth device is woken up again and the user interacts with it via voice, and the user's voice intent is a voice capability supported by the fifth device, then the historical wake-up inertia value of the fifth device is increased by a preset value to obtain the wake-up inertia value of the fifth device, and the historical wake-up inertia value of the fourth device is decreased by a preset value to obtain the wake-up inertia value of the fourth device.
[0088] The preset value can be, for example, 1 point.
[0089] Understandably, if other devices are woken up but the user does not interact with them via voice while the fifth device is being woken up, and the user does not interact with them via voice, then the historical wake-up inertia value of the other woken devices in the smart home environment between the wake-up of the fourth device and the wake-up of the fifth device is reduced by a preset value to obtain the wake-up inertia value of the corresponding device.
[0090] For example, the system acquires records of multiple voice interaction devices being woken up, as well as records of user voice interactions with these devices after being woken up. Based on this information, different wake-up inertia values are dynamically updated. Specifically, this includes: if smart device A is woken up and the user interacts with it via voice, and the user's voice intent is a voice skill supported by smart device A, then 1 point is added to the historical wake-up inertia value of smart device A; if smart device A is woken up and the user interacts with it via voice, but the user's voice intent is a voice skill supported by smart device B, then 1 point is added to the historical wake-up inertia value of smart device B; if smart device A is woken up, but the user does not interact with it via voice, the system continues to wake up smart device A until smart device B is woken up, and the user's current voice intent is a voice skill supported by smart device B, then 1 point is deducted from the historical wake-up inertia value of smart device A, and 1 point is added to the historical wake-up inertia value of smart device B; if smart device C is woken up after smart device A is woken up but before smart device B is woken up, then 1 point is deducted from the historical wake-up inertia value of smart device C.
[0091] S304: The wake-up weight is obtained by weighting and summing the first audio energy value, the voice capability value, and the wake-up inertia value.
[0092] Optionally, the weighting coefficient of the first audio energy value is greater than the weighting coefficient of the speech capability value, the weighting coefficient of the speech capability value is greater than the weighting coefficient of the wake-up inertia value, and the weighting coefficient of the first audio energy value can be, for example, 0.5, the weighting coefficient of the speech capability value can be, for example, 0.3, and the weighting coefficient of the wake-up inertia value can be, for example, 0.2.
[0093] Understandably, the weighted coefficients of the first audio energy value, voice capability value, and wake-up inertia value are data information pre-selected and stored in the smart home system.
[0094] For example, upon receiving the first audio energy values from different smart devices, the voice wake-up module in the smart home system calls upon the voice capability values of the different smart devices, the available wake-up inertia values, and the weighting coefficients of the different smart devices, and calculates the wake-up weights of the different smart devices according to their respective weighting coefficients. If the first audio energy values of smart devices A, B, and C are 80, 85, and 75 respectively, their voice capability values are 70, 65, and 80 respectively, and their wake-up inertia values are 60, 55, and 70 respectively, then according to the different weighting coefficients, the wake-up weight of smart device A is 80.5, the wake-up weight of smart device B is 70.5, and the wake-up weight of smart device C is 60.5. This application does not impose any special restrictions on the weighting coefficients.
[0095] Optionally, when determining the weighting coefficients for the first audio energy value, voice capability value, and wake-up inertia value, the influence of the first audio energy value, voice capability value, and wake-up inertia value on the user's successful wake-up of the corresponding voice interaction device is calculated based on the records of the user's voice interaction with the voice interaction device after it is woken up. The influence is then quantified, and the result is the corresponding weighting coefficient.
[0096] Furthermore, in this smart home environment, operational data on successful wake-up of voice interaction devices and their interaction with users can be acquired in real time. Based on the newly acquired operational data, the weighting coefficients of the first audio energy value, voice capability value, and wake-up inertia value can be updated to continuously optimize the weighting coefficients. This application does not impose any special restrictions on the update frequency of the weighting coefficients.
[0097] Specifically, the first audio energy value, voice capability value, and wake-up inertia value are determined as the influencing indicators of wake-up weight. The calculation process for the corresponding influencing indicators includes, for example, using statistical methods (such as Pearson correlation coefficient, Spearman rank correlation coefficient, etc.) to analyze the correlation between each indicator and the user's successful wake-up of the voice interaction device; alternatively, a multiple regression model can be constructed based on multiple indicators, with the first audio energy value, voice capability value, and wake-up inertia value as independent variables and the user's successful wake-up as the dependent variable, and the degree of independent influence and significance of each indicator can be determined through regression analysis; after determining the degree of influence of different influencing indicators on wake-up weight, since the dimensions and units of each indicator may be different, they need to be normalized so that all indicators are compared within the same numerical range, thereby determining the weighting coefficients corresponding to different influencing indicators.
[0098] S305: If the first target device is the device with the highest wake-up weight among multiple voice interaction devices, then send a wake-up command to the first target device.
[0099] The wake-up command is used to wake up the first target device.
[0100] Based on the wake-up weights of the multiple voice interaction devices acquired at that time, the device with the highest wake-up weight is identified as the first target device, and a wake-up command is sent to the first target device.
[0101] For example, if the wake-up weight of smart device A is calculated to be 80.5, the wake-up weight of smart device B is 70.5, and the wake-up weight of smart device C is 60.5, and the wake-up weight of smart device A is the highest among smart device A, smart device B, and smart device C, and the first target device determined in this instance is smart device A, then the smart home system generates a wake-up command and sends the wake-up command to smart device A.
[0102] Preferably, after the first target device is woken up, if the user does not interact with the first target device via voice, and other devices continue to be woken up, and the user interacts with other devices via voice, and the user's voice intent is the voice capability supported by other devices, then, based on the second audio energy value of the second wake-up audio received by each of the multiple devices in the next instance, the second target device among the multiple voice interaction devices is woken up according to the second audio energy value.
[0103] Understandably, if the user does not interact with the first target device by voice after it is woken up, and other devices continue to be woken up, it indicates that there is a false wake-up in the current wake-up process. In the next device wake-up process, the smart home system will determine the woken-up device only according to the audio energy value received by different devices when waking up the target device.
[0104] The device wake-up method provided in this embodiment involves receiving the first audio energy value of the first wake-up audio received by each of multiple voice interaction devices; determining the voice capability value of each device based on the number and / or types of voice capabilities supported by different voice interaction devices; adjusting the historical wake-up inertia values of different voice interaction devices to obtain the wake-up inertia value of the corresponding device based on records of different voice interaction devices being woken up and records of user voice interaction with the corresponding device after being woken up; weighting and summing the first audio energy value, voice capability value, and wake-up inertia value to obtain the wake-up weight; and if the first target device is the device with the highest wake-up weight among multiple different voice interaction devices, then a wake-up command is sent to the first target device, and the first target device provides services to the user. This method optimizes the device wake-up logic in a smart home environment, improves the accuracy of device wake-up, avoids false wake-ups and resource waste from multiple devices being woken up simultaneously, ensures the coordinated operation of multiple voice interaction devices, and can continuously optimize the device wake-up logic by dynamically adjusting the wake-up inertia value of the device, thereby improving the accuracy of distributed wake-up.
[0105] Figure 4 A schematic diagram of the device wake-up device provided in this application. Figure 4 As shown, this application provides a device wake-up device 400, which includes:
[0106] The processing module 401 is used to receive the first audio energy value of the first wake-up audio received by the first voice interaction device, which is sent by each of the multiple voice interaction devices.
[0107] The determining module 402 is used to determine the wake-up weight of each voice interaction device based on the first audio energy value, voice capability value and wake-up inertia value corresponding to each voice interaction device, wherein the voice capability value is used to characterize the range of voice capabilities supported by the voice interaction device, and the wake-up inertia value is used to characterize the record of the voice interaction being woken up and the record of the user interacting with the voice interaction after being woken up.
[0108] The processing module 401 is further configured to wake up the first target device among the plurality of voice interaction devices based on the wake-up weights of the plurality of voice interaction devices.
[0109] Optionally, the determining module 402 is further configured to determine the voice capability value of the voice interaction device based on the number and / or types of voice capabilities supported by the voice interaction device.
[0110] Optionally, the processing module 401 is further configured to adjust the historical wake-up inertia value of the voice interaction device to obtain the wake-up inertia value of the corresponding device based on the record of the voice interaction device being woken up and the record of the user interacting with the voice interaction device after being woken up, wherein the initial wake-up inertia value of the voice interaction device is a preset wake-up inertia value.
[0111] Optionally, if the first device is woken up and the user interacts with the first device via voice, and the user's voice intent is the voice capability supported by the first device, then the processing module 401 is further configured to increase the historical wake-up inertia value of the first device by a preset value to obtain the wake-up inertia value of the first device.
[0112] If the second device is woken up and the user interacts with the second device via voice, but the user's voice intent is for the voice capability supported by the third device, then the processing module 401 is further configured to add the preset value to the historical wake-up inertia value of the third device to obtain the wake-up inertia value of the third device.
[0113] If the fourth device is woken up, but the user does not interact with the fourth device via voice, and the fifth device is woken up again, and the user interacts with the fifth device via voice, and the user's voice intent is the voice capability supported by the fifth device, then the processing module 401 is further configured to increase the historical wake-up inertia value of the fifth device by the preset value to obtain the wake-up inertia value of the fifth device, and decrease the historical wake-up inertia value of the fourth device by the preset value to obtain the wake-up inertia value of the fourth device.
[0114] Optionally, the processing module 401 is further configured to perform a weighted summation of the first audio energy value, the voice capability value, and the wake-up inertia value to obtain the wake-up weight, wherein the weighting coefficient of the first audio energy value is greater than the weighting coefficient of the voice capability value, and the weighting coefficient of the voice capability value is greater than the weighting coefficient of the wake-up inertia value.
[0115] Optionally, if the first target device is the device with the highest wake-up weight among the plurality of voice interaction devices, the processing module 401 is further configured to send a wake-up command to the first target device, the wake-up command being used to wake up the first target device.
[0116] Optionally, after the first target device is woken up, if the user does not interact with the first target device via voice, and other devices are continuously woken up, and the user interacts with the other devices via voice, and the user's voice intent is the voice capability supported by the other devices, then the processing module 401 is further configured to wake up the second target device among the multiple voice interaction devices based on the second audio energy value when receiving the second audio energy value of the second wake-up audio received by the device from each of the multiple devices in the next instance.
[0117] Figure 5 A schematic diagram of the device wake-up device provided in this application. (For example...) Figure 5 As shown, this application provides a device wake-up device 500, which includes a receiver 501, a transmitter 502, a processor 503, and a memory 504.
[0118] Receiver 501 is used to receive instructions and data;
[0119] Transmitter 502 is used to send commands and data;
[0120] Memory 504 is used to store instructions executed by the computer;
[0121] The processor 503 is used to execute computer execution instructions stored in the memory 504 to implement the various steps of the device wake-up method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing embodiments of the device wake-up method.
[0122] Optionally, the memory 504 can be either standalone or integrated with the processor 503.
[0123] When the memory 504 is set up independently, the electronic device also includes a bus for connecting the memory 504 and the processor 503.
[0124] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0125] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method described in any of the foregoing embodiments.
[0126] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods described in any of the foregoing embodiments.
[0127] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application.
[0128] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The memory may include high-speed RAM, and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk, or optical disc, etc.
[0129] The aforementioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0130] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.
[0131] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0132] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A device wake-up method, characterized in that, include: Receive the first audio energy value of the first wake-up audio received by the voice interaction device from each of the multiple voice interaction devices; Based on the first audio energy value, voice capability value and wake-up inertia value corresponding to each of the voice interaction devices, the wake-up weight of each of the voice interaction devices is determined, wherein the voice capability value is used to characterize the range of voice capabilities supported by the voice interaction device, and the wake-up inertia value is used to characterize the record of the voice interaction device being woken up and the record of the user interacting with the voice interaction device after being woken up. Based on the wake-up weights of the multiple voice interaction devices, the first target device among the multiple voice interaction devices is woken up.
2. The method according to claim 1, characterized in that, The voice capability value of the voice interaction device is determined based on the number and / or types of voice capabilities supported by the voice interaction device.
3. The method according to claim 1, characterized in that, Also includes: Based on the records of the voice interaction device being woken up and the records of the user's voice interaction with the voice interaction device after being woken up, the historical wake-up inertia value of the voice interaction device is adjusted to obtain the wake-up inertia value of the corresponding device, wherein the initial wake-up inertia value of the voice interaction device is a preset wake-up inertia value.
4. The method according to claim 1, characterized in that, The step of adjusting the historical wake-up inertia value of the voice interaction device to obtain the corresponding device's wake-up inertia value based on the records of the voice interaction device being woken up and the records of the user's voice interaction with the voice interaction device after being woken up includes: If the first device is woken up and the user interacts with the first device via voice, and the user's voice intent is the voice capability supported by the first device, then the historical wake-up inertia value of the first device is increased by a preset value to obtain the wake-up inertia value of the first device. If the second device is woken up and the user interacts with the second device via voice, but the user's voice intent is the voice capability supported by the third device, then the historical wake-up inertia value of the third device is increased by the preset value to obtain the wake-up inertia value of the third device. If the fourth device is woken up, but the user does not interact with the fourth device via voice, and the fifth device is woken up again, and the user interacts with the fifth device via voice, and the user's voice intent is the voice capability supported by the fifth device, then the historical wake-up inertia value of the fifth device is increased by the preset value to obtain the wake-up inertia value of the fifth device, and the historical wake-up inertia value of the fourth device is decreased by the preset value to obtain the wake-up inertia value of the fourth device.
5. The method according to any one of claims 1-4, characterized in that, The step of determining the wake-up weight of each voice interaction device based on the first audio energy value, voice capability value, and wake-up inertia value corresponding to each voice interaction device includes: The wake-up weight is obtained by weighting and summing the first audio energy value, the voice capability value, and the wake-up inertia value, wherein the weighting coefficient of the first audio energy value is greater than the weighting coefficient of the voice capability value, and the weighting coefficient of the voice capability value is greater than the weighting coefficient of the wake-up inertia value.
6. The method according to any one of claims 1-4, characterized in that, The step of waking up the first target device among the plurality of voice interaction devices based on their respective wake-up weights includes: If the first target device is the device with the highest wake-up weight among the plurality of voice interaction devices, then a wake-up command is sent to the first target device, and the wake-up command is used to wake up the first target device.
7. The method according to any one of claims 1-4, characterized in that, Also includes: After the first target device is woken up, if the user does not interact with the first target device by voice, and other devices are woken up again, and the user interacts with the other devices by voice, and the user's voice intent is the voice capability supported by the other devices, then in the next instance of receiving the second audio energy value of the second wake-up audio received by each of the multiple devices, the second target device among the multiple voice interaction devices is woken up based on the second audio energy value.
8. A device wake-up device, characterized in that, include: The processing module is used to receive the first audio energy value of the first wake-up audio received by the first voice interaction device from each of the multiple voice interaction devices; The determining module is used to determine the wake-up weight of each of the voice interaction devices based on the first audio energy value, voice capability value and wake-up inertia value corresponding to each of the voice interaction devices, wherein the voice capability value is used to characterize the range of voice capabilities supported by the voice interaction device, and the wake-up inertia value is used to characterize the record of the voice interaction being woken up and the record of the user interacting with the voice interaction after being woken up. The processing module is further configured to wake up the first target device in the interaction of the multiple voice interaction devices based on the wake-up weights of the multiple voice interaction devices.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 7 through the computer program.
Citation Information
Patent Citations
Equipment awakening method and system
CN111540360A
Voice equipment, awakening method and device thereof and storage medium
CN112164405A
Equipment awakening method and device thereof, intelligent terminal and equipment awakening system
CN113889101A
Voice equipment awakening method and device, electronic equipment, storage medium and chip
CN115966204A
Voice interaction device and voice interaction method
CN116129942A