A voice wake-up method and an electronic device

By building a location map and combining face orientation, using a multi-microphone array and image acquisition module, the device that users want to wake up is determined, which solves the problem of inaccurate wake-up in multiple device scenarios and improves the user experience.

CN114566171BActive Publication Date: 2025-07-29HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011362525.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-27
Publication Date
2025-07-29
Estimated Expiration
2040-11-27

AI Technical Summary

Technical Problem

In multi-device scenarios, the user wake-up words are responded to by multiple devices at the same time, resulting in poor user experience. The prior art cannot effectively solve the problem of multiple devices being awakened at the same time. Especially in smart home scenarios, the lack of image acquisition function of the device limits the application of wake-up method based on face orientation.

Method used

By positioning the relative positions of the user and multiple devices in the space, building a position map, and combining the image acquisition module to obtain the user's face orientation, determining the device that the user wants to wake up, using a multi-microphone array to collect sound to determine the user's location, and combining the relative positions and face orientation of the device to accurately wake up the target device.

Benefits of technology

It improves the accuracy and flexibility of device wake-up in multi-device scenarios, avoids the problem of multiple devices responding simultaneously, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114566171B_ABST
    Figure CN114566171B_ABST
Patent Text Reader

Abstract

The present application provides a voice wake-up method and an electronic device, which relates to the field of terminal artificial intelligence. Among them, the method includes: by using the ambient sounds collected by each device, on the one hand, the relative positions of the user and multiple devices in space can be located to construct a position map; on the other hand, the main device with an image acquisition module among the multiple devices can be used to collect the face orientation of the user. In this way, by combining the position map and the face orientation of the user collected by the main device, the device that the user wants to wake up can be determined. This method helps to improve the accuracy of device wake-up in a multi-device scenario, and the application effect is relatively good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technologies, and in particular, to a voice wake-up method and an electronic device. Background Art

[0002] Currently, users can wake up an electronic device by speaking a wake-up word, thereby realizing the interaction between the user and the electronic device. Usually, the wake-up word is pre-set by the user in the electronic device, or the wake-up word is set before the electronic device leaves the factory. In a multi-device scenario (such as a smart home scenario), for the convenience of memory, the user may set the same wake-up word for multiple devices. For example, the user sets the wake-up words for the smart screen, smart speaker, and smart switch to be "Xiaoyi Xiaoyi". As Figure 1 shown, assuming that the user only hopes to wake up the smart screen, but when the user says "Xiaoyi Xiaoyi", both the smart screen and the speaker are woken up and both give the voice response of "I'm here" to the user, which will cause trouble to the user and affect the user experience. Summary of the Invention

[0003] This application provides a voice wake-up method and an electronic device, which helps to improve the accuracy of waking up an electronic device by voice in a multi-device scenario, thereby improving the user experience.

[0004] In a first aspect, an embodiment of this application provides a voice wake-up method, which can be applied to a first electronic device and relates to the field of terminal artificial intelligence (AI). The method includes:

[0005] The first electronic device receives a voice wake-up instruction from the user; the first electronic device also acquires a user image and detects the orientation of the user's face; then, the first electronic device determines a target device facing the user's face from the first electronic device and at least one second electronic device according to the relative positions of the first electronic device and the at least one second electronic device, the user's position, and the orientation of the user's face; finally, the first electronic device instructs the target device to respond to the voice wake-up instruction.

[0006] Among them, the first electronic device may have an image acquisition function, and the first electronic device acquires a user image from an image acquisition module. The first electronic device may also not have an image acquisition function, and the first electronic device acquires a user image from a second electronic device.

[0007] In an embodiment of this application, the first electronic device can determine the device that the user wants to wake up through the relative positions of the first electronic device and at least one second electronic device and the orientation of the user's face collected by the device, which helps to improve the accuracy of device wake-up in a multi-device scenario, and the application effect is relatively good.

[0008] In a possible design, when the second electronic device determines that the number of candidate devices facing the user's face is greater than or equal to two, the first electronic device needs to determine the relative distances between the user and the at least two candidate devices; then, based on the relative distances, determine the priorities of the candidate devices, where the smaller the relative distance of a candidate device, the lower its priority; finally, determine the candidate device corresponding to the highest priority as the target device.

[0009] In the embodiments of the present application, on the basis of directional wake-up, it helps to improve the accuracy of proximity wake-up in a multi-device scenario.

[0010] In another possible design, when the second electronic device determines that the number of candidate devices facing the user's face is greater than or equal to two, the first electronic device needs to determine the relative distances between the user and the at least two candidate devices; finally, determine the candidate device corresponding to the minimum relative distance as the target device.

[0011] In the embodiments of the present application, on the basis of directional wake-up, it helps to improve the accuracy of proximity wake-up in a multi-device scenario.

[0012] In a possible design, the first electronic device can obtain information about the first audio of the first electronic device and obtain information about the second audio from at least one second electronic device; then, based on the information about the first audio and the information about the second audio, determine the user's location.

[0013] In the embodiments of the present application, the sound collected by the multi-microphone array of the electronic device can effectively determine the user's location and ensure the accuracy of the user location positioning result.

[0014] In a possible design, the first electronic device includes a first microphone and a second microphone; the information about the first audio includes: the first arrival time when the voice wake-up instruction reaches the first microphone, the second arrival time when the voice wake-up instruction reaches the second microphone, and the first phase when the voice wake-up instruction reaches the first microphone and the second phase when the voice wake-up instruction reaches the second microphone; at least one second electronic device includes a third microphone and a fourth microphone; the information about the second audio includes: the third arrival time when the voice wake-up instruction reaches the third microphone, the fourth arrival time when the voice wake-up instruction reaches the fourth microphone, and the third phase when the voice wake-up instruction reaches the third microphone and the fourth phase when the voice wake-up instruction reaches the fourth microphone.

[0015] In a possible design, the first electronic device can determine the user's location according to the information about the first audio and the information about the second audio, specifically including the following steps:

[0016] Determine the relative distance between the user and the first electronic device according to the time difference between the first arrival time and the second arrival time, and determine the azimuth angle of the user relative to the first electronic device according to the phase difference between the first phase and the second phase; determine the relative distance between the user and at least one second electronic device according to the time difference between the second arrival time and the third arrival time; and determine the azimuth angle of the user relative to at least one second electronic device according to the phase difference between the third phase and the fourth phase; determine the user position according to the relative distance between the user and the first electronic device, the azimuth angle of the user relative to the first electronic device, the relative distance between the user and at least one second electronic device, and the azimuth angle of the user relative to at least one second electronic device.

[0017] In the embodiments of the present application, the sound collected by the multi-microphone array of the electronic device can effectively determine the user position and ensure the accuracy of the user position positioning result.

[0018] In a possible design, the method further includes: the first electronic device obtains information of historical audio from the first electronic device and the at least one second electronic device;

[0019] The first electronic device obtains the arrival time and phase of the voice wake-up instruction issued by the user N times to different electronic devices from the information of the historical audio, where N is a positive integer; then the first electronic device determines the relative azimuth angle and distance difference corresponding to the voice wake-up instruction issued by the user N times according to the arrival time and phase of the voice wake-up instruction issued by the user N times to different electronic devices;

[0020] The first electronic device establishes an objective function with the relative azimuth angle and distance difference corresponding to the voice wake-up instruction issued by the user N times as observation values; the first electronic device solves the objective function by an exhaustive search method to obtain the relative positions of the first electronic device and at least one second electronic device.

[0021] In the embodiments of the present application, the first electronic device can locate the relative positions of multiple devices in space according to the above method and construct a position map including the relative positions of the devices. So that the devices can also reverse-deduce the relative positions between multiple sound pickup devices through multiple voice recognitions of the user voice in an interference environment. As the number of voice wake-up messages increases, the positioning between the devices will become more accurate.

[0022] In a possible design, the first electronic device and the at least one second electronic device are connected to the same local area network, or the first electronic device and the at least one second electronic device are pre-bound with the same user account, or the first electronic device and the at least one second electronic device are bound with different user accounts, and different user accounts establish a binding relationship.

[0023] In a second aspect, an embodiment of the present application provides a voice wake-up method, which can be applied to a second electronic device. The method includes:

[0024] The second electronic device collects the sound of the surrounding environment and converts it into a second audio. Then, the second electronic device sends the second audio to the first electronic device. When the second electronic device detects a wake-up word in the second audio, it sends a wake-up message to the first electronic device. The first electronic device can determine the user's location based on the information of the first audio and the information of the second audio of the first electronic device. When it is determined that the target device is the second electronic device according to the relative positions of the first electronic device and at least one second electronic device, the user's location, and the user's face orientation, the first electronic device sends a wake-up response to the second electronic device. After receiving the wake-up response from the first electronic device, the second electronic device responds to the user's voice wake-up instruction.

[0025] In another possible case, if the first electronic device determines that the target device is not the second electronic device according to the relative positions of the first electronic device and at least one second electronic device, the user's location, and the user's face orientation, the first electronic device does not send a wake-up response to the second electronic device, or sends a response prohibiting wake-up to the second electronic device, and the second electronic device does not respond to the user's voice wake-up instruction.

[0026] In a third aspect, the present application provides a voice wake-up system, which includes a first electronic device and at least one second electronic device. The first electronic device can implement the method of any possible implementation manner of the first aspect above, and at least two second electronic devices can implement the method of any possible implementation manner of the second aspect above.

[0027] In a fourth aspect, an electronic device provided by an embodiment of the present application includes: one or more processors and a memory, where program instructions are stored in the memory. When the program instructions are executed by the device, the methods of the above-mentioned various aspects of the embodiments of the present application and any possible design involved in each aspect are implemented.

[0028] In a fifth aspect, a chip system provided by an embodiment of the present application is coupled to the memory in the electronic device, so that when the chip system runs, it calls the program instructions stored in the memory to implement the methods of the above-mentioned various aspects of the embodiments of the present application and any possible design involved in each aspect.

[0029] In a sixth aspect, a computer-readable storage medium of an embodiment of the present application stores program instructions. When the program instructions run on an electronic device, the device is caused to execute the methods of the above-mentioned various aspects of the embodiments of the present application and any possible design involved in each aspect.

[0030] In a seventh aspect, a computer program product according to an embodiment of the present application, when the computer program product runs on an electronic device, enables the electronic device to execute a method for implementing each of the above aspects of the embodiments of the present application and any possible design involved in each aspect.

[0031] In addition, for the technical effects brought about by any possible design manner in the fourth aspect to the seventh aspect, reference may be made to the technical effects brought about by different design manners in the relevant method part, which will not be elaborated herein. Description of the Drawings

[0032] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0033] Figure 2 It is a schematic diagram of the structure of a mobile phone provided by an embodiment of the present application;

[0034] Figure 3 It is another schematic diagram of an application scenario provided by an embodiment of the present application;

[0035] Figure 4 It is a schematic diagram of the interaction of a voice wake-up method provided by an embodiment of the present application;

[0036] Figure 5 It is a schematic diagram of a wake-up method provided by an embodiment of the present application;

[0037] Figure 6A It is a schematic diagram of a user location positioning method provided by an embodiment of the present application;

[0038] Figures 6B to 6D It is another schematic diagram of an application scenario provided by an embodiment of the present application;

[0039] Figure 7 It is another schematic diagram of an application scenario provided by an embodiment of the present application;

[0040] Figure 8A It is a schematic diagram of a device location positioning method provided by an embodiment of the present application;

[0041] Figure 8B It is a schematic diagram of a wake-up voice analysis method provided by an embodiment of the present application;

[0042] Figure 8C It is a schematic diagram of a device map provided by an embodiment of the present application;

[0043] Figure 9 It is a schematic diagram of the interaction of a device location positioning method provided by an embodiment of the present application;

[0044] Figure 10 It is a schematic diagram of a set of perception ability layers provided by an embodiment of the present application;

[0045] Figure 11 It is a schematic structural diagram of a device according to an embodiment of the present application;

[0046] Figure 12 It is a schematic structural diagram of another device according to an embodiment of the present application. Detailed implementation manners

[0047] It should be understood that in the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; "and / or" in the present application is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. "At least one" means one or more, and "a plurality" means two or more.

[0048] In the present application, terms such as "exemplary", "in some embodiments", "in other embodiments" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the term "exemplary" is intended to present concepts in a specific manner.

[0049] In addition, terms such as "first" and "second" involved in the present application are only used for the purpose of distinguishing descriptions, and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features, nor can they be construed as indicating or implying an order.

[0050] The electronic device in the embodiment of the present application is an electronic device with a voice wake-up function, that is, the user can wake up the electronic device by voice. Specifically, the user wakes up the electronic device by saying a wake-up word. Among them, the wake-up word can be set in the electronic device by the user according to his own needs in advance, or can be set by the electronic device before leaving the factory. The present application does not limit the setting method of the wake-up word. It should be noted that the user who wakes up the electronic device in the embodiment of the present application can be any user or a specific user. Exemplarily, a specific user can be a user who has stored the voice of the wake-up word in the electronic device in advance, such as the owner of the device.

[0051] Currently, an electronic device triggers device wake-up by detecting whether a wake-up word is included in the audio. Specifically, when the wake-up word is included in the audio, the electronic device is woken up; otherwise, the electronic device is not woken up. After the electronic device is woken up, the user can interact with the electronic device through voice. For example, the wake-up word is "Xiaoyi Xiaoyi". When the electronic device detects that the audio includes "Xiaoyi Xiaoyi", the electronic device is woken up. Among them, the electronic device acquires or receives ambient sound through a multi-microphone array on the device to obtain audio. However, in a multi-device scenario (such as a smart home scenario), the voice including the "wake-up word" spoken by the user may be received or acquired by multiple electronic devices, resulting in two or more electronic devices being woken up, which brings confusion to the user's voice interaction process and affects the user experience.

[0052] Currently, to solve the problem of devices being accidentally woken up in a multi-device scenario, the priorities of each device are usually artificially specified. Assume that it is artificially pre-specified that Figure 1 the smart screen in has a higher priority than the smart speaker. Then, when both the smart screen and the smart speaker collect the user's "Xiaoyi Xiaoyi", only the smart screen is woken up. Although this method can constrain the situation of multiple devices being woken up by setting rules, because the rules need to be artificially set in advance, it is not intelligent enough, and the user can only actively modify the setting rules manually to adjust the priority of the woken-up device according to actual needs, so there is a problem of poor flexibility.

[0053] In addition, there is also a related technology that triggers the device facing the face to be woken up by monitoring the face orientation. Although this method can improve the flexibility and accuracy of the human-computer interaction method. However, since most devices do not have an image acquisition function, the application effect is not good. For example: Devices such as smart speakers and smart voice-controlled switches in a smart home scenario are generally not integrated with an image acquisition module due to cost constraints, so the method of directionally waking up devices based on face orientation cannot be applied to such devices.

[0054] Therefore, when the existing voice wake-up method is applied to a multi-device scenario, it still cannot effectively solve the problem of multiple devices being woken up simultaneously. In view of this, the embodiments of the present application provide a voice wake-up method. On the one hand, it can locate the relative positions of the user and multiple devices in space and construct a position map; on the other hand, it can use the master device with an image acquisition module among the multiple devices to collect the user's face orientation. In this way, by combining the position map and the user's face orientation collected by the master device, the device that the user wants to wake up can be determined. This method helps to improve the accuracy of device wake-up in a multi-device scenario, and the application effect is relatively good.

[0055] Below, taking an electronic device as an example, Figure 2 the structural schematic diagram of the electronic device 200 is shown.

[0056] The voice wake-up method provided by the embodiments of the present application can be applied to an electronic device. In some embodiments, the electronic device may be a portable terminal including functions such as a personal digital assistant and / or a music player, such as a mobile phone, a tablet computer, a wearable device with wireless communication function (such as a smart watch), a vehicle-mounted device, etc. Exemplary embodiments of the portable terminal include, but are not limited to, a portable terminal equipped with the Harmony operating system (OS). Or a portable terminal with other operating systems. The above portable terminal may also be a laptop computer (Laptop) with a touch-sensitive surface (such as a touch panel), etc. It should also be understood that in some other embodiments, the above terminal may also be a desktop computer with a touch-sensitive surface (such as a touch panel).

[0057] Figure 2 A schematic structural diagram of the electronic device 200 is shown.

[0058] The electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0059] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0060] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0061] The electronic device 200 implements the display function through the GPU, the display screen 294, and the application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 294 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change the display information.

[0062] The electronic device 200 can implement the shooting function through the ISP, the camera 293, the video codec, the GPU, the display screen 294, and the application processor, etc.

[0063] The SIM card interface 295 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to achieve contact and separation from the electronic device 200. The electronic device 200 may support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 may support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 295 at the same time. The types of the multiple cards may be the same or different. The SIM card interface 295 can also be compatible with different types of SIM cards. The SIM card interface 295 can also be compatible with external memory cards. The electronic device 200 interacts with the network through the SIM card to achieve functions such as calls and data communication. In some embodiments, the electronic device 200 uses an eSIM, that is, an embedded SIM card.

[0064] The wireless communication function of the electronic device 200 can be implemented by antenna 1, antenna 2, the mobile communication module 250, the wireless communication module 260, the modulation and demodulation processor, the baseband processor, etc. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0065] The mobile communication module 250 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the electronic device 200. The mobile communication module 250 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves through antenna 1, filter, amplify, etc. the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 250 can be disposed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 can be disposed in the same device.

[0066] The wireless communication module 260 can provide solutions for wireless communications such as wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared radiation (IR) technology, etc. applied to the electronic device 200. The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves through antenna 2, frequency-modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 260 can also receive the signals to be sent from the processor 210, frequency-modulate and amplify them, and convert them into electromagnetic waves through antenna 2 for radiation.

[0067] In some embodiments, the antenna 1 of the electronic device 200 is coupled to the mobile communication module 250, and the antenna 2 is coupled to the wireless communication module 260, so that the electronic device 200 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc.

[0068] It can be understood that Figure 2 the components shown do not constitute a specific limitation on the electronic device 200. The electronic device 200 may further include more or fewer components than those shown, or combine certain components, or split certain components, or have different component arrangements. In addition, Figure 2 the combination / connection relationship between the components in

[0069] The software system of the electronic device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, the layered architecture is taken as an example. Among them, the layered architecture may include the Harmony operating system (OS) or other operating systems. The voice wake-up method provided in the embodiments of the present application can be applied to terminals integrated with the above operating systems.

[0070] The above Figure 2 are respectively the hardware structures of the electronic devices applicable to the embodiments of the present application. To solve the problems raised in the background art, the embodiments of the present application provide a voice wake-up method. This method can accurately determine the device that the user wants to wake up by using a location map including the user's location and the device's location and the face orientation of the user collected by the device, and improve the accuracy of the pointing device wake-up result in a multi-device scenario.

[0071] For example, as Figure 3 shown, it is a schematic diagram of a multi-device scenario applicable to the embodiments of the present application. Specifically, in Figure 3In the multi-device scenario shown, it is assumed that the electronic device 10, the electronic device 20, and the electronic device 30 are all sound pickup devices equipped with multi-microphone arrays, and the same wake-up word is preset for all of them. For example, the wake-up word is "Xiaoyi Xiaoyi". When the user says "Xiaoyi Xiaoyi", the electronic device 10, the electronic device 20, and the electronic device 30 can all collect or receive this voice. On the one hand, the electronic device 30 can use this wake-up voice to determine the user's position in the device map obtained through pre-training; on the other hand, the electronic device 30 can use the face information collected by image acquisition to determine the face orientation, and thus determine which one of the electronic device 10, the electronic device 20, and the electronic device 30 is the target device facing the face according to the face orientation and the position map including the user's position.

[0072] It should be noted that Figure 3 This is only an example of a multi-device scenario. The embodiments of the present application do not limit the number of electronic devices in the multi-device scenario, nor do they limit the wake-up words preset in the electronic devices. In addition, it should be noted that in other possible cases, it may also be that the electronic device 30 does not collect face images, but obtains the face image acquisition results from other devices (such as the electronic device 10 or the electronic device 20). The electronic device 30 can be a central device with strong data processing capabilities, such as a smart speaker or a smart screen in a smart home scenario. For the convenience of description below, it is described with the electronic device 30 having a face image acquisition function and being a central device.

[0073] Combined with Figure 3 the multi-device scenario shown, a voice wake-up method according to an embodiment of the present application will be specifically described for the voice wake-up method according to an embodiment of the present application. As Figure 4 shown, the method flow specifically includes the following steps.

[0074] Steps 401a to 401c, the electronic device 10, the electronic device 20, and the electronic device 30 all collect the surrounding environmental sounds in real time and convert the collected surrounding environmental sounds into audio.

[0075] Specifically, the multi-microphone array of the electronic device 10 collects the surrounding environmental sounds and converts the collected surrounding environmental sounds into audio. The multi-microphone array of the electronic device 20 collects the surrounding environmental sounds and converts the collected surrounding environmental sounds into audio. The multi-microphone array of the electronic device 30 collects the surrounding environmental sounds and converts the collected surrounding environmental sounds into audio. At this time, if the user issues a wake-up voice, for example, when the user says "Xiaoyi Xiaoyi", the voices collected by the electronic device 10, the electronic device 20, and the electronic device 30 will include the user's wake-up voice.

[0076] Exemplarily, combined with Figure 1 it is said that the user is facing Figure 1The smart screen shown issues a voice wake-up command of "Xiaoyi Xiaoyi". Since the smart screen, smart speaker, and smart switch all collect the surrounding ambient sound in real time and convert it into audio, the sound collected by the smart screen, smart speaker, and smart switch will include the user's wake-up voice.

[0077] In steps 402a to 402c, the electronic device 10, the electronic device 20, and the electronic device 30 all perform wake-word detection on the generated audio.

[0078] Specifically, the electronic device 10, the electronic device 20, and the electronic device 30 can perform one-dimensional convolution on the data within each sliding window of the audio collected by their own devices to extract the features of different frequency bands in the data. Thus, when an audio segment with the same preset voice features as the user is recognized from the audio, it indicates that the audio segment includes the wake word; otherwise, it does not include the wake word.

[0079] In steps 403a to 403b, when the electronic device 10 and the electronic device 20 detect the wake word, they both send a wake-up message to the electronic device 30 and carry the information of the audio generated by their own devices while sending the wake-up message.

[0080] Specifically, when the electronic device 10 detects the wake word, the electronic device 10 sends a first wake-up message and the information of the audio to the electronic device 30, and the first wake-up message is used to request confirmation of whether to wake up the electronic device 10; when the electronic device 20 detects the wake word, the electronic device 20 sends a second wake-up message and the information of the audio to the electronic device 30, and the second wake-up message is used to request confirmation of whether to wake up the electronic device 20. Among them, the information of the audio can include all the data of the audio, or the information of the audio can include information such as the arrival time and phase related to the voice wake-up command.

[0081] For example, the information of the first audio generated by the electronic device 30 includes: the first arrival time (time of arrival) of the voice wake-up command at the first microphone, the second arrival time of the voice wake-up command at the second microphone, and the first phase of the voice wake-up command at the first microphone and the second phase of the voice wake-up command at the second microphone. It should be noted that the first arrival time refers to the moment when the first microphone first picks up the voice wake-up command, and the second arrival time refers to the moment when the second microphone first picks up the voice wake-up command.

[0082] The electronic device 20 includes a third microphone and a fourth microphone; the information of the second audio generated by the electronic device 20 includes: the third arrival time when the voice wake-up instruction arrives at the third microphone, the fourth arrival time when the voice wake-up instruction arrives at the fourth microphone, as well as the third phase when the voice wake-up instruction arrives at the third microphone and the fourth phase when the voice wake-up instruction arrives at the fourth microphone. It should be noted that the third arrival time refers to the time when the third microphone first picks up the voice wake-up instruction, and the fourth arrival time refers to the time when the fourth microphone first picks up the voice wake-up instruction.

[0083] It should be noted that the electronic devices 10, 20, and 30 may also include other microphones. The embodiments of the present application do not limit the number of microphones, and other microphones can also collect sounds according to the above method.

[0084] In addition, it should be noted that if the electronic device 10 or the electronic device 20 does not detect the wake-up word, there is no need to send a wake-up message to the electronic device 30, and the electronic device 10 or the electronic device 20 only needs to send the information of the audio to the electronic device 30. Similarly, if the electronic device 20 does not detect the wake-up word, there is no need to send a wake-up message to the electronic device 30, and the electronic device 20 only needs to send the information of the audio to the electronic device 30. This embodiment will not be illustrated again. Figure 1 Shown as one.

[0085] Step 403c, when the electronic device 30 also detects the wake-up word, a third wake-up message is also generated; otherwise, no third wake-up message is generated. The third wake-up message is used to request confirmation of whether to wake up the electronic device 30.

[0086] Step 404, the electronic device 30 determines the relative position of the user in the device map according to the pre-trained device map and the audio collected by any two of the electronic devices 10, 20, and 30, so as to generate a position map including the user's position.

[0087] Among them, the electronic devices synchronize information between multiple devices based on multi-device interconnection technology (such as HiLink (a multi-device interconnection technology)). Specifically, the electronic devices 10, 20, and 30 can be connected to the same local area network to achieve mutual communication. Or the electronic devices 10, 20, and 30 can be pre-bound with the same user account (such as a Huawei account). Or the electronic devices 10, 20, and 30 are bound with different user accounts, and different user accounts have a binding relationship (such as pre-binding the user accounts of family members, that is, authorizing one's own device to be able to connect with the devices of family members) to ensure secure communication between devices.

[0088] Exemplarily, such as Figure 5As shown, both the first microphone and the second microphone of the electronic device 30 collect sound and record the information of the first audio. On the one hand, the electronic device 30 can determine the direction angle between the electronic device 30 and the user according to the phase difference between the first phase of the voice wake-up instruction reaching the first microphone of the electronic device 30 and the second phase of the voice wake-up instruction reaching the second microphone of the electronic device 30; the electronic device 30 obtains the information of the second audio from the electronic device 20, so the electronic device 30 can determine the direction angle between the electronic device 20 and the user according to the phase difference between the third phase and the fourth phase of the electronic device 20. On the other hand, since the electronic device 30 can determine the first relative distance between the electronic device 30 and the user according to the time difference between the first arrival time and the second arrival time of the electronic device 30, furthermore, the electronic device 30 can determine the second relative distance between the electronic device 20 and the user according to the time difference between the third arrival time and the fourth arrival time of the electronic device 20. In this way, the user position can be determined by combining the first azimuth angle, the second azimuth angle, the first relative distance, and the second relative distance.

[0089] Exemplarily, as Figure 6A shown, point A refers to the position of the electronic device 20 in the pre-trained device map, point B refers to the position of the electronic device 30 in the pre-trained device map, θA is the first azimuth angle of the electronic device 20 relative to the user, θB is the second azimuth angle of the electronic device 30 relative to the user, PA is the second relative distance of the electronic device 20 relative to the user, and PB is the first relative distance of the electronic device 30 relative to the user. As can be seen from the figure, the intersection point P of the two azimuth angle rays of θA and θB is the position where the user issues the wake-up voice.

[0090] Similarly, according to the above method, the electronic device 30 can also determine the direction angle between the electronic device 20 and the user according to the phase difference between different microphones of the voice wake-up instruction reaching the electronic device 10.

[0091] It should be noted that in this embodiment, specifically, the TDoA (a positioning method using time difference) or MUSIC (a sound source localization method) algorithm can be used to calculate the position of the sound-emitting user and the distance from the device, and the embodiments of the present application do not limit this. The user position determined in the embodiments of the present application refers to the relative position. For example, the user to be located is due south of the electronic device 30 and the distance from the electronic device 30 is 1 meter.

[0092] Step 405, the electronic device 20 collects the user's image and detects the face orientation.

[0093] For example, the user is facing Figure 1The smart screen in it emits a wake-up voice of "Xiaoyi, Xiaoyi", and the camera on the smart screen takes pictures or videos of the user. The smart screen analyzes the face image to determine that the user's face is facing the smart screen. For another example, the user is facing Figure 1 The smart screen next to the smart speaker in it emits a wake-up voice of "Xiaoyi, Xiaoyi", and the camera on the smart screen takes pictures or videos of the user. The smart screen analyzes the face image to determine that the user's face is facing the first azimuth angle (for example, the first azimuth angle is the front left of the user).

[0094] It should be noted that in this embodiment, the smart screen is taken as the main control device (or the central device) for illustration. Since the smart screen has an image acquisition function, the user image collected by the smart screen is preferentially used for face orientation detection. If the main control device (or the central device) does not have an image acquisition function, the user image can also be obtained from other electronic devices with a face acquisition function, and then the obtained user image is analyzed. Examples are not shown one by one here.

[0095] Step 406, the electronic device 30 determines, according to the face orientation and the position map including the user's position, that the target device facing the user's face is the electronic device 10.

[0096] Exemplarily, the one determined by the electronic device 30 and Figure 3 The position map corresponding to the multi-device scenario shown is as Figure 6B shown. In this figure, at the position where the user is located, the field of view range corresponding to the user's face orientation is the range shown by the θ angle in the figure, and the electronic device 10 exists within this range.

[0097] Step 407a, when the electronic device 30 determines that the target device is the electronic device 10, it sends an instruction allowing wake-up to the electronic device to instruct the electronic device 10 to respond to the user's wake-up voice.

[0098] In a possible embodiment, this embodiment may further include the following steps: Step 407a, when the electronic device 30 determines that the target device is the electronic device 10, it may also instruct that the electronic device 30 itself is not woken up, and Step 407b, the electronic device 30 sends an instruction prohibiting wake-up to the electronic device 20, and this instruction prohibiting wake-up is used to instruct the electronic device 20 not to respond to the user's wake-up voice.

[0099] Alternatively, in another possible embodiment, when the electronic device 30 determines that the target device is the electronic device 10, it may not respond to the electronic device 30 itself and the electronic device 20. In this way, the electronic device 30 and the electronic device 20 do not respond to the user's wake-up voice either.

[0100] In summary, the first electronic device (such as electronic device 30) receives a voice wake-up command from the user; the first electronic device 30 obtains a user image from a local or other device and detects the orientation of the user's face; then, based on the relative positions of the first electronic device and at least one second electronic device (such as electronic device 10 and electronic device 20), the user's position (such as the position of point P), and the orientation of the user's face, the target device (such as target device is electronic device 10) that the user's face is facing is determined from the first electronic device and at least one second electronic device, and the target device is instructed to respond to the voice wake-up command. It can be seen that through the above method, the problem of "one call, multiple responses" or "one call, multiple rings" can be effectively improved, so that when the user issues a wake-up command facing electronic device 10, electronic device 10 will make a voice interaction response to it, while other devices will not make a response.

[0101] In this embodiment, it is not required that the electronic devices in the multi-device scenario have an image acquisition function. That is to say, in the multi-device scenario, as long as one device has an image acquisition function and can capture a face image, the device pointed to by the user can be determined by combining the position map and the face image, and the device pointed to can completely not have an image acquisition function. Therefore, to a certain extent, the application scenario of this method is wider.

[0102] In addition, in this embodiment, when the user moves to other positions in the scenario and issues a wake-up voice, the method shown in steps 401a to 404 can still be used to re-locate the user's current position, so that the position after the user moves can still be accurately located in combination with the captured face image. Exemplarily, when the user walks indoors to near the smart speaker, electronic device 30 can locate the user's position moving from position P1 to position P2 according to the above method. It can be seen that this method can real-time locate the current position of the user making the sound. Further, the target device can also be determined by combining the currently captured face image and the user's current position P2 as Figure 6C the shown electronic device 20. It can be seen that the device to be woken up can always be directionally located according to the above method, regardless of whether the user's position has changed.

[0103] Furthermore, in a possible embodiment, when there are more than two candidate devices determined for the device towards which the user's face is facing by combining the user's location and the face orientation, at this time, the relative distance between the candidate devices and the user's location can be further combined to determine the target device, that is, the candidate device with the shortest relative distance is selected from more than two candidate devices as the target device. That is to say, the electronic device 30 calculates the relative distance between the candidate devices within the face orientation range and the user's location, and then determines the wake-up priority of each candidate device according to the magnitude of the relative distance. The smaller the relative distance of a device, the higher the wake-up priority. On the contrary, the larger the relative distance of a device, the lower the wake-up priority. In this way, the electronic device 30 can wake up the candidate device with the highest wake-up priority directionally. Exemplarily, as Figure 6D shown, when the user walks indoors near the smart speaker, the electronic device 30 can locate the user's location from P1 to P2 according to the above method, and then combine the currently collected face image and the user's current location P2 to determine that the candidate devices within the visual field corresponding to the face orientation are the electronic device 20 (the smart speaker shown in the figure) and the electronic device 40 (the smart alarm clock shown in the figure). The electronic device 30 determines that the relative distance between the electronic device 20 and P2 is D2, and determines that the relative distance between the electronic device 40 and P2 is D1. Since D1 is greater than D2, the wake-up priority of the electronic device 20 is high. Therefore, the electronic device 30 determines the electronic device 10 as the target device to be woken up.

[0104] It can be seen that in the above method, it is necessary to rely on a pre-trained device map. That is to say, first, it is necessary to construct a relative position map of each device in a multi-device scenario to determine the user's relative position in the device map. For this reason, the embodiments of the present application provide a method for training a device map. This method performs speech analysis on the user's historical wake-up speech, thereby inversely deducing the relative positions between multiple sound pickup devices. The following combines Figures 7 to 9 to exemplarily give the construction method of the device map.

[0105] Exemplarily, as Figure 7 shown, it is a schematic diagram of a multi-device scenario provided by an embodiment of the present application. There are multiple sound pickup devices deployed in this scenario. Among them, the main device with an image acquisition module is the smart screen device 71 in the figure. There are also multiple sound pickup devices 72a to 72f deployed in the space where the smart screen device 71 is located. It should be noted that in the figure, multiple speakers are used to illustrate the sound pickup devices 72a to 72f. In an actual home scenario, the speakers can also be replaced with devices such as smart alarm clocks, smart cameras, and smart switches, which will not be illustrated one by one with legends here. Only speakers are used in the figure to illustrate the sound pickup devices around the smart screen device 71.

[0106] In Figure 7In the smart home scenario shown, the user may wake up any one of the sound pickup devices at any position in this scenario. For example, the user may sit on the sofa and call "Xiaoyi Xiaoyi" to wake up the smart screen device. Another example is that the user faces the sound pickup device 72a (such as a speaker) and calls "Xiaoyi Xiaoyi" to wake up the sound pickup device 72a. By analogy, during the user's use, each sound pickup device records the user's historical wake-up voice and synchronizes it to the central device. The central device can be Figure 7 the smart screen device shown, or it can also be a smart speaker or a router, etc.

[0107] Such as Figure 8A shown, assume that the sound generation position when the user calls "Xiaoyi Xiaoyi" for the first time is the sound generation point P1, and the sound generation position when the user calls "Xiaoyi Xiaoyi" for the second time is the sound generation point P2. Of course, during the device usage period, the user may also call other sound pickup devices at other positions. Here, take the case where the wake-up voice wakes up the sound pickup device 72a and the sound pickup device 72b twice respectively as an example. The sound pickup device 72a and the sound pickup device 72b convert the collected sound into audio and then synchronize it to the central device. The central device can perform voiceprint analysis on the sounds collected by the sound pickup device 72a and the sound pickup device 72b respectively. Specifically, the sound pickup device 72a and the sound pickup device 72b can perform one-dimensional convolution on the data in each sliding window of the audio collected by their own devices, and extract the voiceprint features of different frequency bands from the data. Thus, when an audio segment with the same voiceprint feature as the preset wake-up voice is recognized from the audio, it means that this audio segment includes the wake-up word issued by this user; otherwise, it does not include the wake-up word issued by this user. Such as Figure 8B shown, according to the above voice detection method, the arrival time t1 of the wake-up voice issued by the user at point P1 reaching the sound pickup device 72a, and the arrival time t2 of the wake-up voice issued by the user at point P1 reaching the sound pickup device 72a are obtained. The audio voiceprints within the Δt shown in the figure are the same, which is the wake-up voice of the same user, so as to obtain the arrival time difference Δτ1 = t2 - t1.

[0108] In a possible embodiment, in some special cases, there may be multiple users simultaneously issuing the same wake-up voice in the same period. For example, the father and the child simultaneously call "Xiaoyi Xiaoyi" in the living room. At this time, although the audio segment corresponding to the wake-up voice in this period can be detected according to the above method, because there are multiple voiceprints and the audio overlaps during this time period, it is not conducive to accurately determining the phase of the wake-up voice. Therefore, the central device can further screen the obtained audio, and screen out the audio with overlapping audio during the same period, so as to improve the accuracy of the device map calculation result.

[0109] Specifically, a polar coordinate system A is established on the sound pickup device 72a, and a polar coordinate system B is established on the sound pickup device 72b. A rectangular coordinate system C is established with points A and B as the X-axis. Such asFigure 8A As shown is the angle between the polar coordinate system A and the rectangular coordinate system C; θA1 is the angle between the sound source point P1 and point A and the x-axis of the rectangular coordinate system C; θB1 is the angle between the sound source point P1 and point B and the x-axis of the rectangular coordinate system C; dA1 is the distance from the sound source point P1 to point A; dB1 is the distance from the sound source point P1 to point B. Additionally, is the included angle between the polar coordinate system B and the rectangular coordinate system C; θA2 is the angle between the sound source point P2 and point A and the x-axis of the rectangular coordinate system C; θB2 is the angle between the sound source point P2 and point B and the x-axis of the rectangular coordinate system C; dA2 is the distance from the sound source point P2 to point A; dB2 is the distance from the sound source point P2 to point B.

[0110] In this embodiment, by using the time difference Δτ1 between the wake-up signal emitted by the user at point P1 and the polar coordinate system A and the polar coordinate system B (i.e., the sound pickup devices 72a and 72b), the distance difference dA1 - dB1 can be calculated, that is:

[0111] dA1 - dB1 = Δτ1·v = a1...... Formula 1

[0112] where dA1 is the distance from the sound source point P1 to point A; dB1 is the distance from the sound source point P1 to point B, Δτ1 is the time difference between the wake-up signal emitted by the user at point P1 and the polar coordinate system A and the polar coordinate system B, that is, Δτ1 is the time difference between the wake-up signal emitted by the user and two different electronic devices, v is the sound transmission speed, and a1 is the distance difference dA1 - dB1, that is, a1 is the distance difference between the user's sound source position and two different electronic devices.

[0113] Additionally, by using the phase of the wake-up signal emitted by the user at point P1 to the polar coordinate system A and the phase of the wake-up signal emitted by the user at point P1 to the polar coordinate system B, the following formula can be obtained:

[0114] 180° - θA1 - θB1 = b1...... Formula 2

[0115] where θA1 is the angle between the sound source point P1 and point A and the x-axis of the rectangular coordinate system C; θB1 is the angle between the sound source point P1 and point B and the x-axis of the rectangular coordinate system C, and b1 is the included angle between the distance from P1 to point A and the distance from P1 to point B.

[0116] Specifically, θA1 and θB1 can be obtained by using Figure 6A the time difference and phase of the wake-up voice reaching different microphones as shown, and the calculation process is not repeated here.

[0117] Additionally, and The difference between them is an unknown constant, which is used to characterize the included angle between the sound pickup direction corresponding to the sound pickup device 72a and the sound pickup direction corresponding to the sound pickup device 72b. That is and satisfy the following formula:

[0118]

[0119] Furthermore, using the cosine theorem, the following formula can be obtained:

[0120]

[0121] Further deriving the above formula four, the following formula five can be obtained:

[0122]

[0123] Combining the above formula one, formula two, formula three and formula five, and using the following letters to simplify the constants, the following equation can be obtained:

[0124]

[0125] Among them,

[0126] Using formula two and formula three to further simplify, the following formula six is obtained, where x1 is the unknown, representing dA1.

[0127] And so on. Since the sound generation position when the user calls "Xiaoyi Xiaoyi" for the second time is the sound generation point P2, based on that θA2 is the included angle between the sound generation point P2 and point A and the x-axis of the rectangular coordinate system C; θB2 is the included angle between the sound generation point P2 and point B and the x-axis of the rectangular coordinate system C; dA2 is the distance from the sound generation point P2 to point A; dB2 is the distance from the sound generation point P2 to point B, according to the above calculation method, the following equations can be combined:

[0128]

[0129] Similarly, it can be obtained where x2 is the unknown, representing dA2.

[0130] And so on. Since the sound generation position when the user calls "Xiaoyi Xiaoyi" for the nth time is the sound generation point Pn, based on that θAn is the included angle between the sound generation point Pn and point A and the x-axis of the rectangular coordinate system C; θBn is the included angle between the sound generation point Pn and point B and the x-axis of the rectangular coordinate system C; dAn is the distance from the sound generation point Pn to point A; dBn is the distance from the sound generation point Pn to point B, according to the above calculation method, the following equations can be combined:

[0131]

[0132] Similarly, we can get Among them, xn is an unknown number, which refers to dAn.

[0133] That is, based on the user's historical wake-up voice, the following equations can be constructed using the above method:

[0134]

[0135] Use letters to simplify constants and combine the simplified equations:

[0136]

[0137] In this way, by using the analysis results of a large number of historical wake-up voices, multiple equations can be jointly formulated and solved by algorithms (such as numerical solutions) to obtain the optimal device distance and device angle, thereby calculating the relative position between the devices. For example, Figure 8C As shown. In other words, the side length calculated using the distance divergence attenuation formula is substituted into the above equation, and the device distance and the angle between devices are searched through the grid search method. The least squares method is used to determine the device spacing that minimizes the loss function of all equations. The device spacing and device angle that minimize the loss function are the precise relative positions of the devices obtained through offline learning at night. In this way, through the grid method and exhaustive search method, the optimal relative position between devices can be obtained. As the accumulated user wake-up voice history increases, the device will continue to record the call data when the user calls Xiaoyi. The device can use this saved data for offline learning and training at night, achieving a smarter effect with use, and the error calculated by the exhaustive search method will be smaller.

[0138] In summary, if Figure 9 As shown, a device map training method of an embodiment of the present application is used to locate the relative position of each device in a multi-device scene. Figure 9 As shown, the method flow specifically includes the following steps.

[0139] In steps 901a to 901c, the electronic device 10, the electronic device 20, and the electronic device 30 all collect ambient sounds in real time and convert the collected ambient sounds into audio.

[0140] Specifically, the multi-microphone array of the electronic device 10 collects the ambient sound, and converts the collected ambient sound into audio. The multi-microphone array of the electronic device 20 collects the ambient sound, and converts the collected ambient sound into audio. The multi-microphone array of the electronic device 30 collects the ambient sound, and converts the collected ambient sound into audio. At this time, if the user issues a wake-up voice, for example, when the user issues a voice command of "Xiaoyi Xiaoyi", the sound collected by the electronic device will include the user's wake-up voice.

[0141] Exemplarily, the user faces Figure 1 the shown smart screen and issues a wake-up voice of "Xiaoyi Xiaoyi". Since the smart screen, the smart speaker, and the smart switch all collect the ambient sound in real time and convert it into audio, the sound collected by the smart screen, the smart speaker, and the smart switch will include the user's wake-up voice.

[0142] In steps 902a to 902b, the electronic device 10 and the electronic device 20 synchronize the generated audio to the electronic device 30.

[0143] That is to say, the electronic device 10 synchronizes the generated audio to the electronic device 30; the electronic device 20 synchronizes the generated audio to the electronic device 30. It should be noted that the electronic device 10 and the electronic device 20 can synchronize the generated audio to the electronic device 30 at fixed intervals, or both can synchronize the generated audio to the electronic device 30 at fixed time points (such as 1 o'clock in the morning). In this regard, the present application does not make any limitations.

[0144] In addition, this embodiment takes the multi-device scenario including the electronic device 10, the electronic device 20, and the electronic device 30, and the electronic device 30 is the central device as an example for illustration. In other possible embodiments, there may also be other electronic devices, or the central device is other electronic devices, and so on. Other electronic devices can also refer to the above-mentioned electronic device 10 to electronic device 30, and will not be elaborated here one by one.

[0145] Step 903, the electronic device 30 analyzes the audio generated by the electronic device 10, the electronic device 20, and the electronic device 30 itself according to the method shown above Figure 8A and establishes a system of equations, and finally solves to obtain the relative position map of each electronic device.

[0146] In this embodiment, the calculation process of the relative positions between multiple devices relies on the information of historical wake-up voices. By obtaining the time difference information of the arrival times of the wake-up voices received by multiple devices multiple times, as well as the time difference of the arrival times of the wake-up voice at different microphones of the same device, the relative positions between the devices are obtained. Since the above method can continuously use the historical user wake-up voices in the recent period for device positioning, even during the use of the device, if the position of the device changes, for example, the smart speaker is moved from the living room to the dining room, according to the above method, by accumulating the information of the historical wake-up voices for a period of time after the smart speaker is moved, the latest relative positions between the devices can still be updated. That is to say, by positioning according to the above method, the latest position of the current device can be updated in real time, giving the user the feeling that the device becomes smarter with use.

[0147] In addition, in this embodiment, there is no requirement for the usage environment of the sound pickup devices. That is to say, even when the devices are in an environment with noise interference, the relative positions between multiple sound pickup devices can be deduced in reverse through multiple voice recognitions of the user's voice.

[0148] To implement the above training process of the device map, in the embodiments of the present application, in each device in the multi-device scenario, a multi-layer perception ability layer of the device is constructed, as Figure 10 shown, which specifically includes: a basic perception ability layer, a secondary perception ability layer, and a high-order perception ability layer.

[0149] Among them, the basic perception ability layer refers to the ability after some existing functions of devices or software are simply encapsulated by the intelligent perception framework. For example, the chip layer provides the ability of face orientation, the sound source localization ability of multi-microphone devices, and the call status monitoring ability provided by the Android layer, etc.

[0150] The secondary perception ability layer refers to the ability obtained after the calculation results obtained by processing the basic perception ability are encapsulated by the framework. For example, the positioning ability between devices and the user voice localization ability here are both calculation results obtained by processing the sound localization data (basic perception ability) reported by different underlying devices.

[0151] The high-order perception ability layer refers to the complex calculation results reported to the upper layer calculated through complex calculations, fusions, models, and rules. For example, the directional wake-up ability, which is an advanced ability calculated by fusing multiple basic perception abilities and secondary perception abilities.

[0152] In addition, during the software implementation process of this application, a fence mechanism is also provided, which refers to a virtual fence enclosing a virtual boundary. When the mobile phone enters or leaves a specific geographical area / a specific action, etc., the upper-layer application of the mobile phone can receive automatic notifications and warnings. The fence concept is the core of intelligent perception, and each capability will be paired with a corresponding fence to trigger. When a capability is paired with a fence, it forms a platform that provides the ability to perceive user behavior for the upper layer, and the upper-layer application can receive reports when the user enters and exits specific events.

[0153] As Figure 11 shown, it is a schematic diagram of software modules in a multi-device scenario provided by an embodiment of this application. The cooperation of the modules in slave device 1 to slave device n and the master device can implement the device map training method provided by the embodiment of this application. Among them, the master device refers to the central device for positioning the device location in the above text, such as electronic device 30, and the slave devices refer to the devices for collecting wake-up voices in the above text, such as electronic device 10 and electronic device 20, etc.

[0154] Figure 11 In the scenario shown, each slave device includes an audio collection module 1101, and the master device includes an audio collection module 1101, an audio processing module 1102, and an audio recognition module 1103.

[0155] Among them, the audio collection module 1101 is used to collect sounds in the environment using a multi-microphone array and convert the collected sounds into audio.

[0156] The audio collection module 1101 of each device can send the audio corresponding to each sampling period to the audio processing module 1102 of the master device (such as electronic device 30) for processing. Among them, the master device can be Figure 3 the smart screen device in

[0157] the smart home scenario, or a device with strong computing power in the smart home scenario, such as a router or a smart speaker, etc.

[0158] The audio processing module 1102 of the master device is used to preprocess the audio corresponding to each sampling period, such as channel conversion, smoothing processing, noise reduction processing, etc., so as to facilitate the subsequent detection of wake-up voices by the audio recognition module 1103.

[0159] The device map calculation module 1104 is configured to calculate the relative positions between different devices based on the arrival times of the same wake-up voice at different devices and the arrival times of the wake-up signal at different microphones of the same device.

[0160] In addition, the cooperation of the modules in the slave device 1 to the slave device n and the master device can implement the voice wake-up method provided in the embodiments of the present application. To implement this voice wake-up method, Figure 12 In the scene shown, each slave device may further include a user position positioning module 1105, a face orientation recognition module 1106, and a directional wake-up module 1107.

[0161] Among them, the user position positioning module 1105 is configured to calculate the arrival time of the wake-up voice currently emitted by the user at the multi-microphone device by using the multi-microphone array in a single device, and calculate the user's voice position based on the arrival times of the wake-up voices collected by at least two devices at the multi-microphone device.

[0162] The face orientation recognition module 1106 is configured to calculate the face direction of the user in the position map within the wide-angle camera range of the camera of the device by using the face orientation recognition ability at the chip layer.

[0163] The directional wake-up module 1107 is configured to obtain the target device to be woken up by using the device map calculated and processed by the device map calculation module 1104, the user position located by the user position positioning module 1105, and the face orientation.

[0164] It should be understood that Figure 11 The software structure shown is only an example. The electronic device in the embodiments of the present application may have more or fewer modules than the electronic device shown in the figure, and two or more modules may be combined, etc. Each module shown in the figure may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.

[0165] It should be noted that Figure 11 The audio processing module 1102, the audio recognition module 1103, the device map calculation module 1104, the user position positioning module 1105, the face orientation recognition module 1106, and the directional wake-up module 1107 shown in Figure 2In one or more processing units in the illustrated processor 210, for example, some or all of the audio processing module 1102, the audio recognition module 1103, the device map calculation module 1104, the user location positioning module 1105, the face orientation recognition module 1106, and the directional wake-up module 1107 may be integrated in one or more processors such as an application processor or a dedicated processor. It should be noted that the dedicated processor in the embodiments of the present application may be a DSP, an application specific integrated circuit (ASIC) chip, etc.

[0166] The following embodiments can all be implemented in an electronic device having the above-mentioned hardware structure and / or software structure.

[0167] Based on the same concept, Figure 12 Shown is a device 1200 provided by the present application. The device 1200 includes at least one processor 1210, a memory 1220, and a transceiver 1230. Among them, the processor 1210 is coupled to the memory 1220 and the transceiver 1230. The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms, and is used for information interaction between devices, units, or modules. In the embodiments of the present application, the connection medium between the above-mentioned transceiver 1230, processor 1210, and memory 1220 is not limited. For example, in the embodiments of the present application Figure 12 it is possible that the memory 1220, the processor 1210, and the transceiver 1230 are connected through a bus, and the bus can be divided into an address bus, a data bus, a control bus, etc.

[0168] Specifically, the memory 1220 is used to store program instructions.

[0169] The transceiver 1230 is used to receive or send data.

[0170] The processor 1210 is used to call the program instructions stored in the memory 1220, so that the device 1200 executes the steps performed by the above-mentioned electronic device 30, or executes the steps performed by the electronic device 10 or electronic device 20 in the above-mentioned method.

[0171] In an embodiment of the present application, the processor 1210 may be a general-purpose processor, a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0172] In an embodiment of the present application, the memory 1220 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or may also be a volatile memory, such as a random-access memory (RAM). The memory is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0173] It should be understood that the device 1200 may be used to implement the method shown in the embodiments of the present application. For related features, reference may be made to the above, and details are not described herein again.

[0174] Those skilled in the art can clearly understand that the embodiments of the present application can be implemented by hardware, firmware, or a combination thereof. When implemented in software, the above functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transmission of a computer program from one place to another. The storage media can be any available medium that can be accessed by a computer. By way of example but not limitation: the computer-readable medium can include RAM, ROM, electrically erasable programmable read only memory (EEPROM), compact disc read-Only memory (CD-ROM), or other optical disc storage, magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In addition, any connection can appropriately become a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, wireless, and microwave are included in the definition of the medium. As used in the embodiments of the present application, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where a disk typically replicates data magnetically, while a disc replicates data optically with a laser. The above combinations should also be included within the scope of protection of the computer-readable medium.

[0175] In summary, the above are only the embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made in accordance with the disclosure of the present application shall be included within the protection scope of the present application.

Claims

1. A voice wake-up method, applied to a first electronic device, characterized in that, The method includes: Receiving a voice wake-up instruction from a user; Obtaining information of a first audio of the first electronic device, and obtaining information of a second audio from at least one second electronic device; According to the first audio information and the second audio information, determining the time difference and phase difference of the voice wake-up instruction reaching different microphones of the first electronic device, and the time difference and phase difference of the voice wake-up instruction reaching different microphones of the at least one second electronic device; According to the time difference and phase difference of the voice wake-up instruction reaching different microphones of the first electronic device, determining a first relative distance and a first azimuth angle between the user and the first electronic device; according to the time difference and phase difference of the voice wake-up instruction reaching different microphones of the at least one second electronic device, determining a second relative distance and a second azimuth angle between the user and the at least one second electronic device; Based on a relative position map of the first electronic device and the at least one second electronic device, determining the user position according to the first relative distance, the first azimuth angle, the second relative distance and the second azimuth angle; wherein, the relative position map is pre-constructed based on information of historical audio and dynamically adjusted based on information of real-time audio; Obtaining a user image and detecting the orientation of the user's face; Determining a target device facing the user's face from the first electronic device and the at least one second electronic device according to the relative position map, the user position and the orientation of the user's face; Instructing the target device to respond to the voice wake-up instruction.

2. The method according to claim 1, characterized in that, Determining a target device facing the user's face according to the relative positions of the first electronic device and the at least one second electronic device, the user position and the orientation of the user's face includes: Determining candidate devices facing the user's face from the first electronic device and the at least one second electronic device according to the relative positions of the first electronic device and the at least one second electronic device, the user position and the orientation of the user's face; When the number of candidate devices is greater than or equal to two, determining the relative distances between the user and the at least two candidate devices; Determining the priorities of the candidate devices according to the relative distances, wherein the candidate device with a smaller relative distance has a higher priority; Determining the candidate device corresponding to the highest priority as the target device.

3. The method according to claim 1, wherein The first electronic device includes a first microphone and a second microphone; the information of the first audio includes: a first arrival time of the voice wake-up instruction reaching the first microphone, a second arrival time of the voice wake-up instruction reaching the second microphone, a first phase of the voice wake-up instruction reaching the first microphone, and a second phase of the voice wake-up instruction reaching the second microphone; The at least one second electronic device includes a third microphone and a fourth microphone; the information of the second audio includes: a third arrival time when the voice wake-up instruction arrives at the third microphone, a fourth arrival time when the voice wake-up instruction arrives at the fourth microphone, a third phase when the voice wake-up instruction arrives at the third microphone, and a fourth phase when the voice wake-up instruction arrives at the fourth microphone.

4. The method according to claim 3, wherein Determining the time difference and phase difference of the voice wake-up instruction arriving at different microphones of the first electronic device, and the time difference and phase difference of the voice wake-up instruction arriving at different microphones of the at least one second electronic device according to the first audio information and the second audio information includes: Determining the time difference between the first microphone and the second microphone of the first electronic device at which the voice wake-up instruction arrives according to the first arrival time and the second arrival time; determining the phase difference between the first microphone and the second microphone of the first electronic device at which the voice wake-up instruction arrives according to the first phase and the second phase; Determining the time difference between the third microphone and the fourth microphone of the at least one second electronic device at which the voice wake-up instruction arrives according to the third arrival time and the fourth arrival time; determining the phase difference between the third microphone and the fourth microphone of the at least one second electronic device at which the voice wake-up instruction arrives according to the third phase and the fourth phase.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtaining information of historical audio from the first electronic device and the at least one second electronic device; Obtaining the arrival time and phase at which the voice wake-up instructions issued by the user N times arrive at different electronic devices from the information of the historical audio, where N is a positive integer; Determining the relative azimuth angle and distance difference corresponding to the voice wake-up instructions issued by the user N times according to the arrival time and phase at which the voice wake-up instructions issued by the user N times arrive at different electronic devices; Establishing an objective function with the relative azimuth angle and distance difference corresponding to the voice wake-up instructions issued by the user N times as observation values; Solving the objective function by an exhaustive search method to obtain a relative position map of the first electronic device and the at least one second electronic device.

6. The method according to claim 5, wherein Determining the relative azimuth angle and distance difference corresponding to the voice wake-up instructions issued by the user N times according to the arrival time and phase at which the voice wake-up instructions issued by the user N times arrive at different electronic devices includes: For the Kth voice wake-up instruction among the N voice wake-up instructions issued by the user, where K is any one of the N times, the distance difference corresponding to the Kth voice wake-up instruction issued by the user satisfies the following formula one: Δτ1·v=a1……Formula One Where, Δτ1 is the time difference between the wake-up signal issued by the user and two different electronic devices, v is the sound transmission speed, and a1 is the distance difference between the user's voice generation position and two different electronic devices; The relative azimuth angle corresponding to the Kth voice wake-up instruction issued by the user satisfies the following formula two: 180°-θA1-θB1=b1……Formula Two Among them, θA1 is the azimuth angle of the user's occurrence position relative to the second electronic device, and θB1 is the azimuth angle of the user's voice - sounding position relative to the first electronic device; θB1 is the angle between the line from the sound - generating point P1 to point B and the x - axis of the rectangular coordinate system C, and b1 is the relative azimuth angle of the user's voice - sounding position relative to two different electronic devices.

7. According to the method described in any one of claims 1 to 4, characterized in that, The first electronic device and the at least one second electronic device are connected to the same local area network, or the first electronic device and the at least one second electronic device are pre - bound with the same user account, or the first electronic device and the at least one second electronic device are bound with different user accounts, and different user accounts establish a binding relationship.

8. The method according to any one of claims 1 to 4, characterized in that The first electronic device has an image - acquisition function, or the at least one second electronic device has an image - acquisition function.

9. A voice wake-up system, characterized in that, The system includes a first electronic device and at least one second electronic device; The second electronic device is configured to collect the sound of the surrounding environment through a multi - microphone array and convert it into a second audio. The sound of the surrounding environment includes the voice wake - up instruction issued by the user. The second electronic device is further configured to send the information of the second audio to the first electronic device. The first electronic device is configured to obtain the information of the first audio of the first electronic device and receive the information of the second audio obtained from the at least one second electronic device; according to the information of the first audio and the information of the second audio, determine the time difference and phase difference of the voice wake - up instruction reaching different microphones of the first electronic device, and the time difference and phase difference of the voice wake - up instruction reaching different microphones of the at least one second electronic device. According to the time difference and phase difference of the voice wake - up instruction reaching different microphones of the first electronic device, determine the first relative distance and the first azimuth angle between the user and the first electronic device. According to the time difference and phase difference of the voice wake - up instruction reaching different microphones of the at least one second electronic device, determine the second relative distance and the second azimuth angle between the user and the at least one second electronic device. Based on the relative - position map between the first electronic device and the at least one second electronic device, the first relative distance, the first azimuth angle, the second relative distance, and the second azimuth angle, determine the user position; among them, the relative - position map is pre - constructed based on the information of historical audio and dynamically adjusted based on the information of real - time audio. According to the relative positions of the first electronic device and the at least one second electronic device, the user position, and the user's face orientation, determine the target device towards which the user's face is oriented from the first electronic device and the at least one second electronic device. The first electronic device is further configured to instruct the target device to respond to the voice wake - up instruction.

10. The system according to claim 9, wherein The first electronic device determines the target device towards which the user's face is oriented according to the relative positions of the first electronic device and the at least one second electronic device, the user position, and the user's face orientation. Specifically, it is used for: Determine a candidate device towards which the user's face is oriented from the first electronic device and at least one second electronic device according to the relative positions of the first electronic device and the at least one second electronic device, the user position, and the orientation of the user's face; When the number of candidate devices is greater than or equal to two, determine the relative distances between the user and the at least two candidate devices; Determine the priorities of the candidate devices according to the relative distances, where the candidate device with a smaller relative distance has a higher priority; Determine the candidate device corresponding to the highest priority as the target device.

11. The system according to claim 9, wherein, The first electronic device includes a first microphone and a second microphone; the information of the first audio includes: the first arrival time when the voice wake-up instruction arrives at the first microphone, the second arrival time when the voice wake-up instruction arrives at the second microphone, and the first phase when the voice wake-up instruction arrives at the first microphone, the second phase when the voice wake-up instruction arrives at the second microphone; The at least one second electronic device includes a third microphone and a fourth microphone; the information of the second audio includes: the third arrival time when the voice wake-up instruction arrives at the third microphone, the fourth arrival time when the voice wake-up instruction arrives at the fourth microphone, and the third phase when the voice wake-up instruction arrives at the third microphone, the fourth phase when the voice wake-up instruction arrives at the fourth microphone.

12. The system according to claim 11, wherein The first electronic device determines the time differences and phase differences of the voice wake-up instruction arriving at different microphones of the first electronic device and the time differences and phase differences of the voice wake-up instruction arriving at different microphones of the at least one second electronic device according to the first audio information and the second audio information, specifically for: Determine the time difference between the first microphone and the second microphone of the first electronic device that the voice wake-up instruction arrives at according to the first arrival time and the second arrival time; determine the phase difference between the first microphone and the second microphone of the first electronic device that the voice wake-up instruction arrives at according to the first phase and the second phase; Determine the time difference between the third microphone and the fourth microphone of the at least one second electronic device that the voice wake-up instruction arrives at according to the third arrival time and the fourth arrival time; determine the phase difference between the third microphone and the fourth microphone of the at least one second electronic device that the voice wake-up instruction arrives at according to the third phase and the fourth phase.

13. The system according to any one of claims 9 to 12, characterized in that The first electronic device is further configured to: Obtain the information of historical audio from the first electronic device and the at least one second electronic device; Obtain the arrival times and phases when the user's N voice wake-up instructions arrive at different electronic devices from the information of the historical audio, where N is a positive integer; Determine the relative azimuth angles and distance differences corresponding to the user's N voice wake-up instructions according to the arrival times and phases when the user's N voice wake-up instructions arrive at different electronic devices; Taking the relative azimuth angle and distance difference corresponding to the voice wake-up instruction issued by the user N times as observation values, a target function is established; By solving the target function through an exhaustive search method, a relative position map of the first electronic device and at least one second electronic device is obtained.

14. An electronic device, characterized in that, The electronic device includes a processor and a memory; Program instructions are stored in the memory; When the program instructions are executed, the electronic device is caused to execute the method according to any one of claims 1 to 8.

15. A chip system, characterized in that, The chip system is coupled to the memory in the electronic device, so that the chip calls the program instructions stored in the memory during operation, and implements the method according to any one of claims 1 to 8.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes program instructions, and when the program instructions are run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Electronic equipment control method and device, terminal and storage medium

    CN111176744A

  • Electronic device for estimating position of sound source

    US20190293746A1