Device wake-up methods, apparatus, electronic devices and storage media
By identifying the voice signal energy and acoustic characteristics of the device to be woken up, candidate and target devices are determined, solving the problem of low efficiency in waking up multiple devices and achieving fast and accurate device wake-up.
Patent Information
- Application Number
- CN202111556561.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-12-17
AI Technical Summary
In scenarios involving multiple smart devices, how can we improve the efficiency of device wake-up and solve the problem of multiple devices responding simultaneously?
By acquiring the voice signal energy and acoustic feature information of the device to be woken up, candidate devices are identified and the target device is determined based on the acoustic features. A wake-up command is then sent, reducing the number of devices involved in the decision-making process.
It enables faster device wake-up feedback, improving the efficiency and accuracy of device wake-up.
Smart Images

Figure CN114360550B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech signal processing technology, and in particular to a device wake-up method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of artificial intelligence and the increasing maturity of 5G technology, more and more smart devices are being deployed in home environments, and scenarios with multiple smart devices in the home are now very common. Voice wake-up is the entry point for smart interaction, but multiple smart devices face the problem of multiple wake-up responses. Therefore, how to improve the efficiency of device wake-up is a technical problem that urgently needs to be solved. Summary of the Invention
[0003] This application aims to at least partially address one of the technical problems in the related art.
[0004] Therefore, this application proposes a device wake-up method, apparatus, electronic device, and storage medium to achieve faster response feedback.
[0005] One embodiment of this application proposes a device wake-up method applied to a server, including:
[0006] Acquire wake-up requests sent by multiple devices to be woken up; wherein the wake-up request carries the first speech energy and acoustic feature information of the speech signal corresponding to the wake-up word;
[0007] Based on the first voice energy sent by each of the devices to be woken up, candidate devices to be woken up are determined from a plurality of devices to be woken up;
[0008] Based on the acoustic feature information sent by the candidate devices to be woken up, the target device to be woken up is determined from the candidate devices to be woken up;
[0009] Send response information carrying a wake-up command to the target device to be woken up.
[0010] Another embodiment of this application proposes a device wake-up method, applied to a first device to be woken up, including:
[0011] Acquire a speech signal containing a wake word; filter the speech signal to obtain a target speech signal;
[0012] Based on the target speech signal, determine the first speech energy corresponding to the target speech signal;
[0013] A wake-up request carrying the first speech energy and acoustic feature information is sent to the server, so that the server determines candidate wake-up devices based on the first speech energy, and determines and wakes up the target wake-up device based on the acoustic feature information corresponding to the candidate wake-up devices; wherein, the acoustic feature information is determined based on the speech signal or the target speech signal.
[0014] Another aspect of this application provides a device wake-up device, comprising:
[0015] The acquisition module is used to acquire wake-up requests sent by multiple devices to be woken up; wherein the wake-up request carries the first speech energy and acoustic feature information of the speech signal corresponding to the wake-up word;
[0016] The first determining module is used to determine candidate wake-up devices from a plurality of wake-up devices based on the first voice energy sent by each wake-up device;
[0017] The second determining module is used to determine the target device to be woken up from the candidate devices to be woken up based on the acoustic feature information sent by the candidate devices to be woken up;
[0018] The sending module is used to send response information carrying a wake-up command to the target device to be woken up.
[0019] Another aspect of this application provides a device wake-up device, comprising:
[0020] The acquisition module is used to acquire a voice signal containing a wake-up word; the processing module is used to filter the voice signal to obtain a target voice signal.
[0021] The determining module is used to determine the first speech energy corresponding to the target speech signal based on the target speech signal;
[0022] The sending module is used to send a wake-up request carrying the first voice energy and acoustic feature information to the server, so that the server determines the candidate wake-up device based on the first voice energy, and determines the target wake-up device and wakes it up based on the acoustic feature information corresponding to the candidate wake-up device; wherein, the acoustic feature information is determined based on the voice signal or the target voice signal.
[0023] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method described in one aspect or the method described in another aspect.
[0024] A third aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in one aspect above or the method described in another aspect above.
[0025] The fourth aspect of this application provides a computer program product having a computer program stored thereon, which, when executed by a processor, implements the method described in one aspect above or the method described in another aspect above.
[0026] The device wake-up method, apparatus, electronic device, and storage medium proposed in this application identify candidate wake-up devices by recognizing the magnitude of the voice energy transmitted by the device to be woken up, and then determine the target wake-up device based on acoustic characteristics. This eliminates the need to wait for wake-up requests from all devices to arrive before making a decision, reducing the number of devices involved in the decision-making process and achieving faster response feedback.
[0027] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0028] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0029] Figure 1 A schematic flowchart illustrating a device wake-up method provided in an embodiment of this application;
[0030] Figure 2 A flowchart illustrating another device wake-up method provided in an embodiment of this application;
[0031] Figure 3 A flowchart illustrating another device wake-up method provided in an embodiment of this application;
[0032] Figure 4 A flowchart illustrating another device wake-up method provided in an embodiment of this application;
[0033] Figure 5 This is a schematic diagram of the structure of a device wake-up device provided in an embodiment of this application;
[0034] Figure 6 This is a schematic diagram of another device wake-up device provided in an embodiment of this application;
[0035] Figure 7 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0036] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0037] The following description, with reference to the accompanying drawings, outlines a device wake-up method, apparatus, electronic device, and storage medium according to embodiments of this application.
[0038] Figure 1 This is a schematic flowchart of a device wake-up method provided in an embodiment of this application.
[0039] The device wake-up method in this embodiment is executed by a device wake-up device, which can be located in a server. The server can be a single server or a server cluster, a local server or a cloud server. This embodiment does not impose any limitations.
[0040] like Figure 1 As shown, the method may include the following steps:
[0041] Step 101: Obtain wake-up requests sent by multiple devices to be woken up, wherein the wake-up requests carry the speech energy and acoustic feature information of the speech signal corresponding to the wake-up word.
[0042] The device to be woken up is a smart electronic device. For example, in a home setting, the device to be woken up can be a smart home appliance such as an air conditioner, a color TV, a light, a speaker, or a robot vacuum cleaner, or a smart wearable device.
[0043] The wake-up request is a process whereby each device to be woken up, upon receiving a wake-up word, determines the corresponding speech energy and acoustic feature information based on the speech signal corresponding to the wake-up word, generates a wake-up request based on the corresponding speech energy and acoustic feature information, and sends it to the server. The server then uses the wake-up request to determine whether the device that sent the wake-up request can be woken up.
[0044] In this embodiment, the speech energy of the speech signal corresponding to the wake-up word for each device to be woken up is determined by statistical analysis of the speech signal corresponding to the wake-up word. This process removes interference noise from the speech energy, improving its accuracy. The speech energy indicates the distance between the device to be woken up and the target object that emitted the wake-up word's speech signal. In other words, the closer the device to the target object, the higher the energy of the speech signal corresponding to the wake-up word; conversely, the farther the device is from the target object, the lower the energy of the speech signal. The acoustic characteristics of the speech signal are also used to determine the sound energy; this determined energy is the sound energy before removing interference noise.
[0045] Step 102: Based on the first voice energy sent by each device to be woken up, determine the candidate devices to be woken up from the multiple devices to be woken up.
[0046] There can be one or more candidate devices to be woken up.
[0047] In this embodiment, the server determines multiple wake-up devices whose first voice energy is greater than a set threshold based on the first voice energy received from each wake-up device. Since the larger the first voice energy, the closer the wake-up device is to the user, at least one wake-up device with a first voice energy greater than the set threshold is used as a candidate wake-up device, which reduces the number of subsequent screening and determination of target wake-up devices and improves the efficiency of target wake-up device determination.
[0048] It should be noted that the first speech signal is speech energy after removing interference noise, which can be environmental interference noise or reverberation signal.
[0049] Step 103: Determine the target device to be woken up from the candidate devices based on the acoustic feature information sent by the candidate devices.
[0050] In this embodiment of the application, the acoustic feature information includes the frame length of the wake word and the time-domain signal information converted by the candidate device after receiving the wake word.
[0051] In one implementation of this application, acoustic feature information is used to determine the second speech energy. The second speech energy can be unfiltered speech energy or filtered speech energy. The interference noise can be environmental interference noise signal or reverberation signal. In other words, the second speech energy can be the first speech energy.
[0052] In one implementation of this application, the second speech energy is determined based on the unfiltered raw speech signal as an example. Each candidate wake-up device uses a speech activity detection algorithm to obtain a wake-up word of length T frames, and then calculates its own second speech energy, i.e., average energy E, during the wake-up period:
[0053]
[0054] Where x n (t) represents the time-domain signal collected by the nth candidate device to be woken up, where t is the sampling time.
[0055] Furthermore, in a scenario where there is only one candidate device to be woken up, that device will be selected as the target device to be woken up.
[0056] In another scenario, if there are multiple candidate wake-up devices, the second voice energy corresponding to each candidate wake-up device is calculated, and the candidate device with the largest second voice energy is determined. The candidate wake-up device with the largest second voice energy is then taken as the target wake-up device.
[0057] In another implementation of this application, the acoustic feature information can be a correlation function corresponding to the microphone that collects the voice signal. The magnitude of this correlation function value is used to characterize the magnitude of the received signal. A larger received signal indicates that the microphone is closer to the user, and the higher the probability that the device containing the microphone is the target device to be woken up.
[0058] Step 104: Send response information carrying a wake-up command to the target device to be woken up.
[0059] The wake-up command is generated based on the wake-up word and is used to wake up the target device and perform the corresponding operation. Different wake-up commands are used in different scenarios. For example, the wake-up command is to turn on the air conditioner or turn on the stereo.
[0060] In the device wake-up method of this application embodiment, wake-up requests sent by multiple devices to be woken up are obtained. The wake-up requests carry the speech energy and acoustic feature information of the first speech signal corresponding to the wake-up word. Candidate devices to be woken up are determined based on the first speech energy sent by each device to be woken up. The target device to be woken up is determined from the candidate devices based on the acoustic feature information sent by the candidate devices to be woken up. Response information carrying a wake-up command is sent to the target device to be woken up. Candidate devices to be woken up are determined by identifying the magnitude of the speech energy sent by the devices to be woken up. Then, the target device to be woken up is determined based on the acoustic features. This eliminates the need to wait for the wake-up requests of all devices to be woken up before making a decision, reduces the number of devices involved in the decision-making process, and achieves faster response feedback.
[0061] Based on the previous embodiment, this application provides another device wake-up method, which shows that candidate wake-up devices can be determined based on the first voice energy sent by each device to be woken up, thereby reducing the number of devices involved in the decision-making process and improving the efficiency and accuracy of device wake-up.
[0062] Figure 2 A flowchart illustrating another device wake-up method provided in this application embodiment is shown below. Figure 2 As shown, the method includes the following steps:
[0063] Step 201: Obtain wake-up requests sent by multiple devices to be woken up, wherein the wake-up requests carry the first speech energy and acoustic feature information of the speech signal corresponding to the wake-up word.
[0064] Step 201 can be explained in the foregoing embodiments, and the principle is the same. It is not limited in this embodiment.
[0065] Step 202: Monitor the first voice energy corresponding to each device to be woken up.
[0066] Step 203: Determine whether there is a device to be woken up whose first voice energy is greater than the set energy threshold. If it exists, proceed to step 204; otherwise, proceed to step 206.
[0067] In this embodiment of the application, the first voice energy corresponding to each device to be woken up is monitored. If it is determined that the first voice energy corresponding to one device to be woken up is greater than the set energy threshold, step 204 is executed. If it is determined that the first voice energy corresponding to all the devices to be woken up is not greater than the set threshold, step 205 is executed.
[0068] Step 204: Obtain the location information and / or device set information corresponding to a device to be woken up.
[0069] The location information includes location coordinates or location area. Location coordinates indicate the location of the device to be woken up, and can be coordinates determined through positioning, such as latitude and longitude coordinates. The location area indicates the area, or space, where the device is located. For example, in a home setting, the location area could be the living room or bedroom; in an office setting, it could be the office or meeting room. The division of location areas varies depending on the scenario, and this embodiment does not list or limit them.
[0070] The location information and / or device set information sent by each device to be woken up, which the server obtains, are determined by each device to be woken up. This will be explained in detail in the subsequent implementation of the method by the devices to be woken up, and will not be repeated here.
[0071] Step 205: Determine candidate devices to be woken up based on the location information and / or device set information corresponding to a device to be woken up.
[0072] In the first implementation of this application, the location information includes a location region. Based on the location region corresponding to a device to be woken up, candidate devices to be woken up within that location region are determined. Specifically, if a device to be woken up exists in or near the space where the user who issued the wake-up word is located, and a set energy threshold is set, devices with a first voice energy greater than the set energy threshold can be identified. Based on the location region corresponding to the device to be woken up, the correspondence between the location region and the device to be woken up is queried to determine candidate devices to be woken up within that location region. As an example, a device to be woken up with a first voice energy greater than the set energy threshold is identified as device A. The location region of device A is determined to be X. By querying, devices B and C, which are located in the same location region X as device A, can be identified. In other words, device B and device C, which are candidate devices to be woken up within location region X, can be identified, reducing the number of devices to be woken up involved in the decision-making process and achieving faster response feedback.
[0073] In the second implementation of this application, the location information includes location coordinates. Based on the location coordinates of a device to be woken up, a target location coordinate that matches the location coordinates is queried, and the device to be woken up corresponding to the target location coordinate is selected as a candidate device to be woken up. In other words, by determining the target location coordinates through location coordinate matching, candidate devices to be woken up whose distance from the device to be woken up is within a set distance are identified, reducing the number of devices to be woken up that need to be processed subsequently and improving processing efficiency.
[0074] In the third implementation of this application, candidate devices to be woken up are determined based on the identifiers of each device to be woken up carried in the device set information. The identifier of each device to be woken up is used to uniquely identify the corresponding device. In one scenario, each device to be woken up carried in the device set information refers to devices located in the same location area. This same location area can be a definite location area or an indefinite location area. For an indefinite location area, for example, in a home scenario, it can be determined that the television and the robot vacuum cleaner are in the same location area, but it is not certain whether this same location area is the living room, a bedroom, or the balcony; that is, it is uncertain which specific location area it is.
[0075] In another scenario, the devices to be woken up carried in the device set information refer to devices whose corresponding location coordinates are less than or equal to a set threshold. By using the devices to be woken up carried in the device set, the number of devices to be woken up participating in the decision-making process is reduced, resulting in faster response feedback. In the fourth implementation of this application, candidate devices to be woken up are determined based on the location information corresponding to a device to be woken up and the device set information. That is, the location information corresponding to the device to be woken up can be obtained at the same time as the device set information, which can quickly determine the candidate devices to be woken up at the corresponding location information, thus improving efficiency.
[0076] Step 206: Select multiple devices to be woken up as candidate devices to be woken up.
[0077] In another scenario of this application embodiment, if it is determined that the first voice energy corresponding to each device to be woken up is less than or equal to a set energy threshold, it indicates that the location information of the user who issued the wake-up word and the devices to be woken up do not match, that is, the distance between the user who issued the wake-up word and the devices to be woken up is far. In this case, all devices to be woken up are considered as candidate devices to be woken up. For example, if the user who issued the wake-up word is in the bathroom, and the user wants to control a smart device in the living room to turn on, such as a TV, then if the first voice energy corresponding to each device to be woken up is detected to be less than or equal to the set energy threshold, then all devices that have received the wake-up request are considered as candidate devices to be woken up.
[0078] Step 207: Determine the target device to be woken up from the candidate devices based on the acoustic feature information sent by the candidate devices to be woken up.
[0079] Step 208: Send response information carrying a wake-up command to the target device to be woken up.
[0080] Steps 207 and 208 can be explained in the foregoing embodiments, as the principle is the same, and will not be repeated in this embodiment.
[0081] In the device wake-up method of this application embodiment, wake-up requests sent by multiple devices to be woken up are obtained. Each wake-up request carries the speech energy and acoustic feature information of the speech signal corresponding to the wake-up word. By comparing the first speech energy sent by each device to be woken up with a set energy threshold, the location information and / or device set information corresponding to a device whose first speech energy is greater than the set energy threshold is determined. Based on the location information and / or device set information, candidate devices to be woken up are determined. Based on the acoustic feature information of the candidate devices, the target device to be woken up is determined. This eliminates the need to wait for wake-up requests from all devices to arrive before making a decision, reducing the number of devices involved in the decision-making process and achieving faster response feedback. However, if it is determined that there is no device whose first speech energy is greater than the set energy threshold (i.e., the distance between the user and the device to be woken up is far), then a decision is made for all devices to be woken up, meeting the needs of different scenarios.
[0082] Based on the above embodiments, this application provides another device wake-up method. Figure 3 A flowchart illustrating another device wake-up method provided in this application embodiment is shown below. Figure 3 As shown, the method includes the following steps:
[0083] Step 301: Obtain the speech signal containing the wake word.
[0084] The execution subject of this application embodiment is a first device to be woken up, which is a smart electronic device, such as a smart speaker or a smart robot. The first device to be woken up is equipped with a microphone, which can be a single microphone or a microphone array; this embodiment does not impose any limitations.
[0085] The wake word can be set text, such as "Xiao Mi" or "Xiao Bai, sweep the floor".
[0086] In this embodiment of the application, the first device to be woken up collects voice signals through a microphone and obtains a wake-up word by parsing the voice signals.
[0087] It should be understood that the first device to be woken up is a device to be woken up in the scenario. In this embodiment, the first device to be woken up and the second device to be woken up are just identifiers. The first device to be woken up can also be the second device to be woken up, and the second device to be woken up can also be the first device to be woken up.
[0088] Step 302: Filter the speech signal to obtain the target speech signal.
[0089] In this embodiment, the acquired speech signal is first processed by framing to obtain multiple speech frames. For each speech signal frame, windowing is applied, and a short-time Fourier transform is performed on the windowed speech signal of that frame to finally obtain the frequency domain signal of that frame. Thus, the frequency domain signal corresponding to the speech signal can be obtained. The filtering process can be implemented in the following ways:
[0090] In one implementation of this application, a filter is used to remove interference signals in the environment from the frequency domain signal of the speech signal in order to obtain a target speech signal with noise removed.
[0091] In another implementation of this application, a Kalman filter is used to estimate the reverberation spectrum of the obtained frequency domain signal. Based on the estimated reverberation spectrum, the frequency domain signal is reverberated and reverberated to obtain the target speech signal of the dereverberated signal. In other words, the target speech signal is the direct speech signal from the sound source to the device to be woken up.
[0092] It is important to understand that different spaces, due to their different spatial structures, materials, and objects placed there, will produce different levels of reverberation interference when sound signals propagate in different spaces. Therefore, when identifying the target space based on the first speech energy corresponding to each device, removing the interference of the reverberation signal can improve the reliability of the first speech energy corresponding to the target speech signal, thereby improving the accuracy of the target space determination.
[0093] Step 303: Determine the first speech energy corresponding to the target speech signal based on the target speech signal.
[0094] In this embodiment, for a target speech signal with reverberation information removed, the corresponding first speech energy is determined. The method for determining the speech energy can be referred to the description in the foregoing embodiments, and will not be repeated in this embodiment.
[0095] Step 304: A wake-up request carrying the first voice energy and acoustic feature information is sent to the server, so that the server determines the candidate wake-up device based on the first voice energy, and determines the target wake-up device and wakes it up based on the acoustic feature information corresponding to the candidate wake-up device.
[0096] Among them, acoustic feature information is determined based on speech signals or target speech signals.
[0097] In this embodiment, the acoustic feature information can be the energy of the speech signal collected by the microphone, or the correlation function corresponding to the microphone collecting the speech signal. The magnitude of the correlation function value is used to characterize the magnitude of the received signal. A larger received signal indicates that the microphone is closer to the user, and the higher the probability that the device containing the microphone is the target device to be woken up.
[0098] In this embodiment, the first device to be woken up sends a wake-up request carrying the first voice energy and acoustic feature information to the server via a wireless network, such as a WIFI network. The method executed by the server can be referred to the explanation in the foregoing embodiment, and the principle is the same. It is not limited in this embodiment.
[0099] In the device wake-up method of this application embodiment, a voice signal containing a wake-up word is acquired, the voice signal is de-reverberated to obtain a target voice signal, a first voice energy corresponding to the target voice signal is determined based on the target voice signal, and a wake-up request carrying the first voice energy and acoustic feature information is sent to a server, so that the server determines candidate devices to be woken up based on the first voice energy, and determines and wakes up the target device to be woken up based on the acoustic feature information corresponding to the candidate devices to be woken up. The acoustic feature information is determined based on the voice signal or the target voice signal. By filtering the acquired voice signal, the accuracy of the voice signal is improved, thereby improving the accuracy of the calculated first voice energy.
[0100] Based on the above embodiments, this application provides another device wake-up method. Figure 4 A flowchart illustrating another device wake-up method provided in this application embodiment is shown below. Figure 4 As shown, the method includes the following steps:
[0101] Step 401: Obtain the location information of the first device to be woken up.
[0102] The location information includes location coordinates and location area.
[0103] When the location information is a location area, in one implementation of this application embodiment, when the first device to be woken up connects to the wireless network, it responds to the user's configuration operation and selects the location area corresponding to the first device to be woken up in the network configuration application interface to determine the location area of the first device to be woken up, thus realizing the server's determination of the location area of each device to be woken up. For example, the location area configured for the TV and Xiao Ai speaker is the living room, and the location area configured for the air conditioner and smart alarm clock is the master bedroom.
[0104] When the location information is location coordinates, in one implementation of this application embodiment, a positioning module can be set in the first device to be woken up, and the location coordinates of the first device to be woken up can be determined according to the positioning module. The positioning module can be a positioning module based on Global Positioning System (GPS) technology or a positioning module based on a mobile network. In this embodiment, the implementation method of positioning is not limited.
[0105] Step 402: Broadcast a location determination request.
[0106] In this embodiment of the application, the first device to be woken up can broadcast a location determination request. As one implementation, this can be done by broadcasting based on a beacon frame. Other devices to be woken up, referred to as second devices to be woken up for easy distinction, listen to the beacon frames of the surrounding first devices to be woken up or other second devices to be woken up in order to obtain the location determination request.
[0107] Step 403: Receive a response from at least one second device to be woken up.
[0108] The application response is the reply sent after listening to the location determination request broadcast by the first device to be woken up. The location determination request may carry the identification information of the first device to be woken up. At least one device to be woken up that receives the location determination request will, based on the identification information of the first device to be woken up, either forward the reply to the server or send the reply directly to the corresponding first device to be woken up.
[0109] The response carries the identification information of the second device to be woken up that sent the response. Therefore, the first device to be woken up can determine the identification of the second device to be woken up that is in the same location area as itself, or the identification of any second device to be woken up whose location coordinates are less than or equal to a set threshold, based on the received response.
[0110] It should be noted that the "same location area" referred to in this step can be an uncertain location area. For example, in a home setting, it can be determined that the TV and the robot vacuum cleaner are in the same location area, but it is not certain whether the same location area is the living room, a bedroom, or the balcony. In other words, it is uncertain which specific location area the first device to be woken up and at least one second device to be woken up are in, but it can be determined that the first device to be woken up and at least one second device to be woken up are in the same location area.
[0111] Step 404: Generate device set information corresponding to the first device to be woken up based on the identifier of the first device to be woken up and the identifier of at least one second device to be woken up.
[0112] The device set information includes multiple devices to be woken up whose location information matches. Location information matching means that multiple devices to be woken up are in the same location area, or the distance between their location coordinates is less than a set distance.
[0113] Step 405: Send the location information and / or device set information corresponding to the first device to be woken up to the server.
[0114] The first device to be woken up sends its corresponding location information and / or device set information to the server via a wireless network. The method by which the server determines candidate devices to be woken up based on the location information and / or device set information is the same as explained in the signed embodiment, and is not limited thereto in this embodiment.
[0115] It should be noted that the steps in this embodiment can be executed after step 104 or before any step in the embodiment, and this embodiment does not impose any limitations.
[0116] It should be noted that the second device to be woken up can also be woken up in the same way. By broadcasting, other devices in the same location area can be identified. Through broadcasting, they can sense each other and determine which devices are in the same location area. This automatic sensing reduces the time required for manual operation and lowers the configuration cost.
[0117] In the device wake-up method of this application embodiment, a voice signal containing a wake-up word is acquired, the voice signal is de-reverberated to obtain a target voice signal, a first voice energy corresponding to the target voice signal is determined based on the target voice signal, and a wake-up request carrying the first voice energy and acoustic feature information is sent to a server, so that the server determines candidate devices to be woken up based on the first voice energy, and determines and wakes up the target device to be woken up based on the acoustic feature information corresponding to the candidate devices to be woken up. The acoustic feature information is determined based on the voice signal or the target voice signal. By filtering the acquired voice signal, the accuracy of the voice signal is improved, thereby improving the accuracy of the calculated first voice energy.
[0118] To implement the above embodiments, this application also proposes a device wake-up device.
[0119] Figure 5 This is a schematic diagram of a device wake-up device provided in an embodiment of this application.
[0120] like Figure 5 As shown, the device may include: an acquisition module 51, a first determination module 52, a second determination module 53, and a transmission module 54.
[0121] The acquisition module 51 is used to acquire wake-up requests sent by multiple devices to be woken up; wherein the wake-up request carries the first speech energy and acoustic feature information of the speech signal corresponding to the wake-up word.
[0122] The first determining module 52 is used to determine candidate wake-up devices from multiple wake-up devices based on the first voice energy sent by each wake-up device.
[0123] The second determining module 53 is used to determine the target device to be woken up from the candidate devices to be woken up based on the acoustic feature information sent by the candidate devices to be woken up.
[0124] The sending module 54 is used to send response information carrying a wake-up command to the target device to be woken up.
[0125] Furthermore, in one implementation of this application embodiment, the first determining module 52 is specifically used for:
[0126] Monitor the first voice energy received from each of the devices to be woken up;
[0127] If it is determined that the first voice energy corresponding to a device to be woken up is greater than a set energy threshold, the location information and / or device set information corresponding to the device to be woken up are obtained.
[0128] Based on the location information and / or device set information corresponding to the device to be woken up, candidate devices to be woken up are determined.
[0129] In one implementation of this application embodiment, the location information includes a location region, and the first determining module 52 is specifically used for:
[0130] Based on the location region corresponding to the device to be woken up, the candidate device to be woken up in the location region is determined.
[0131] In one implementation of this application, the location information includes location coordinates, and the first determining module 52 is specifically used for:
[0132] Based on the location coordinates corresponding to the device to be woken up, query the target location coordinates that match the location coordinates;
[0133] The device to be woken up corresponding to the target location coordinates is selected as the candidate device to be woken up.
[0134] In one implementation of this application, the first determining module 52 is specifically used to: determine the candidate devices to be woken up based on the identifiers of each device to be woken up carried in the device set information.
[0135] In one implementation of this application, the first determining module 52 is specifically used to: when it is determined that there is no voice energy corresponding to a device to be woken up that is greater than a set energy threshold, select the plurality of devices to be woken up as the candidate devices to be woken up.
[0136] In one implementation of this application embodiment, the second determining module 53 is specifically used for:
[0137] Based on the acoustic feature information sent by the candidate device to be woken up, the second voice energy corresponding to the candidate device to be woken up is determined;
[0138] The candidate device to be woken up with the highest second voice energy is selected as the target device to be woken up.
[0139] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and will not be repeated here.
[0140] In the device wake-up apparatus of this application embodiment, wake-up requests sent by multiple devices to be woken up are acquired. The wake-up requests carry the speech energy and acoustic feature information of the first speech signal corresponding to the wake-up word. Based on the first speech energy sent by each device to be woken up, candidate devices to be woken up are determined. Based on the acoustic feature information sent by the candidate devices to be woken up, a target device to be woken up is determined from the candidate devices. Response information carrying a wake-up command is sent to the target device to be woken up. By identifying the magnitude of the speech energy sent by the devices to be woken up, candidate devices to be woken up are determined, and then the target device to be woken up is determined based on the acoustic features. This eliminates the need to wait for the wake-up requests of all devices to be woken up before making a decision, reduces the number of devices involved in the decision-making process, and achieves faster response feedback.
[0141] To implement the above embodiments, this application also proposes a device wake-up device.
[0142] Figure 6 This is a schematic diagram of another device wake-up device provided in an embodiment of this application.
[0143] like Figure 6 As shown, the device may include: an acquisition module 61, a processing module 62, a determination module 63, and a sending module 64.
[0144] Acquisition module 61 is used to acquire voice signals containing wake words.
[0145] The processing module 62 is used to filter the speech signal to obtain the target speech signal.
[0146] The determining module 63 is used to determine the first speech energy corresponding to the target speech signal based on the target speech signal.
[0147] The sending module 64 is used to send a wake-up request carrying the first voice energy and acoustic feature information to the server, so that the server determines the candidate wake-up device based on the first voice energy, and determines the target wake-up device and wakes it up based on the acoustic feature information corresponding to the candidate wake-up device; wherein, the acoustic feature information is determined based on the voice signal or the target voice signal.
[0148] In one implementation of this application embodiment, the processing module 62 is specifically used for:
[0149] The speech signal is de-reverberated to obtain the target speech signal.
[0150] In one implementation of this application, the apparatus further includes:
[0151] The second sending module is used to send the location information and / or device set information of the first device to be woken up to the server.
[0152] In one implementation of this application, the apparatus further includes:
[0153] The broadcast module is used to broadcast location determination requests;
[0154] A receiving module is used to receive a response sent by at least one second device to be woken up;
[0155] The generation module is used to generate device set information corresponding to the first device to be woken up based on the identifier of the first device to be woken up and the identifier of the at least one second device to be woken up.
[0156] In one implementation of this application embodiment, the acquisition module 61 is further configured to:
[0157] Obtain the location information of the first device to be woken up.
[0158] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of this embodiment, and will not be repeated here.
[0159] In the device wake-up device of this application embodiment, a voice signal containing a wake-up word is acquired, the voice signal is de-reverberated to obtain a target voice signal, a first voice energy corresponding to the target voice signal is determined based on the target voice signal, and a wake-up request carrying the first voice energy and acoustic feature information is sent to a server, so that the server determines candidate devices to be woken up based on the first voice energy, and determines and wakes up the target device to be woken up based on the acoustic feature information corresponding to the candidate devices to be woken up. The acoustic feature information is determined based on the voice signal or the target voice signal. By filtering the acquired voice signal, the accuracy of the voice signal is improved, thereby improving the accuracy of the calculated first voice energy.
[0160] To implement the above embodiments, this application also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method described in the foregoing method embodiments.
[0161] To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing method embodiments.
[0162] To implement the above embodiments, this application also proposes a computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in the foregoing method embodiments.
[0163] Figure 7 This is a block diagram of an electronic device provided in an embodiment of this application. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0164] Reference Figure 7 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.
[0165] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 208 and processing component 802.
[0166] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of this data include instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0167] Power component 806 provides power to various components of electronic device 800. Power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0168] Multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0169] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0170] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0171] Sensor assembly 814 includes one or more sensors for providing state assessments of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or a component of electronic device 800, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0172] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0173] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0174] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0175] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0176] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0177] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0178] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0179] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0180] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0181] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0182] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A device wake-up method, characterized in that, Applied to servers, including: Receive location information and / or device set information from multiple devices to be woken up; Acquire wake-up requests sent by multiple devices to be woken up; wherein the wake-up request carries the first speech energy and acoustic feature information of the speech signal corresponding to the wake-up word, the first speech energy is the speech energy after removing interference noise, and the acoustic feature information includes the frame length of the wake-up word and the time domain signal information converted by the device to be woken up after receiving the wake-up word; The first voice energy corresponding to each of the devices to be woken up is monitored and received. The magnitude of the first voice energy indicates the distance between the device to be woken up and the target object that issued the voice signal corresponding to the wake-up word. If it is determined that the first voice energy corresponding to a device to be woken up is greater than a set energy threshold, the location information and / or device set information corresponding to the device to be woken up are obtained. Based on the location information and / or device set information corresponding to the device to be woken up, candidate devices to be woken up are determined, wherein the candidate devices to be woken up are in the same location area as the device to be woken up. Based on the acoustic feature information sent by the candidate devices to be woken up, the target device to be woken up is determined from the candidate devices to be woken up; Send response information carrying a wake-up command to the target device to be woken up; The location information includes a location region. The step of determining candidate devices to be woken up based on the location information corresponding to the device to be woken up includes: Based on the location region corresponding to the device to be woken up, the device to be woken up located in the location region is selected as the candidate device to be woken up.
2. The method as described in claim 1, characterized in that, The location information includes location coordinates. The step of determining candidate devices to be woken up based on the location information corresponding to the device to be woken up includes: Based on the location coordinates corresponding to the device to be woken up, query the target location coordinates that match the location coordinates; The device to be woken up corresponding to the target location coordinates is selected as the candidate device to be woken up.
3. The method as described in claim 1, characterized in that, The step of determining candidate wake-up devices based on the device set information corresponding to the wake-up device includes: The candidate devices to be woken up are determined based on the identifiers of each device to be woken up carried in the device set information.
4. The method as described in claim 1, characterized in that, After monitoring and receiving the voice energy corresponding to each of the devices to be woken up, the method further includes: If it is determined that there is no device to be woken up whose voice energy is greater than the set energy threshold, the plurality of devices to be woken up will be selected as candidate devices to be woken up.
5. The method as described in claim 1, characterized in that, The step of determining the target device to be woken up from the candidate devices to be woken up based on the acoustic feature information sent by the candidate devices to be woken up includes: Based on the acoustic feature information sent by the candidate device to be woken up, the second voice energy corresponding to the candidate device to be woken up is determined; The candidate device to be woken up with the highest second voice energy is selected as the target device to be woken up.
6. A device wake-up method, characterized in that, Applied to a first device to be woken up, the method includes: Acquire the voice signal containing the wake word; The speech signal is filtered to obtain the target speech signal; Based on the target speech signal, determine the first speech energy corresponding to the target speech signal; A wake-up request carrying the first voice energy and acoustic feature information is sent to the server, so that the server determines candidate wake-up devices based on the first voice energy, and determines and wakes up the target wake-up device based on the acoustic feature information corresponding to the candidate wake-up devices; wherein, the acoustic feature information is determined based on the voice signal or the target voice signal; The server uses the method described in any one of claims 1-5 to determine the target device to be woken up.
7. The method as described in claim 6, characterized in that, The step of filtering the speech signal to obtain the target speech signal includes: The speech signal is de-reverberated to obtain the target speech signal.
8. The method as described in claim 6, characterized in that, The method further includes: Send the location information and / or device set information of the first device to be woken up to the server.
9. The method as described in claim 8, characterized in that, Before sending the device set information of the first device to be woken up to the server, the following steps are included: Broadcast location determination request; Receive a response from at least one second device to be woken up; Based on the identifier of the first device to be woken up and the identifier of the at least one second device to be woken up, device set information corresponding to the first device to be woken up is generated.
10. The method as described in claim 8, characterized in that, Before sending the location information of the first device to be woken up to the server, the following steps are included: Obtain the location information of the first device to be woken up.
11. A device wake-up device, characterized in that, Applied to servers, including: The acquisition module is used to acquire wake-up requests sent by multiple devices to be woken up; wherein the wake-up request carries the first speech energy and acoustic feature information of the speech signal corresponding to the wake-up word, the first speech energy is the speech energy after removing interference noise, and the acoustic feature information includes the frame length of the wake-up word and the time domain signal information converted by the device to be woken up after receiving the wake-up word; A first determining module is configured to monitor the first voice energy corresponding to each of the devices to be woken up, wherein the magnitude of the first voice energy indicates the distance between the device to be woken up and the target object that emits the voice signal corresponding to the wake-up word; if it is determined that the first voice energy corresponding to a device to be woken up is greater than a set energy threshold, the module acquires the location information and / or device set information corresponding to the device to be woken up; and determines candidate devices to be woken up based on the location information and / or device set information corresponding to the device to be woken up, wherein the candidate devices to be woken up are located in the same location area as the device to be woken up. The second determining module is used to determine the target device to be woken up from the candidate devices to be woken up based on the acoustic feature information sent by the candidate devices to be woken up; The sending module is used to send response information carrying a wake-up command to the target device to be woken up; The first determining module is specifically used for: Based on the location area corresponding to the device to be woken up, the device to be woken up in the location area is selected as the candidate device to be woken up; The device also performs: Receive location information and / or device set information from multiple devices to be woken up.
12. A device wake-up device, characterized in that, Applied to a first device to be woken up, the method includes: The acquisition module is used to acquire a voice signal containing a wake-up word; the processing module is used to filter the voice signal to obtain a target voice signal. The determining module is used to determine the first speech energy corresponding to the target speech signal based on the target speech signal; The sending module is used to send a wake-up request carrying the first voice energy and acoustic feature information to the server, so that the server determines candidate wake-up devices based on the first voice energy, and determines and wakes up the target wake-up device based on the acoustic feature information corresponding to the candidate wake-up devices; wherein, the acoustic feature information is determined based on the voice signal or the target voice signal; The server uses the method described in any one of claims 1-5 to determine the target device to be woken up.
13. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method as described in any one of claims 1-5, or implements the method as described in any one of claims 6-10.
14. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5, or the method as described in any one of claims 6-10.
Citation Information
Patent Citations
Intelligent equipment awakening method and device, intelligent loudspeaker box and storage medium
CN111192591A
Voice equipment awakening method and device
CN112634872A
Near-awakening method for voice equipment in whole-house intelligent system
CN113506570A