Multi-device based voice processing method, medium, electronic device and system

By selecting suitable pickup devices in multi-device scenarios and utilizing noise reduction processing, the problem of low voice recognition accuracy caused by noise interference near the device is solved, the voice assistant's pickup effect and recognition accuracy are improved, and the user experience is improved.

CN114255763BActive Publication Date: 2025-09-16HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010955837.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-11
Publication Date
2025-09-16
Estimated Expiration
2040-09-11

AI Technical Summary

Technical Problem

In a multi-device scenario, if there is strong external noise near the selected device or the device has poor sound pickup capability, the voice recognition accuracy will be low, affecting the voice assistant's sound pickup effect and recognition accuracy.

Method used

By selecting the electronic device that is closest to the user, farthest from external noise sources, and has internal noise reduction capabilities as the sound pickup device among multiple devices, and using the audio information of other devices for noise reduction processing, the sound pickup effect and voice recognition accuracy are improved.

Benefits of technology

It alleviates the impact of electronic device deployment location, internal noise interference or external noise interference on the voice assistant's pickup effect in multi-device scenarios, and improves the environmental robustness of voice recognition and user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114255763B_ABST
    Figure CN114255763B_ABST
Patent Text Reader

Abstract

The present application relates to speech processing technology in the field of artificial intelligence, and in particular to a speech processing method, medium, electronic device and system based on multiple devices, which can alleviate the impact of the internal noise of electronic devices that are playing audio in multi-device scenarios on the sound pickup effect of voice assistants, ensure the sound pickup effect of voice assistants based on multiple devices, and thus help to ensure the accuracy of voice recognition of voice assistants, and improve the environmental robustness of speech recognition in multi-device scenarios. The scheme includes: a first electronic device among multiple electronic devices picks up sound to obtain a speech to be recognized; the first electronic device receives audio information related to the audio played by the second electronic device among the multiple electronic devices; the first electronic device performs noise reduction processing on the picked-up speech to be recognized based on the received audio information. The scheme is specifically applied to the process of voice assistants picking up sound based on multiple devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to speech processing technology in the field of artificial intelligence, and in particular to a multi-device based speech processing method, medium, electronic device and system. Background Art

[0002] A voice assistant is an application (APP) built on artificial intelligence (AI). Smart devices such as mobile phones receive and recognize user voice commands through voice assistants, providing users with voice control functions such as interactive dialogue, information query, and device control. With the widespread popularity of smart devices with voice assistants, there are usually multiple devices with voice assistants installed in the user's environment (such as the user's home). In this multi-device scenario, if there is a device with the same wake-up word among multiple devices, then after the user says the wake-up word, the voice assistants of the devices with the same wake-up word will be awakened, and will recognize and respond to the user's subsequent voice commands.

[0003] In existing technologies, in multi-device scenarios, multiple devices can collaborate to select the device closest to the user from among multiple devices with the same wake-up word to wake up their voice assistant, so that the device can pick up, recognize, and respond to the user's voice command. However, if there is strong external noise near the selected device, or the device's sound pickup capability is poor, the selected device's recognition accuracy for the voice command in the above-mentioned automatic voice recognition process will be low, and thus the selected device will not be able to accurately execute the operation indicated by the voice command. Summary of the Invention

[0004] The embodiments of the present application provide a multi-device-based voice processing method, medium, electronic device, and system. The sound pickup device selected from the multiple devices may have one or more favorable factors such as the closest distance to the user, the farthest distance from the external noise source, and the ability to reduce internal noise. This can alleviate the impact of the electronic device deployment location, internal noise interference, or external noise interference on the voice assistant's sound pickup effect and voice recognition accuracy in multi-device scenarios, thereby improving the user interaction experience and the environmental robustness of voice recognition in multi-device scenarios.

[0005] In the first aspect, an embodiment of the present application provides a multi-device based voice processing method, the method comprising: a first electronic device among multiple electronic devices picks up sound to obtain a first voice to be recognized; the first electronic device receives audio information related to the audio played externally by a second electronic device among multiple electronic devices; the first electronic device performs noise reduction processing on the first voice to be recognized obtained by picking up sound according to the received audio information to obtain a second voice to be recognized. It can be understood that the electronic device for picking up sound (i.e., the first electronic device) is the sound pickup device hereinafter, such as an electronic device with a better sound pickup effect selected from the multiple devices. The above-mentioned electronic device that plays audio externally (i.e., the second electronic device) is the internal noise device among the multiple devices, and the audio information of the audio played externally by the second electronic device is the noise reduction information of the internal noise device described below. Specifically, the first electronic device uses the audio information of the audio played by the second electronic device to perform noise reduction processing on the first voice to be recognized picked up to obtain the second voice to be recognized. This can alleviate the impact of the internal noise of the electronic device that is playing audio externally in a multi-device scenario on the voice pickup effect of the voice assistant, ensure the voice assistant's voice pickup effect based on multiple devices, and thus help to ensure the voice recognition accuracy of the voice assistant, and improve the environmental robustness of voice recognition in multi-device scenarios.

[0006] In one possible implementation of the first aspect, the audio information includes at least one of the following: audio data of the externally played audio, and voice activation detection (VAD) information corresponding to the audio. It will be appreciated that the audio information of the audio can reflect the audio itself, and that by using this audio information to perform noise reduction processing on internal noise generated by the externally played audio, the impact of this internal noise on other voice data (e.g., voice data picked up by the user, such as voice data corresponding to the second voice to be recognized) can be eliminated, thereby improving the quality of the picked-up voice data.

[0007] In a possible implementation of the first aspect above, the method further includes: the first electronic device sends a second voice to be recognized to a third electronic device for recognizing voice among multiple electronic devices; or, the first electronic device recognizes the second voice to be recognized. The electronic device for recognizing voice (i.e., the third electronic device) may be the answering device mentioned below. It can be understood that in the multi-device scenario of an embodiment of the present application, the electronic device for recognizing voice and the electronic device for picking up sound may be the same or different, that is, the third electronic device may use the first electronic device (or the microphone module of the first electronic device) as a peripheral to pick up the user's voice commands, so that the peripheral resources of multiple electronic devices equipped with microphone modules and voice assistants can be effectively aggregated.

[0008] In a possible implementation of the first aspect above, before the first electronic device among multiple electronic devices picks up sound to obtain the first voice to be recognized, the method further includes: the first electronic device sends the sound pickup selection information of the first electronic device to the third electronic device, wherein the sound pickup selection information of the first electronic device is used to indicate the sound pickup situation of the first electronic device; the first electronic device is the electronic device for picking up sound selected by the third electronic device from multiple electronic devices based on the sound pickup selection information of the multiple electronic devices obtained. For example, in the multi-device scenario of an embodiment of the present application, after the user speaks a voice command, the user does not need to specifically operate a certain electronic device to pick up the voice command to be recognized (such as the voice command corresponding to the second voice data below), but the answering device (i.e., the third electronic device) automatically uses the sound pickup device (i.e., the second electronic device) as a peripheral to pick up the user's voice command, and then realizes the voice control function through the response of the answering device to the user's voice command.

[0009] In one possible implementation of the first aspect, the method further includes: the first electronic device receiving a sound pickup instruction (hereinafter referred to as a sound pickup instruction) sent by a third electronic device, wherein the sound pickup instruction is used to instruct the first electronic device to pick up sound and send the noise-reduced speech to be recognized to the third electronic device. In this way, under the instruction of the sound pickup instruction, the first electronic device can be informed that it needs to send the picked-up speech to be recognized (such as the second speech to be recognized) to the third electronic device without performing subsequent processing such as recognition of the speech to be recognized.

[0010] In a possible implementation of the first aspect, the sound pickup election information includes at least one of the following: acoustic echo cancellation AEC capability information, microphone module information, device status information, voice information corresponding to the wake-up word obtained by sound pickup, and voice information corresponding to the voice instruction obtained by sound pickup; wherein, the voice instruction is obtained after the sound pickup obtains the wake-up word; the device status information includes at least one of the following: network connection status information, headphone connection status information, microphone occupancy status information, and scenario mode information. It can be understood that different information in the sound pickup election information represents different factors that affect the sound pickup effect of the electronic device. In this way, the embodiment of the present application can comprehensively consider different factors affecting the sound pickup effect of the electronic device to select a sound pickup device, such as selecting the electronic device with the best sound pickup effect for sound pickup, that is, as the sound pickup device among multiple devices.

[0011] In the second aspect, an embodiment of the present application provides a multi-device voice processing method, the method comprising: a second electronic device among multiple electronic devices plays audio outward; the second electronic device sends audio information related to the audio to a first electronic device used for picking up sound among multiple electronic devices, wherein the audio information can be used by the first electronic device to perform noise reduction processing on the audio to be recognized obtained by the first electronic device picking up sound. Specifically, since the electronic device that is playing audio outward can provide the audio information of the audio, the first electronic device used for picking up sound performs noise reduction processing on the first voice to be recognized obtained by picking up sound according to the audio information, thereby eliminating the influence of the internal noise generated by the audio on the sound pickup, so as to improve the sound pickup effect of the first electronic device, that is, to improve the quality of the voice data obtained by picking up sound (that is, the voice data of the second voice to be recognized). Thus, the influence of the internal noise of the electronic device that is playing audio outward in the multi-device scenario on the sound pickup effect of the voice assistant can be alleviated, and the sound pickup effect of the voice assistant based on multiple devices can be ensured, which is conducive to ensuring the voice recognition accuracy of the voice assistant and improving the environmental robustness of voice recognition in the multi-device scenario.

[0012] In a possible implementation of the second aspect, the audio information includes at least one of the following: audio data of the audio, and voice activity detection (VAD) information corresponding to the audio.

[0013] In a possible implementation of the second aspect, the method further includes: the second electronic device receiving a sharing instruction (hereinafter referred to as a noise reduction instruction) from a third electronic device used for voice recognition among the plurality of electronic devices; or the second electronic device receiving a sharing instruction from the first electronic device; wherein the sharing instruction is used to instruct the second electronic device to send the audio information to the first electronic device. It is understood that the electronic device sending the sharing instruction (such as the first electronic device or the third electronic device) can monitor whether the second electronic device is playing audio outward, and only sends the sharing instruction to the second electronic device when the second electronic device is playing audio outward.

[0014] In a possible implementation of the second aspect above, before the second electronic device sends audio information related to the external audio to the first electronic device for picking up sound among multiple electronic devices, the method further includes: the second electronic device sends the second electronic device's sound pickup election information to the third electronic device, wherein the sound pickup election information of the second electronic device is used to indicate the sound pickup situation of the second electronic device; the first electronic device is the electronic device for picking up sound selected by the third electronic device from multiple electronic devices based on the sound pickup election information of the multiple electronic devices obtained. For example, the third electronic device, as the answering device below, can select the electronic device with the best audio quality (i.e., the electronic device with the best sound pickup) for picking up voice commands as the sound pickup device (such as the first electronic device) to support the answering device to complete the voice interaction process with the user through the voice assistant, for example, the sound pickup device can be the electronic device closest to the user and with better SE processing capability. In this way, the peripheral resources of multiple electronic devices equipped with microphone modules and voice assistants can be effectively aggregated, alleviating the impact of the electronic device deployment location on the voice assistant recognition accuracy in multi-device scenarios, and improving the user interaction experience and the environmental robustness of voice recognition in multi-device scenarios.

[0015] In a third aspect, an embodiment of the present application provides a multi-device voice processing method, the method comprising: a third electronic device among a plurality of electronic devices detects that there is a second electronic device among a plurality of electronic devices that is playing audio outward; when the second electronic device is different from the third electronic device, the third electronic device sends a sharing instruction to the second electronic device, wherein the sharing instruction is used to instruct the second electronic device to send audio information related to the audio played outward by the second electronic device to the first electronic device used for picking up sound among the plurality of devices; when the second electronic device is the same as the third electronic device, the third electronic device sends the audio information to the first electronic device; wherein the audio information can be used by the first electronic device to perform noise reduction processing on the first to-be-recognized voice picked up by the first electronic device to obtain a second to-be-recognized voice. Specifically, since the second electronic device that is playing audio outward under the instruction of the third electronic device can provide the audio information of the audio, the first electronic device used for picking up sound performs noise reduction processing on the first to-be-recognized voice picked up according to the audio information, thereby eliminating the influence of the internal noise generated by the audio on the sound pickup, so as to improve the sound pickup effect of the first electronic device, that is, to improve the quality of the voice data picked up (that is, the voice data of the second to-be-recognized voice). In this way, the impact of the internal noise of electronic devices that are playing audio in multi-device scenarios on the voice assistant's sound pickup effect can be alleviated, ensuring the voice assistant's sound pickup effect based on multiple devices, which is conducive to ensuring the voice assistant's voice recognition accuracy and improving the environmental robustness of voice recognition in multi-device scenarios.

[0016] In a possible implementation of the third aspect, the audio information includes at least one of the following: audio data of the audio, and voice activity detection (VAD) information corresponding to the audio.

[0017] In a possible implementation of the third aspect above, the first electronic device is different from the third electronic device, and the method further includes: the third electronic device obtains the second voice to be recognized picked up by the first electronic device from the first electronic device; the first electronic device recognizes the second voice to be recognized. Furthermore, it is beneficial to improve the accuracy of voice recognition during voice control and improve user experience. In this way, even if the selected answering device in the multi-device scenario (such as the third electronic device closest to the user) has a poor sound pickup effect, or there is noise generated by an electronic device that is playing audio, multiple devices can collaboratively pick up and recognize voice data with better audio quality without the user having to move or manually control a specific electronic device to pick up sound.

[0018] In a possible implementation of the third aspect above, before the third electronic device sends a sharing instruction to the second electronic device, the method further includes: the third electronic device obtains sound pickup election information of multiple electronic devices, wherein the sound pickup election information of the multiple electronic devices is used to represent the sound pickup conditions of the multiple electronic devices; the third electronic device selects at least one electronic device from the multiple electronic devices as the first electronic device based on the sound pickup election information of the multiple devices. In this way, the peripheral resources of multiple electronic devices equipped with microphone modules and voice assistants can be effectively aggregated, and the impact of various factors such as the electronic device deployment location, internal noise interference, and external noise interference on the voice assistant recognition accuracy in multi-device scenarios can be alleviated, thereby improving the user interaction experience and the environmental robustness of voice recognition in multi-device scenarios.

[0019] In one possible implementation of the third aspect, the method further includes: the third electronic device sending a sound pickup instruction to the first electronic device, wherein the sound pickup instruction is used to instruct the first electronic device to pick up a sound and transmit the picked-up second speech to be recognized to the third electronic device. It will be understood that under the instruction of the sound pickup instruction, the first electronic device is informed that the picked-up speech to be recognized needs to be transmitted to the third electronic device, without performing subsequent processing such as recognition of the speech to be recognized.

[0020] In a possible implementation of the third aspect above, the above-mentioned sound pickup selection information includes at least one of the following: echo cancellation AEC capability information, microphone module information, device status information, voice information corresponding to the wake-up word obtained by sound pickup, and voice information corresponding to the voice command obtained by sound pickup; wherein, the voice command is obtained after the sound pickup obtains the wake-up word; the device status information includes at least one of the following: network connection status information, headphone connection status information, microphone occupancy status information, and scenario mode information.

[0021] In a possible implementation of the third aspect, the third electronic device selects at least one electronic device from the plurality of electronic devices as the first electronic device based on the sound pickup selection information of the plurality of electronic devices, including at least one of the following: when the third electronic device is in a preset network state, the third electronic device determines the third electronic device as the first electronic device; when the third electronic device is connected to headphones, the third electronic device determines the third electronic device as the first electronic device; the third electronic device determines at least one of the electronic devices in a preset scenario mode among the plurality of electronic devices as the first electronic device. It is understood that if the electronic device is in a device state that is not conducive to sound pickup, such as a poor network connection state of the electronic device, a wired or wireless headset is connected, the microphone is occupied or in airplane mode, it means that the sound pickup effect of the electronic device is difficult to guarantee, or the electronic device cannot normally cooperate with other devices to pick up sound, such as being unable to normally send the picked-up voice data to other electronic devices. In this way, according to the sound pickup device selection step, a sound pickup device with better sound pickup effect (such as the first electronic device) can be selected.

[0022] In a possible implementation of the third aspect, the third electronic device selects at least one electronic device from the plurality of electronic devices as the first electronic device based on the sound pickup selection information of the plurality of electronic devices, including at least one of the following: the third electronic device selects at least one of the plurality of electronic devices with AEC enabled as the first electronic device; the third electronic device selects at least one of the plurality of electronic devices with a noise reduction capability greater than that satisfying a predetermined noise reduction condition as the first electronic device; the third electronic device selects at least one of the plurality of electronic devices with a distance from the user less than a first predetermined distance as the first electronic device; the third electronic device selects at least one of the plurality of electronic devices with a distance from the external noise source greater than a second predetermined distance as the first electronic device. For example, the predetermined noise reduction condition indicates that the electronic device has a good SE processing effect, such as AEC being enabled or having internal noise reduction capability; the first predetermined distance (such as 0.5m) indicates that the electronic device is closer to the user; the second predetermined distance (such as 3m) indicates that the electronic device is farther from the user. It is understood that, generally speaking, the closer the electronic device is to the user, the better the sound pickup effect, and the farther the electronic device is from external noise, the better the sound pickup effect; the better the noise reduction performance of the microphone module or the electronic device with AEC enabled, the better the SE processing effect of the electronic device, that is, the better the sound pickup effect of the electronic device. Therefore, by comprehensively considering these factors, a sound pickup device (ie, the first electronic device mentioned above) with better sound pickup effect can be selected from multiple devices.

[0023] In a possible implementation of the third aspect above, the preset network state includes at least one of the following: a network with a network communication rate less than or equal to a predetermined rate, and a network wire frequency greater than or equal to a predetermined frequency; the preset scenario mode includes at least one of the following: subway mode, flight mode, driving mode, and travel mode. Among them, if the network communication rate is less than or equal to the predetermined rate, and the network wire frequency is greater than or equal to the predetermined frequency, it means that the network communication rate of the electronic device is poor, and the specific values ​​of the predetermined rate and the predetermined frequency can be determined according to actual needs. It can be understood that the electronic device in the preset network state is generally not suitable for participating in the election of the sound pickup device or serving as a sound pickup device (such as the first electronic device for picking up sound).

[0024] In one possible implementation of the third aspect, the third electronic device selects the first electronic device from multiple electronic devices using a neural network algorithm or a decision tree algorithm. It will be appreciated that the sound pickup selection information from the multiple devices can be used as input to the neural network algorithm or the decision tree algorithm, which then outputs a decision indicating that the first electronic device is the sound pickup device.

[0025] In a fourth aspect, the present application provides a multi-device voice processing method, the method comprising: a third electronic device among multiple electronic devices obtains sound pickup election information of multiple electronic devices, wherein the sound pickup election information is used to indicate the sound pickup status of the multiple electronic devices; the third electronic device selects at least one electronic device from the multiple electronic devices as a first electronic device for sound pickup based on the sound pickup election information of the multiple devices, wherein the first electronic device is the same as or different from the third electronic device; the third electronic device obtains the speech to be recognized obtained by the first electronic device from the first electronic device; and the third electronic device recognizes the acquired speech to be recognized. Thus, even if the third electronic device selected in the multi-device scenario (such as the electronic device closest to the user) has a poor sound pickup effect, multiple devices can collaboratively pick up and recognize voice data with better audio quality without the user having to move the position or manually control a specific electronic device to pick up the sound. Furthermore, this is conducive to improving the accuracy of speech recognition during voice control and improving the user experience. In addition, it can alleviate the impact of various factors such as the deployment location of electronic devices and external noise interference on the voice assistant's sound pickup effect and speech recognition accuracy in multi-device scenarios, thereby improving the user interaction experience and the environmental robustness of speech recognition in multi-device scenarios.

[0026] In a possible implementation of the fourth aspect, the pickup selection information includes at least one of the following: acoustic echo cancellation AEC capability information, microphone module information, device status information, voice information corresponding to the wake-up word obtained by pickup, and voice information corresponding to the voice command obtained by pickup; wherein, the voice command is obtained after the wake-up word is picked up; the device status information includes at least one of the following: network connection status information, headphone connection status information, microphone occupancy status information, and scenario mode information.

[0027] In a possible implementation of the fourth aspect above, the third electronic device selects at least one electronic device as the first electronic device from multiple electronic devices based on the sound pickup selection information of multiple electronic devices, including at least one of the following: when the third electronic device is in a preset network state, the third electronic device determines the third electronic device as the first electronic device; when the third electronic device is connected to headphones, the third electronic device determines the third electronic device as the first electronic device; the third electronic device determines at least one of the multiple electronic devices that are in a preset scenario mode as the first electronic device.

[0028] In a possible implementation of the fourth aspect, the third electronic device selects at least one electronic device as the first electronic device from a plurality of electronic devices based on the sound pickup selection information of the plurality of electronic devices, including at least one of the following: the third electronic device selects at least one of the electronic devices in which AEC is effective as the first electronic device; the third electronic device selects at least one of the electronic devices in which noise reduction capability is greater than that of the electronic devices that meet the predetermined noise reduction conditions as the first electronic device; the third electronic device selects at least one of the electronic devices in which the distance from the user is less than the first predetermined distance as the first electronic device; the third electronic device selects at least one of the electronic devices in which the distance from the external noise source is greater than the second predetermined distance as the first electronic device.

[0029] In a possible implementation of the fourth aspect, the preset network status includes at least one of the following: a network communication rate less than or equal to a predetermined rate, and a network line frequency greater than or equal to a predetermined frequency; and the preset scenario mode includes at least one of the following: subway mode, flight mode, driving mode, and travel mode.

[0030] In a possible implementation of the fourth aspect, the third electronic device selects the first electronic device from a plurality of electronic devices using a neural network algorithm or a decision tree algorithm.

[0031] In a possible implementation of the fourth aspect above, the method further includes: a third electronic device detects that there is a second electronic device among multiple electronic devices that is playing audio outwardly; the third electronic device sends a sharing instruction to the second electronic device, wherein the sharing instruction is used to instruct the second electronic device to send audio information related to the audio played outwardly by the second electronic device to the first electronic device, wherein the audio information can be used by the first electronic device to perform noise reduction processing on the audio to be recognized picked up by the first electronic device.

[0032] In a possible implementation of the fourth aspect above, the third electronic device is different from the first electronic device, and the method further includes: the third electronic device plays audio; the third electronic device sends audio information related to the audio played by the third electronic device to the first electronic device, wherein the audio information can be used by the first electronic device to perform noise reduction processing on the audio to be recognized picked up by the first electronic device.

[0033] In a possible implementation of the fourth aspect, the audio information includes at least one of the following: audio data of an external audio player, and voice activation detection (VAD) information corresponding to the audio.

[0034] In a sixth aspect, the present application provides a device, which is included in an electronic device, and the device has the function of implementing the above aspects and the electronic device behavior in the possible implementation methods of the above aspects. The function can be implemented by hardware, or it can be implemented by hardware executing the corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a sound pickup unit or module (such as a microphone or microphone array), a receiving unit or module (such as a transceiver), a noise reduction module or unit (such as a processor with the function of the module or unit), etc. For example, the sound pickup unit or module is used to support a first electronic device among multiple electronic devices to pick up sound to obtain a first voice to be recognized; the receiving unit or module (such as a transceiver) is used to support the first electronic device to receive audio information related to the audio played by the second electronic device from the second electronic device among the multiple electronic devices; the noise reduction module or unit is used to support the first electronic device to perform noise reduction processing on the first voice to be recognized picked up according to the audio information received by the receiving unit or module to obtain a second voice to be recognized.

[0035] In a sixth aspect, the present application provides a readable medium having instructions stored thereon, which, when executed on an electronic device, causes the electronic device to execute the multi-device-based voice processing method in the first to fourth aspects above.

[0036] In a seventh aspect, the present application provides an electronic device comprising: one or more processors; one or more memories; the one or more memories storing one or more programs, which, when executed by the one or more processors, enable the electronic device to perform the multi-device-based voice processing method described in the first to fourth aspects above. In one possible implementation, the electronic device may further include a transceiver (which may be a separate or integrated receiver and transmitter) for receiving and transmitting signals or data.

[0037] In an eighth aspect, the present application provides an electronic device comprising: a processor, a memory, a communication interface and a communication bus; the memory is used to store at least one instruction, and the at least one processor, the memory and the communication interface are connected through the communication bus. When the at least one processor executes the at least one instruction stored in the memory, the electronic device executes the multi-device-based voice processing method in the above-mentioned first to fourth aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A schematic diagram of a multi-device voice processing scenario provided in an embodiment of the present application;

[0039] Figure 2 A schematic diagram of a voice assistant interactive conversation process provided in an embodiment of the present application;

[0040] Figure 3 A schematic diagram of another scenario of multi-device voice processing provided in an embodiment of the present application;

[0041] Figure 4 A flowchart of a multi-device voice processing method provided in an embodiment of the present application;

[0042] Figure 5 A flowchart of another method for multi-device voice processing provided in an embodiment of the present application;

[0043] Figure 6 A schematic diagram of another scenario of multi-device voice processing provided in an embodiment of the present application;

[0044] Figure 7 A flowchart of another method for multi-device voice processing provided in an embodiment of the present application;

[0045] Figure 8 A schematic diagram of another scenario of multi-device voice processing provided in an embodiment of the present application;

[0046] Figure 9 A flowchart of another method for multi-device voice processing provided in an embodiment of the present application;

[0047] Figure 10 A schematic diagram of another scenario of multi-device voice processing provided in an embodiment of the present application;

[0048] Figure 11 A flowchart of another method for multi-device voice processing provided in an embodiment of the present application;

[0049] Figure 12 According to some embodiments of the present application, a schematic structural diagram of an electronic device is shown. DETAILED DESCRIPTION

[0050] The illustrative embodiments of the present application include but are not limited to a multi-device-based voice processing method, medium, and electronic device. The following describes in detail the multi-device scenario of the multi-device-based voice processing application provided by the embodiments of the present application in conjunction with the accompanying drawings.

[0051] Figure 1 The figure shows a multi-device scenario of a multi-device voice processing application provided by an embodiment of the present application. Figure 1 As shown, for ease of explanation, the multi-device scenario 10 only shows three electronic devices, such as electronic device 101, electronic device 102, and electronic device 103. However, it can be understood that the multi-device scenario to which the technical solution of the present application is applicable can include any number of electronic devices, not limited to three.

[0052] Specifically, continue to refer to Figure 1 After the user says the wake-up word, an answering device can be selected from multiple electronic devices, for example, electronic device 101 is selected as the answering device. The answering device then selects the sound pickup device with the best sound pickup effect (such as the electronic device with the best voice enhancement effect) from multiple devices. For example, electronic device 101 selects electronic device 103 as the sound pickup device. Furthermore, after the sound pickup device (such as electronic device 103) picks up the voice data corresponding to the user's voice command, the answering device (such as electronic device 101) can receive, recognize and respond to the voice data, so that the quality of the voice data processed by the answering device is better. In addition, in this scenario, if there is an internal noise device that plays audio outside near the sound pickup device, the voice data picked up by the sound pickup device can be noise-reduced based on the noise reduction information of the internal noise device, further improving the quality of the voice data processed by the answering device. Thus, even if the answering device selected in the multi-device scenario (such as the electronic device closest to the user) has a poor sound pickup effect, or there is noise generated by an electronic device that is playing audio outside, multiple devices can collaboratively pick up and recognize voice data with better audio quality without the user having to move or manually control a specific electronic device to pick up the sound. Furthermore, it is helpful to improve the accuracy of voice recognition during voice control and enhance the user experience.

[0053] In some embodiments, the electronic devices 101-103 in the multi-device scenario 10 are interconnected via a wireless network, for example, Wi-Fi (such as Wireless Fidelity), Bluetooth (BT), Near Field Communication (NFC), and other wireless networks, but not limited thereto. As an example, in order to achieve interconnection between the electronic devices 101-103 via a wireless network, the electronic devices 101-103 meet at least one of the following conditions:

[0054] 1) Connect to the same wireless access point (such as a Wi-Fi access point);

[0055] 2) Logged in with the same account;

[0056] 3) Being set in the same group of devices, for example, the same group of devices all have identification information of each device, so that the devices in the group can communicate with each other according to their respective identification information.

[0057] It is understood that different electronic devices can transmit information in a broadcast or point-to-point manner through an interconnected wireless network, but is not limited thereto.

[0058] According to some embodiments of the present application, the types of wireless networks between different electronic devices in a multi-device scenario can be the same or different. For example, electronic device 101 is connected to electronic device 102 via a Wi-Fi network, while electronic device 101 is connected to electronic device 103 via Bluetooth.

[0059] In each embodiment of the present application, the types of electronic devices in the multi-device scenario may be the same or different. For example, the electronic devices applicable to the present application may include but are not limited to mobile phones, tablet computers, desktop computers, laptop computers, handheld computers, notebook computers, desktop computers, ultra-mobile personal computers (UMPCs), netbooks, as well as cellular phones, personal digital assistants (PDAs), augmented reality (AR) and virtual reality (VR) devices, media players, smart TVs, smart speakers, smart watches, smart headphones, etc. As an example, Figure 1 The types of electronic devices 101-103 shown are all different, and are respectively shown as mobile phones, tablet computers, and smart TVs. In addition, the embodiments of the present application do not impose any special restrictions on the specific form of the electronic devices. The specific structure of the electronic devices can be referred to in the following Figure 12 The corresponding description is not repeated here.

[0060] It is understood that in some embodiments of the present application, all electronic devices in the multi-device scenario have voice control functions, for example, all are installed with voice assistants with the same wake-up word, such as the wake-up word is "Xiaoyi Xiaoyi". In addition, all electronic devices in the multi-device scenario are within the effective working range of the voice assistant, such as the distance between them and the user (i.e., the pickup distance) is less than or equal to the preset distance (such as 5m), the screen is in use (such as the screen is placed face up or the screen cover is not closed), Bluetooth is not turned off, and the Bluetooth communication range is not exceeded, etc., but not limited to this.

[0061] As you can understand, a voice assistant is an AI-based application (APP) that uses speech and semantic recognition algorithms to interact with users through instant question-and-answer voice interactions, helping them complete operations such as information search, device control, and text input. Voice assistants can be system applications in electronic devices or third-party applications.

[0062] like Figure 2 As shown, voice assistants usually adopt a staged cascade processing, which realizes the above functions through processes such as voice wake-up, speech enhancement processing (SE) (or voice front-end processing), automatic speech recognition (ASR), natural language understanding (NLU), dialogue management (DM), natural language generation (NLG), text to speech (TTS), and response output. For example, when the user says the wake-up word "Xiaoyi Xiaoyi" to wake up the voice assistant, after the user says the voice command "What's the weather like in Beijing tomorrow?" or "Play music", the voice command goes through SE, ASR, NLU, DM, NLG, TTS and other processes, which can trigger the electronic device to respond to the voice command.

[0063] It can be understood that the voice data picked up by the electronic device in this application is the voice data directly collected by the microphone, or the voice data processed by SE after collection, which is used to be input into the ASR for processing. Among them, the text processing results of the voice data output by the ASR are the basis for the voice assistant to accurately complete subsequent operations such as recognizing and responding to voice data. Therefore, the quality of the voice data picked up by the voice assistant and input into the ASR will affect the accuracy of the voice assistant in recognizing and responding to the voice data.

[0064] To address the issue of electronic devices' sound pickup being susceptible to various factors and improve the quality of voice data picked up by electronic devices, the present embodiment comprehensively considers various factors and implements a multi-device voice processing process in a multi-device scenario. Generally, factors affecting electronic devices' sound pickup include environmental factors 1)-3) and device factors 4)-6), as shown below:

[0065] 1) The distance or position of the electronic device to the user, that is, the placement of the electronic device. Generally, the closer the electronic device is to the user, the better the sound pickup effect.

[0066] 2) Is there any external noise near the electronic device, such as air conditioning fans, irrelevant human voices, etc. It can be understood that the noise around the electronic device refers to other sounds other than the user's voice commands. Generally, electronic devices farther away from external noise have better sound pickup effects.

[0067] 3) Whether there is internal noise in the electronic device, such as the internal noise represented by the audio played through the electronic device's speakers. Generally speaking, the internal noise of one electronic device can become external noise to other electronic devices, affecting the sound pickup effect of other electronic devices.

[0068] 4) Information about the electronic device's microphone module, such as whether the microphone module is a single microphone or a microphone array, a near-field microphone array or a far-field microphone array, and the microphone module's cutoff frequency. Generally, microphone arrays provide better sound pickup than single microphones. When the user is farther away from the device, far-field microphone arrays provide better sound pickup than near-field microphone arrays. Furthermore, the higher the cutoff frequency of the microphone module, the better the sound pickup.

[0069] 5) The SE capabilities of the electronic device, such as the noise reduction performance of the electronic device's microphone module and the electronic device's AEC capabilities, such as whether the electronic device's AEC is effective. Generally, electronic devices with better microphone module noise reduction performance or effective AEC indicate better SE processing effects, that is, better sound pickup effects. For example, microphone arrays have better noise reduction performance than single microphones.

[0070] 6) Device status of the electronic device, such as one or more factors such as the device's network connection status, headphone connection status, microphone occupancy status, and scenario mode information. For example, if the electronic device is in a device state that is not conducive to the electronic device picking up sound, such as the electronic device's network connection status is poor, a wired or wireless headphone is connected, the microphone is occupied or in airplane mode, it means that the electronic device's sound pickup effect is difficult to guarantee, or the electronic device cannot properly cooperate with other devices to pick up sound, such as being unable to properly transmit the picked-up voice data to other electronic devices.

[0071] Figures 3 to 11 In response to the above-mentioned different influencing factors, various embodiments of collaborative voice processing among multiple electronic devices are proposed.

[0072] Example 1

[0073] Figure 3 The figure shows a scenario where multiple electronic devices at different deployment locations collaborate to process voice. Figure 3 As shown, in this multi-device scenario (referred to as multi-device scenario 11), mobile phone 101a, tablet computer 102a, and smart TV 103a are interconnected via a wireless network and deployed at different distances from the user, for example, 0.3 meters (m), 1.5 meters (m), and 3.0 meters (m) from the user, respectively. In this scenario, mobile phone 101a is held by the user, tablet computer 102a is placed on a desk, and smart TV 103a is mounted on a wall.

[0074] In this multi-device scenario 11, it is assumed that multiple electronic devices are in a low-noise environment with an ambient noise level of ≤20 decibels (dB). In addition, there is no internal noise generated by electronic devices that play audio externally. Therefore, the impact of external and internal noise on the sound pickup effect of electronic devices can be ignored. Instead, the deployment location of the electronic devices, such as which electronic device is closest to the user, is considered. This factor affects the impact of multi-device voice processing.

[0075] Figure 4 yes Figure 3 The specific process of the method for collaborative voice processing in the scenario shown. Figure 4 As shown, the process of the method for collaboratively processing voice by the mobile phone 101a, the tablet computer 102a and the smart TV 103a includes:

[0076] Step 401: The mobile phone 101a, the tablet computer 102a and the smart TV 103a respectively pick up the first voice data corresponding to the wake-up word spoken by the user.

[0077] For example, the pre-registered wake-up word in mobile phone 101a, tablet computer 102a, and smart TV 103a is "Xiaoyi Xiaoyi". After the user says the wake-up word "Xiaoyi Xiaoyi", mobile phone 101a, tablet computer 102a, and smart TV 103a can detect the voice corresponding to "Xiaoyi Xiaoyi" and then determine whether to wake up the corresponding voice assistant.

[0078] It is understood that if a user speaks within the pickup range of an electronic device, the electronic device can detect the corresponding voice data through the microphone and cache it. Specifically, electronic devices such as mobile phone 101a, tablet computer 102a, and smart TV 103a, without other software or hardware using the microphone to pick up voice data, can use the microphone to monitor in real time whether the user has input voice data and cache the picked-up voice data, such as the first voice data described above.

[0079] Step 402: The mobile phone 101a, the tablet computer 102a, and the smart TV 103a respectively verify the picked-up first voice data to determine whether the corresponding first voice data is a pre-registered wake-up word.

[0080] If the mobile phone 101a, tablet computer 102a, and smart TV 103a all successfully verify the first voice data, it indicates that the first voice data picked up is the wake-up word, and the following step 403 can be executed. If the mobile phone 101a, tablet computer 102a, and smart TV 103a all fail to verify the first voice data, it indicates that the first voice data picked up is not the wake-up word, and the following step 409 can be executed.

[0081] In some embodiments, the electronic devices that have successfully verified the first voice data corresponding to the wake-up word can be recorded in a list. For example, if the mobile phone 101a, the tablet computer 102a and the smart TV 103a all successfully verify the first voice data, the mobile phone 101a, the tablet computer 102a and the smart TV 103a are recorded in a list (such as a candidate answering device list). Then, the devices in the above-mentioned candidate answering device list will be used to participate in the following multi-device answering election to elect an electronic device that wakes up the voice assistant and recognizes the user's voice (i.e., the answering device hereinafter). It can be understood that in an embodiment of the present application, the multi-device answering election is performed between multiple devices that have successfully detected the wake-up word, that is, between the electronic devices that have successfully verified the above-mentioned first voice data.

[0082] Step 403: The mobile phone 101a, tablet computer 102a and smart TV 103a select the smart TV 103a as the answering device.

[0083] In some embodiments, the answering device is generally an electronic device that the user is accustomed to or inclined to use, or an electronic device that has a higher probability of successfully identifying and responding to the user's voice data. Specifically, in a multi-device scenario, the answering device is used to identify and respond to the user's voice data, such as performing ASR, NLU and other processing steps on the voice data. There is usually only one answering device in a multi-device scenario, such as an electronic device in the candidate answering device list. In addition, after the electronic device (such as the smart TV 103a) wakes up the voice assistant as an answering device, it can play a wake-up response tone, such as "I'm here". In the multi-device scenario, electronic devices other than the answering device, such as mobile phones 101a and tablet computers 102a, do not respond according to the candidate pickup indication, that is, they do not output a wake-up response tone.

[0084] The selection of the answering device may be performed using various existing technologies, which will be described in detail below.

[0085] In some embodiments, the answering device (such as the smart TV 103a) in a multi-device scenario can perform a collaborative sound pickup election to select a sound pickup device, and specifically perform the following step 404.

[0086] Step 404: the smart TV 103a obtains the sound pickup election information corresponding to the mobile phone 101a, the tablet computer 102a and the smart TV 103a respectively, and selects the mobile phone 101a as the sound pickup device according to the sound pickup election information.

[0087] Among them, the sound pickup election information can be a parameter used to determine the quality of the sound pickup effect of each electronic device. For example, in some embodiments, the sound pickup election information may include the sound information of the detected user voice (such as the sound information of the first voice data mentioned above), the microphone module information of each electronic device, the device status information of each electronic device, and at least one of the AEC capability information of each electronic device. In addition, it can be understood that the information used for sound pickup device election can also include other information, as long as the information can be used to evaluate the sound pickup function of the electronic device, and there is no limitation here.

[0088] Sound information can include signal-to-noise ratio (SNR), sound intensity (or energy), and reverberation parameters (such as reverberation delay). Furthermore, the higher the SNR, the higher the sound intensity, and the lower the reverberation delay of the user's voice picked up by the electronic device, the better the audio quality of the user's voice, and therefore, the better the sound pickup effect of the electronic device. Therefore, the sound information of the user's voice can be used to select the pickup device.

[0089] In addition, the microphone module information is used to indicate whether the microphone module of the electronic device is a single microphone or a microphone array, whether it is a near-field microphone array or a far-field microphone array, and what the cutoff frequency of the microphone module is. Generally, when the distance between the human and the machine is far, the noise reduction capability of the far-field microphone is higher than that of the near-field microphone, so the far-field microphone has a better sound pickup effect than the near-field microphone. The noise reduction capability of the single microphone, linear array microphone, and ring array microphone increases in turn, and the sound pickup effect of the corresponding electronic device increases in turn. In addition, the higher the cutoff frequency of the microphone module, the better the noise reduction capability, and the better the sound pickup effect of the corresponding electronic device. Therefore, the microphone module information can also be used to select the sound pickup device.

[0090] Device status information refers to the device status that can affect the sound pickup effect of multiple electronic devices in collaborative sound pickup, such as network connection status, headphone connection status, microphone occupancy status, scenario mode information, etc. Among them, scenario modes include: driving mode, riding mode (such as bus mode, high-speed rail mode or airplane mode, etc.), walking mode, sports mode, home mode and other modes. These scenario modes can be automatically determined by the electronic device reading and analyzing the sensor information, short messages or emails, setting information or historical operation records of the electronic device. The sensor information is the Global Positioning System (GPS), inertial sensor, camera or microphone, etc. It can be understood that if the headphone connection status is in the occupied state, it means that the electronic device is being used by the user, then the headphone microphone closer to the user is supported to pick up sound; if the microphone occupancy status indicates that the microphone module is in the occupied state, then the electronic device may not be able to pick up sound through the microphone module; if the network connection status indicates that the wireless network of the electronic device is poor, then the success rate of the electronic device transmitting information through the wireless network, such as the success rate of sending sound pickup election information to the answering device, is affected. If the scenario mode is driving mode, riding mode, or other scenario modes, then the stability and / or connection rate of the electronic device's wireless network connection may be low, which in turn affects the success rate of the electronic device's participation in the sound pickup election process or the collaborative sound pickup process. Therefore, the above-mentioned device status information can also be used to select the sound pickup device.

[0091] The AEC capability information is used to indicate whether the electronic device has the AEC capability and whether the AEC of the electronic device is effective. Among them, the AEC capability is specifically the AEC capability of the microphone module in the electronic device. It can be understood that compared with electronic devices where AEC is not effective or does not have AEC capability, the electronic devices where AEC is effective have better SE processing capabilities and better noise reduction performance, and thus better sound pickup effects. Therefore, the above-mentioned AEC capability information can also be used to select sound pickup devices. In addition, electronic devices where AEC is effective are usually electronic devices that are playing audio externally.

[0092] As you can understand, AEC is a voice enhancement technology that eliminates the noise generated by the air return path between the microphone and the speaker through sound wave interference. This effectively alleviates the noise interference caused by the speaker playing audio or the spatial reflection of sound waves, thereby improving the quality of the voice data picked up by the electronic device. In addition, SE is used to pre-process the user voice data collected by the electronic device's microphone through hardware or software means, using audio signal processing algorithms such as reverberation elimination, AEC, blind source separation, and beamforming to improve the quality of the obtained voice data.

[0093] The smart TV 103a can select a sound pickup device based on the sound pickup election information of each electronic device. The specific election scheme will be described in detail below. For ease of explanation, it is assumed that the smart TV 103a selects the mobile phone 101a as the sound pickup device.

[0094] It can be understood that in the embodiment of the present application, remote peripheral virtualization technology can be used to use the sound pickup device or the microphone of the sound pickup device as a virtual peripheral node of the answering device, which is called by the voice assistant running on the answering device to complete the subsequent cross-device sound pickup process.

[0095] In addition, in some embodiments, after the answering device determines that an electronic device is a sound pickup device, it can send a sound pickup instruction to the electronic device to instruct the electronic device to pick up the user's voice data. Similarly, the answering device can send a stop sound pickup instruction to other electronic devices in the multi-device scenario except the sound pickup device to instruct these electronic devices to stop picking up the user's voice data. Alternatively, if the other electronic devices in the multi-device scenario except the sound pickup device do not receive any instructions within a period of time (e.g., 5 seconds) after sending the sound pickup selection information to the answering device, these electronic devices determine that they are not sound pickup devices.

[0096] Step 405: The mobile phone 101a picks up the second voice data corresponding to the voice command spoken by the user.

[0097] It is understood that in subsequent applications, mobile phone 101 acts as a sound pickup device to pick up various voice commands spoken by the user. For example, if a user speaks the voice command "What's the weather like in Beijing tomorrow?", mobile phone 101a directly collects the voice command through the microphone module to obtain second voice data, or the microphone module in mobile phone 101a collects the voice command and obtains the second voice data after SE processing.

[0098] For the convenience of description, the "voice command" that appears separately in the embodiments of the present application can be a voice command corresponding to a certain event or operation received by the electronic device after waking up the voice assistant. For example, the user's voice command is the above-mentioned "What's the weather like tomorrow?" or "Play music". In addition, the names "voice", "voice command" and "voice data" can sometimes be used interchangeably in this document. It should be pointed out that when the distinction is not emphasized, the meanings they intend to express are the same.

[0099] Step 406: The mobile phone 101a sends the second voice data to the smart TV 103a.

[0100] It can be understood that the mobile phone 101a, as a sound pickup device, directly forwards the voice data of the voice instruction to the answering device after picking up the voice instruction issued by the user, but does not make any recognition or response to the voice instruction issued by the user.

[0101] In addition, it is understood that in other embodiments, if the answering device and the sound pickup device are the same device, this step is not required. After picking up the user's voice command, the answering device or the sound pickup device directly performs voice recognition on the voice data.

[0102] Step 407: The smart TV 103a recognizes the second voice data.

[0103] Specifically, after the smart TV 103a as an answering device receives the voice data picked up by the mobile phone 101a, it can recognize the second voice data after noise reduction processing through the ASR, NLU, DM, NLG, and TTS cascade processing flow.

[0104] For example, for the voice command mentioned above, "What will the weather be like in Beijing tomorrow?", ASR can convert the second voice data processed by SE into corresponding text (or words), and perform text processing such as normalization, error correction, and writing on the spoken text, for example, to obtain the text "What will the weather be like in Beijing tomorrow?".

[0105] Step 408: The smart TV 103a responds to the user's voice command or controls other electronic devices to respond to the user's voice command based on the recognition result.

[0106] It is understood that in the embodiment of the present application, for a recognized user's voice command, if the answering device can execute it or can only execute it, the answering device will respond to the voice command. For example, for the aforementioned voice command "What's the weather like in Beijing tomorrow?", the smart TV 103a responds "It will be sunny in Beijing tomorrow." For the voice command "Please turn off the TV," the smart TV 103a executes the shutdown function.

[0107] It's understandable that the voice message "Tomorrow Beijing will be sunny" is the response voice output by the answering device via TTS. Furthermore, the answering device can also control system software, display screens, vibration motors, and other hardware and software to perform response operations, such as displaying the NLG-generated response text on the display screen.

[0108] For voice commands to other electronic devices, the response device can send the command to the corresponding electronic device after recognizing the voice command. For example, for the voice command "open the curtains", after the smart TV 103a recognizes that the response operation is to open the curtains, it can send the operation instruction to open the curtains to the smart curtains, so that the smart curtains can complete the action of opening the curtains through hardware.

[0109] It is understood that the aforementioned other electronic devices may be Internet of Things (IoT) devices, such as smart home devices such as smart refrigerators, smart water heaters, and smart curtains. In some embodiments, the aforementioned other electronic devices do not have voice control functions, such as not having a voice assistant installed. The other electronic devices execute the operations corresponding to the user's voice commands when triggered by the answering device.

[0110] Furthermore, in a multi-device scenario, after the user speaks the voice command corresponding to the second voice data, they can continue to speak subsequent voice command data streams, such as the voice command "What should I wear tomorrow?" The multi-device scenario's collaborative voice processing flow for these data streams can be found in the description of the second voice data above and will not be further elaborated here.

[0111] Step 409: the mobile phone 101a, the tablet computer 102a, and the smart TV 103a do not respond to the first voice data, and delete the cached first voice data.

[0112] For example, when the mobile phone 101a, tablet computer 102a, and smart TV 103a execute step 409, they will not output the wake-up response voice "I'm here" to the user. Of course, if the user continues to speak a voice command, such as "What's the weather like in Beijing tomorrow?", these devices will not respond to the voice data corresponding to the voice command.

[0113] It is understood that if some of the electronic devices among the mobile phone 101a, tablet computer 102a, and smart TV 103a successfully verify the first voice data, while others fail to verify the first voice data, then only the former will continue to execute the subsequent multi-device collaborative sound pickup process. For example, if the mobile phone 101a and tablet computer 102a successfully verify the first voice data, while the smart TV 103a fails to verify the first voice data, then the execution entities of the above-mentioned step 403 will be replaced by the mobile phone 101a and tablet computer 102a, and the execution entity of step 409 will be replaced by the smart TV 103a.

[0114] As described above, in the multi-device scenario of the embodiment of the present application, after the user speaks a voice command, the user does not need to specifically operate a certain electronic device to pick up the voice command (such as the voice command corresponding to the second voice data). Instead, the answering device automatically uses the pickup device as a peripheral device to pick up the user's voice command, and then realizes the voice control function through the answering device's response to the user's voice command.

[0115] The multi-device voice processing method provided in the embodiment of the present application can select the electronic device with the best audio quality for picking up voice commands as the pickup device through the interactive collaboration of multiple electronic devices, so as to support the answering device to complete the voice interaction process with the user through the voice assistant. For example, the pickup device can be the electronic device closest to the user and has better SE processing capabilities. In this way, the peripheral resources of multiple electronic devices equipped with microphone modules and voice assistants can be effectively aggregated, alleviating the impact of the deployment location of electronic devices on the recognition accuracy of voice assistants in multi-device scenarios, thereby improving the user interaction experience and the environmental robustness of voice recognition in multi-device scenarios.

[0116] The following will specifically introduce the selection scheme of the answering device and the selection scheme of the pickup device in the embodiment of the present application.

[0117] Answering machine election

[0118] In some embodiments, for step 403 above, the electronic devices in the multi-device scenario may perform multi-device response election to select a response device according to at least one of the following response election strategies:

[0119] Answering strategy 1) Select the electronic device closest to the user as the answering device.

[0120] For example, for Figure 3 In the scenario shown, mobile phone 101a can be selected as the answering device. The distance between the electronic device and the user can be represented by the sound information of the voice data corresponding to the wake-up word picked up by the electronic device. For example, the higher the signal-to-noise ratio of the first voice data, the higher the sound intensity, and the lower the reverberation delay, the closer the electronic device is to the user.

[0121] Answering strategy 2) Selecting the electronic device actively used by the user as the answering device.

[0122] It can be understood that if the electronic device is actively used by the user, for example, the user has recently lifted the screen, it means that the user may be using the electronic device and the user is more inclined to use it to recognize and respond to the user's voice data.

[0123] In some embodiments, device usage record information can be used to indicate whether an electronic device is actively used by a user. The device usage record information includes at least one of the following: screen on time, screen on frequency, frequency of using a voice assistant, etc. It is understood that the longer the screen on time, the higher the screen on frequency, and the higher the frequency of using a voice assistant, the more actively the user uses the electronic device. For example, based on the device usage record information of the mobile phone 101a, tablet computer 102a, and smart TV 103a, the smart TV 103a that is actively used by the user can be selected as the answering device.

[0124] Answering strategy 3) Select an electronic device equipped with a far-field microphone array as the answering device.

[0125] It can be understood that most electronic devices equipped with far-field microphone arrays are public devices, that is, electronic devices that users prefer to use to recognize and respond to their voice data. Among them, public devices are usually used by users at a long distance (such as 1-3m) and in multiple orientations, and support shared use by multiple people, such as smart TVs or smart speakers. Compared with small electronic devices such as mobile phones and tablets, electronic devices equipped with far-field microphone arrays usually have better speaker performance and larger screen size, so the response voice output or displayed response information for the user's voice command is better. Therefore, electronic devices equipped with far-field microphone arrays are suitable as answering devices.

[0126] In some embodiments, whether an electronic device is equipped with a far-field microphone array is represented by microphone module information. For example, based on the microphone module information, the mobile phone 101a, tablet computer 102a, and smart TV 103a select the smart TV 103a equipped with a far-field microphone array as the answering device.

[0127] Response strategy 4) Elect the public device as the response device.

[0128] In some embodiments, whether an electronic device is a public device can also be indicated by public device indication information. As an example, if the public device indication information of smart TV 103a indicates that smart TV 103a is a public device, multi-device scenario 11 selects smart TV 103a as the response device. Similarly, the remaining descriptions of response strategy 4) can be found in the description of response strategy 3) and are not repeated here.

[0129] If two or more electronic devices in a multi-device scenario all meet the same answer selection policy, any one of these electronic devices may be selected as the answering device.

[0130] It is understood that for the description of a response device simultaneously satisfying response strategies 1) to 4) above, reference can be made to the description of the response device satisfying each response strategy in response conditions 1) to 4) above, and no further details will be given. In some embodiments, different priorities can be pre-set for different response strategies. If, in a multi-device scenario, one electronic device satisfies the highest priority response condition and another electronic device satisfies a lower priority response condition, the former electronic device will be used as the response device.

[0131] In other embodiments, in addition to the answering selection strategies listed above, any electronic device that successfully verifies the first voice data, that is, any electronic device in the candidate answering device list, may be selected as the answering device.

[0132] In some embodiments, any electronic device in a multi-device scenario can act as a master device to perform the steps of electing an answering device. For example, the mobile phone 101a, as the master device, elects the smart TV 103a as the answering device and sends an answer indication to the smart TV 103a to instruct the smart TV 103a to subsequently recognize and respond to the voice data corresponding to the user's voice command. In addition, the master device can send candidate pickup indications to other electronic devices in the multi-device scenario except the answering device to instruct these electronic devices not to recognize the user's voice command. Alternatively, if other electronic devices in the multi-device scenario except the answering device do not receive any indication within a preset time (such as 10 seconds) after the first voice data is successfully verified, these electronic devices determine that they are not answering devices.

[0133] In addition, in other embodiments, each electronic device in a multi-device scenario can perform an operation to elect a response device. For example, the mobile phone 101a, tablet computer 102a, and smart TV 103a all perform a multi-device response election and each elect the smart TV 103a as the response device. Then, the smart TV 103a can determine that it is the response device and then wake up the voice assistant to recognize and respond to the voice data corresponding to the user's voice command. Similarly, the mobile phone 101a and tablet computer 102a each determine that they are not response devices and do not recognize or respond to the user's voice command.

[0134] In some embodiments, the electronic device that performs multi-device response election obtains response election information of each electronic device in the multi-device scenario and selects a response device according to the response election information.

[0135] For example, the response selection information of an electronic device includes at least one of the following: sound information of the first voice data, device usage record information, microphone module information, and public device indication information, but is not limited thereto.

[0136] In addition, after obtaining the answering election information of each electronic device, the answering device can cache the information.

[0137] Selection of pickup equipment

[0138] Specifically, in the above step 404, the smart TV 103a can receive the corresponding sound pickup election information sent by the mobile phone 101a and the tablet computer 102a respectively, and read its own sound pickup election information.

[0139] It should be noted that the embodiment of the present application does not limit the order in which the sound pickup election information corresponding to the mobile phone 101a and the tablet computer 102a is sent, and the order in which different information in each sound pickup election information is sent. Any feasible sending order can be used.

[0140] In addition, for the sound pickup selection information of each electronic device, if the answering device has calculated and cached some information of each electronic device in the above step 403, such as the sound information of the first voice data, then the cached information can be read in step 404 without recalculating the information.

[0141] Specifically, the embodiments of the present application can comprehensively consider different information in the sound pickup election information corresponding to the electronic device, that is, different factors affecting the sound pickup effect of the electronic device, and set a sound pickup election strategy to select the electronic device with better sound pickup effect in the multi-device scenario as the sound pickup device.

[0142] It can be understood that in an embodiment of the present application, the multi-device sound pickup election is performed between multiple devices that successfully detect the wake-up word, that is, between electronic devices that successfully verify the above-mentioned first voice data. Specifically, the devices in the above-mentioned candidate answering device list can be used to participate in the multi-device sound pickup election to elect a sound pickup device. At this time, the candidate answering device list can be referred to as a candidate sound pickup device list. Specifically, in the process of executing the multi-device sound pickup election, the electronic devices in the above-mentioned candidate sound pickup device list can all be used as candidate sound pickup devices, such as the above-mentioned mobile phone 101a, tablet computer 102a and smart TV 103a can all be used as candidate sound pickup devices, that is, electronic devices that perform sound pickup elections based on the sound pickup election information.

[0143] In some embodiments, an end-to-end method such as an artificial neural network or an expert system can be used to employ a sound pickup election strategy to select electronic devices with better sound pickup effects from the candidate sound pickup device list as sound pickup devices. Specifically, the sound pickup election information corresponding to each electronic device in the candidate sound pickup device list is used as input to the artificial neural network or expert system, and the output of the artificial neural network or expert system is the sound pickup device. For example, if the sound pickup election information corresponding to mobile phone 101a, tablet computer 102a, and smart TV 103a is used as input to the artificial neural network or expert system, the output of the artificial neural network or expert system is mobile phone 101a, i.e., mobile phone 101a is selected as the sound pickup device.

[0144] The above-mentioned artificial neural network can be a deep neural network (DNN), a convolutional neural network (CNN), a long short-term memory network (LSTM) or a recurrent neural network (RNN), etc., and the embodiments of the present application do not make specific limitations on this.

[0145] In addition, in other embodiments, a staged cascade processing method can be used to implement a sound pickup election strategy to select electronic devices with better sound pickup effects in the candidate sound pickup device list as sound pickup devices. Specifically, each parameter vector in the sound pickup election information corresponding to each electronic device in the candidate sound pickup device list (i.e., each sound pickup election information) can be feature extracted or numerically quantized, and then a decision tree, logistic regression, or other algorithm can be used to output the sound pickup device selection result. For example, through a staged cascade processing method, each parameter vector in the sound pickup election information corresponding to the mobile phone 101a, tablet computer 102a, and smart TV 103a can be feature extracted or numerically quantized, and then a decision tree, logistic regression, or other algorithm can be used to output the sound pickup device selection result as mobile phone 101a, i.e., mobile phone 101a is selected as the sound pickup device.

[0146] Specifically, in some embodiments, the answering device can perform a multi-device pickup election process through at least one of the first type of collaborative pickup strategy and the second type of collaborative pickup strategy. For example, the process may include a two-part process. The first part of the process is that the answering device first removes some inferior devices that are obviously unsuitable for participating in subsequent collaborative pickup from the candidate pickup device list through the first type of collaborative pickup strategy, or directly decides to select the answering device as the most suitable pickup device. The second part of the process is that the answering device selects an electronic device with better pickup effect as the pickup device based on the pickup election information corresponding to each electronic device in the candidate pickup device list through the second type of collaborative pickup strategy. It can be understood that if the pickup device is not determined by executing the first part of the process, the second part of the process will be executed to select the pickup device.

[0147] In some embodiments, the first type of collaborative sound pickup strategy may include at least one of the following strategies a1) to a6).

[0148] a1) An electronic device that is connected to a headset and is not an answering device is determined as a non-candidate voice pickup device.

[0149] Among them, the status of the electronic device being connected to headphones is indicated by the headphone connection status information. Specifically, if an electronic device is connected to a wired or wireless headset and is not an answering device, since the electronic device only supports close-range sound pickup through the headphone microphone, then the electronic device has a high probability of being far away from the user or not currently being used by the user. Selecting this device as a sound pickup device may cause speech recognition failure. Therefore, the electronic device is marked as a non-candidate sound pickup device that is not suitable for participating in the multi-device sound pickup election and is removed from the candidate sound pickup device list. It can be understood that non-candidate sound pickup devices will not participate in the multi-device sound pickup election, that is, they will not be selected as sound pickup devices.

[0150] a2) electronic devices that are in a preset network state (ie, a poor network state) and are not answering devices are determined as non-candidate voice pickup devices.

[0151] Among them, the electronic device is in a preset network state indicated by the network connection state information. Specifically, if the network state of an electronic device is poor (such as low network communication rate, weak wireless network signal, frequent network disconnection in recent times, etc.) and it is not an answering device, in order to avoid data loss or delay when the electronic device is called by the answering device, thereby affecting the subsequent collaborative voice pickup and voice interaction process, the electronic device is marked as a non-candidate electronic device that is not suitable for participating in the multi-device voice pickup election and is removed from the candidate voice pickup device list.

[0152] a3) An electronic device whose microphone module is in an occupied state and is not an answering device is determined as a non-candidate pickup device.

[0153] Among them, the microphone module is in an occupied state indicated by the microphone occupancy information. If the microphone module of an electronic device is occupied by an application other than a voice assistant (such as a recorder) and is not an answering device, it will be treated as a non-candidate pickup device and removed from the candidate pickup device list. Specifically, if the microphone module of an electronic device is occupied by other applications, it means that the electronic device may not be able to use the microphone module for pickup, so the electronic device is marked as a device that is not suitable for participating in collaborative pickup.

[0154] a4) Determine the answering device in the preset network state as the sound pickup device.

[0155] Among them, if the network connection status of the answering device is poor, in order to avoid the failure of the answering device to call other candidate pickup devices, the answering device is directly selected as the most suitable pickup device, and the answering device calls the local microphone module as the pickup device for subsequent pickup.

[0156] a5) The answering device connected to the headset is determined as the sound pickup device.

[0157] If the answering device is connected to a wired or wireless headset, then the answering device is most likely to be the device closest to the user or the device the user is currently using, and thus the answering device is directly selected as the sound pickup device.

[0158] a6) determining the electronic device in the preset scene mode as a sound pickup device.

[0159] If the answering device is in a preset scenario mode (such as subway mode, airplane mode, driving mode, or travel mode), the corresponding electronic device can be directly selected as the sound pickup device to ensure system performance. For example, in driving mode, to avoid interference from driving noise, an electronic device with a microphone with good noise reduction capabilities can be fixedly selected as the sound pickup device. For another example, in travel mode, to avoid increased device communication power consumption and reduced battery life, the answering device can be fixedly selected as the sound pickup device.

[0160] In addition, the second type of voice pickup election strategy may include at least one of strategies b1) to b4).

[0161] b1) Use the electronic device with AEC enabled as the sound pickup device.

[0162] That is, the electronic device whose AEC capability information indicates that AEC is enabled in the candidate sound pickup device list is used as the sound pickup device, and the electronic device with AEC enabled has a better sound pickup effect.

[0163] It's understandable that AEC is typically used on electronic devices that are currently playing audio. Furthermore, if an electronic device playing audio does not have AEC capabilities or is not enabled, it can cause significant interference to the device itself, such as severely interfering with its sound pickup. Of course, if the electronic device currently playing audio has internal noise reduction capabilities and AEC is enabled, the impact of internal noise generated by the audio being played on its sound pickup can be eliminated.

[0164] b2) Use electronic devices with good noise reduction capabilities as sound pickup devices.

[0165] That is, electronic devices in the candidate sound pickup device list whose microphone model parameters indicate a microphone module with better noise reduction capabilities are selected as sound pickup devices. For example, when the distance between the user and the device is far or the first voice data is weak, electronic devices equipped with far-field microphone arrays are selected as sound pickup devices. Specifically, by determining whether the microphone module of the electronic device is a near-field microphone or a far-field microphone, the sound pickup device with a far-field microphone with better noise reduction capabilities can be selected.

[0166] b3) Use the electronic device closest to the user as the sound pickup device.

[0167] That is, the electronic device closest to the user in the candidate sound pickup device list is selected as the sound pickup device. Generally, if the voice data corresponding to the user's voice picked up by the candidate sound pickup device (such as the first voice data) has the highest sound intensity, the highest signal-to-noise ratio, and / or the lowest reverberation delay, it means that the electronic device is closest to the user and has the best sound pickup effect.

[0168] b4) Use the electronic device farthest from the external noise source as the sound pickup device.

[0169] That is, the electronic device in the candidate sound pickup device list that is farthest from the external noise source is selected as the sound pickup device. Typically, the voice data corresponding to the user's voice picked up by the candidate sound pickup device in the candidate sound pickup device list (e.g., the first voice data) has the highest sound intensity, the highest signal-to-noise ratio, and / or the lowest reverberation delay, indicating that the electronic device is farthest from the external noise source and has the best sound pickup effect.

[0170] It is understood that the above-mentioned sound pickup election strategies (such as the first-category sound pickup election strategy or the second-category sound pickup election strategy) include, but are not limited to, the above-mentioned examples. Specifically, for the description of simultaneously satisfying multiple of strategies a1) through a6) and strategies b1) through b4) in the above-mentioned sound pickup election strategy, reference can be made to the above-mentioned description of the sound pickup device satisfying each sound pickup election strategy separately, and no further details will be given here.

[0171] In some embodiments, different priorities can be pre-set for different sound pickup election strategies, and the sound pickup device is selected based on the sound pickup election strategy with the highest priority. Of course, the priority of the sound pickup election strategy can be the priority of a single sound pickup election strategy, or it can be the priority of a combination of multiple sound pickup election strategies. For example, if the priority of the combination of strategies b1) and b3) is greater than the priority of strategy b3), in this case, if one electronic device in the candidate sound pickup device list meets strategies b1) and b3), and another electronic device meets strategy b3), the former electronic device can be selected as the sound pickup device.

[0172] For example, in multi-device scenario 11, smart TV 103a acts as the answering device. In a low-noise environment free from external noise interference, mobile phone 101a, tablet computer 102a, and smart TV 103a can be selected based on strategies b2) and b3) above to select mobile phone 101a, which has better SE processing capabilities, the highest voice intensity or signal-to-noise ratio, and the lowest reverberation delay, as the sound pickup device. In this case, mobile phone 101a is the electronic device closest to the user, at a distance of 0.3 meters. This prevents the deployment location of electronic devices in the multi-device scenario from affecting their sound pickup performance.

[0173] Example 2

[0174] Because users will subconsciously increase the volume when uttering the voice wake-up word, and due to factors such as the movement of the user or the electronic device, the sound information such as the sound intensity and signal-to-noise ratio of the voice data of the wake-up word is difficult to accurately express the audio quality of the electronic device's subsequent pickup of the user's voice command. Therefore, the sound information of the voice data corresponding to the voice command spoken by the user after saying the wake-up word can be used as the pickup selection information to select the pickup device.

[0175] Figure 5 A flowchart showing another method for multi-device voice processing is shown. Figure 4 The method flow shown is that a link is added to select a sound pickup device based on the sound information of the voice data of the voice command spoken by the user. Specifically, Figure 5 As shown, the method flow includes:

[0176] Steps 501 to 503 are the same as steps 401 to 403 above, and will not be repeated here.

[0177] Step 504: the smart TV 103a obtains the sound pickup election information corresponding to the mobile phone 101a, the tablet computer 102a and the smart TV 103a respectively.

[0178] Step 505: The smart TV 103a picks up the voice commands spoken by the user from the mobile phone 101a, the tablet computer 102a and the smart TV 103a respectively, and selects the voice data corresponding to the voice within the first duration in the voice command as the third voice data.

[0179] For example, the first duration is X seconds (e.g., 3 seconds). In some embodiments, the voice corresponding to the third voice data is any voice of the first duration in the voice instruction spoken by the user. For example, the voice within the first duration is the voice within the first X seconds of the voice instruction spoken by the user.

[0180] It can be understood that, in the case where the voice within the above-mentioned first duration is the voice within the first X seconds spoken by the user, in step 505, the mobile phone 101a, the tablet computer 102a and the smart TV 103a only pick up the third voice data respectively, and do not pick up the voice data corresponding to the voice command spoken by the user after the first X seconds. At this time, the above-mentioned third voice data can be a voice command spoken by the user before speaking the voice command corresponding to the above-mentioned second voice data (such as "What is the weather like in Beijing tomorrow?"). In this way, while the answering device obtains the third voice data more quickly, it avoids each electronic device in the multi-device scenario from executing the step of picking up the voice command spoken by the user for a long time, which leads to a waste of resources of the electronic device.

[0181] In other embodiments, in step 505, the mobile phone 101a, tablet computer 102a, and smart TV 103a each capture voice data corresponding to a complete voice command spoken by the user, such as second voice data corresponding to "What's the weather like in Beijing tomorrow?", and then select third voice data preceding X seconds from the second voice data. In this case, the third voice data may be a segment of the voice command starting with the user speaking the voice command corresponding to the second voice data (e.g., "What's the weather like in Beijing tomorrow?"), for example, the third voice data is "tomorrow."

[0182] Step 506: The smart TV 103a obtains the sound information of the third voice data picked up by the mobile phone 101a, the tablet computer 102a and the smart TV 103a respectively.

[0183] For example, the sound information of the third voice data includes at least one of the following: signal-to-noise ratio, sound intensity (or energy value), reverberation parameters, etc. Generally speaking, the higher the signal-to-noise ratio, the higher the sound intensity, and the lower the reverberation delay of the third voice data detected by the electronic device, the better the quality of the third voice data. The third voice data is closer to the user's voice command, which in turn indicates that the electronic device is closer to the user. Therefore, the sound information of the third voice data can serve as the sound pickup selection information of the selection pickup device.

[0184] In some embodiments, the mobile phone 101a and the tablet computer 102a can each calculate the sound information of the third voice data and then send the sound information of the third voice data to the smart TV 103a. Alternatively, the mobile phone 101a and the tablet computer 102a can each send the detected third voice data to the smart TV 103a, and the smart TV 103a can then calculate the sound information of the third voice data corresponding to the mobile phone 101a and the tablet computer 102a, respectively.

[0185] Step 507: the smart TV 103a adds the sound information of the third voice data to the sound pickup election information, and selects the mobile phone 101a as the sound pickup device according to the sound pickup election information corresponding to the mobile phone 101a, the tablet computer 102a and the smart TV 103a respectively.

[0186] Steps 504 to 507 are similar to step 404 above, and the similarities are not repeated here. The difference is that in step 507 of this embodiment, the smart TV 103a additionally obtains voice data corresponding to the user's voice command spoken for the first X seconds (i.e., third voice data), thereby enabling the smart TV 103a to determine that the sound pickup device is the mobile phone 101a based on the sound information of the third voice data detected by each electronic device.

[0187] Specifically, in step 507, based on the sound information of the third voice data corresponding to the mobile phone 101a, the tablet computer 102a, and the smart TV 103a, the smart TV 103a is determined to determine whether the smart TV 103a satisfies the above-mentioned sound pickup selection strategies b3) and / or b4. Specifically, if the sound information of the third voice data indicates that the smart TV 103a is the electronic device closest to the user or farthest from the noise, then the smart TV 103a is selected as the sound pickup device.

[0188] It is understood that generally, if the third voice data detected in the candidate pickup device list has the highest sound intensity, the highest signal-to-noise ratio, and / or the lowest reverberation delay, it means that the electronic device is closest to the user. In this case, the electronic device detects the best quality of the third voice data and has the best pickup effect.

[0189] Steps 508 to 512 are similar to the above steps 405 to 409 and will not be repeated here.

[0190] In an embodiment of the present application, in a multi-device scenario, not only can a pickup device be selected based on information such as the sound information of the first language data corresponding to the user's wake-up word, but a pickup device can also be selected based on the sound information of the third language data corresponding to the voice within the first duration (such as the first X seconds) of the voice command spoken by the user. In this way, by comprehensively considering the influence of factors such as the movement of the user or electronic device, and by increasing the selection information of the sound information corresponding to the user's voice command to select the pickup device, the accuracy of the pickup device can be further improved, thereby improving the accuracy of voice recognition in a multi-device scenario.

[0191] Example 3

[0192] In some multi-device scenarios, if there is external noise, especially when some electronic devices are at the same distance from the user, or even the same type of electronic devices (such as mobile phones), the distance between the electronic device and the external noise, that is, the impact of external noise on the sound pickup effect of the electronic device, can be considered to select the sound pickup device. It is understandable that if different electronic devices are of the same type and at the same distance from the user, then the sound pickup effect of these electronic devices is the same.

[0193] Specifically, Figure 6 A multi-device scenario based on multi-device voice processing under external noise interference is shown. In this multi-device scenario (recorded as multi-device scenario 12), mobile phone 101b, mobile phone 102b, and smart TV 103b are interconnected through a wireless network and deployed at positions of 1.5m, 1.5m, and 3.0m away from the user, respectively. At this time, mobile phone 101b and mobile phone 102b can be placed idle on the desktop, and smart TV 103b can be wall-mounted on the wall. Among them, in multi-device scenario 12, there is an external noise source 104 near mobile phone 102b. For example, the external noise source can be a running air conditioner or other device that plays audio externally. Therefore, this scenario mainly considers the impact of external noise source 104 on the sound pickup effect of each electronic device, and performs a multi-device-based voice processing process.

[0194] Figure 7 is based on Figure 6 The specific process of the method for collaborative voice processing. Figure 7 As shown, the process of the method for collaboratively processing voice among the mobile phone 101b, the mobile phone 102b, and the smart TV 103b includes:

[0195] Steps 701 to 709 are similar to the above-mentioned steps 401 to 409, and the similarities are not repeated here.

[0196] The only difference is that the execution entities have changed. In multi-device scenario 12, the electronic devices connected via the wireless network have changed from mobile phone 101a, tablet computer 102a, and smart TV 103a to mobile phone 101b, mobile phone 102b, and smart TV 103b. Specifically, the answering device selected in step 703 is smart TV 103b, and the sound pickup device selected in step 704 is mobile phone 101b.

[0197] In multi-device scenario 12, mobile phone 101b and mobile phone 102b are both 1.5 meters away from the user. Compared to smart TV 103b, which is 3 meters away from the user, mobile phone 101b and mobile phone 102b are both the electronic devices closest to the user. However, in an environment where external noise source 104 is present near mobile phone 102b, the only factor that distinguishes the sound pickup performance of mobile phone 101b and mobile phone 102b is the distance from external noise source 104. Clearly, mobile phone 101b is farther away from external noise source 104 than mobile phone 102b. Therefore, unlike multi-device scenario 11, in step 704 of multi-device scenario 12, smart TV 103b acts as the answering device. When there is interference from an external noise source, it can select mobile phone 101b as the sound pickup device based on a sound pickup selection strategy (e.g., strategy b4), which is farther away from the external noise source, has the highest voice intensity or signal-to-noise ratio, and the lowest reverberation delay. In this way, the influence of external noise on the sound pickup effect of electronic devices in a multi-device scenario can be avoided.

[0198] Similarly, refer to Figure 5 In steps 505-507 shown, the smart TV 103b in the multi-device scenario 12 can also obtain the sound information of the third voice data corresponding to the voice within the first duration in the voice command spoken by the user picked up by each electronic device, and add these sound information to the various sound pickup selection information in step 705, and select the mobile phone 101b that is farther away from the external noise source 104 as the sound pickup device. No further details will be given here.

[0199] Thus, in the multi-device voice processing method provided by the embodiments of the present application, the pickup device can have one or more favorable factors such as being closest to the user, having internal noise reduction capabilities (such as SE processing capabilities), and being far away from external noise sources. This can alleviate the impact of external noise interference on the voice assistant's recognition accuracy in multi-device scenarios, improving the user interaction experience and the environmental robustness of voice recognition in multi-device scenarios.

[0200] Example 4

[0201] In multi-device scenarios, internal noise, such as the noise generated by an external audio device, can reach 60-80dB and significantly interfere with other nearby devices' ability to pick up voice commands. Consider the impact of this internal noise on multi-device collaborative audio pickup. For example, consider using the external audio device as the pickup device, enabling multi-device audio selection.

[0202] Specifically, Figure 8 The diagram shows a scenario of multi-device voice processing under internal noise interference. In this multi-device scenario (denoted as multi-device scenario 13), a mobile phone 101c, a tablet computer 102c, and a smart TV 103c are interconnected via a wireless network and are respectively deployed at positions 0.3m, 1.5m, and 3.0m away from the user.

[0203] Smart TV 103c is playing audio externally and has internal noise reduction (i.e., noise reduction capabilities) or AEC capabilities. For example, if the volume of the audio played by Smart TV 103c is 60-80dB, it will significantly interfere with the sound pickup of mobile phone 101c and tablet computer 102c. Therefore, in this scenario, the impact of the internal noise of Smart TV 103c on the sound pickup of electronic devices is primarily considered, and a multi-device voice processing process is implemented.

[0204] Figure 9 It is Figure 8 The process of the method for collaborative voice processing in a multi-device scenario is shown in Figure 9. As shown in Figure 9, the process of the method for collaborative voice processing by the mobile phone 101c, the tablet computer 102c, and the smart TV 103c includes:

[0205] Steps 901 to 905 are similar to the above-mentioned steps 401 to 405, and the similarities are not repeated here.

[0206] The only difference is that the execution entities have changed. In multi-device scenario 13, the electronic devices connected via the wireless network have changed from mobile phone 101a, tablet computer 102a, and smart TV 103a to mobile phone 101c, tablet computer 102c, and smart TV 103c. The answering device selected in the collaborative answering election in step 903 is smart TV 103c, and the sound pickup device selected in the collaborative sound pickup election in step 904 is also smart TV 103c. That is, the sound pickup device and the answering device are the same.

[0207] Specifically, multi-device scenario 13 adds consideration to the impact of internal noise on the effects of electronic devices. When smart TV 103c is in the state of playing audio externally, smart TV 103c can select smart TV 103c with relatively high voice signal-to-noise ratio and noise reduction capability as the sound pickup device based on strategy b1) and strategy b2) of the second type of sound pickup election strategy in the above embodiment. For example, in multi-device scenario 13, smart TV 103c has internal noise reduction capability or AEC is effective, while mobile phone 101c and tablet computer 102c do not have internal noise reduction capability, their internal noise reduction capability is lower than that of smart TV 103c, they do not have AEC capability, or AEC is not effective.

[0208] It can be understood that electronic devices with SE processing capabilities, such as electronic devices with internal noise reduction capabilities or AEC capabilities, can eliminate the impact of internal noise (i.e., the audio) on the sound pickup effect through the noise reduction information of the audio when playing audio externally, and obtain better quality voice data.

[0209] In addition, in this embodiment, after the multiple devices collaboratively select the answering device and the sound pickup device, the answering device can also query the internal noise device that is playing audio externally, so that the internal noise device can share its noise reduction information.

[0210] Step 906: The smart TV 103c searches the mobile phone 101c, the tablet computer 102c, and the smart TVs 103c to find the smart TV 103c that is playing audio externally, and provides noise reduction information as an internal noise device.

[0211] The smart TV 103c determines which electronic device is playing audio by querying the speaker occupancy status or audio / video software status of each device (such as whether the audio / video software is open and the volume of the electronic device). For example, if the smart device 103c finds that its own speakers are occupied, the volume is high (such as above 60% of the maximum volume), or the audio / video software is open, then the smart TV 103c is determined to be playing audio and will share the noise reduction information.

[0212] Specifically, the mobile phone 101c and the tablet computer 102c can report information about whether they are in an external audio state to the smart TV 103c via a wireless network, such as reporting information about speaker occupancy, volume, and / or audio / video software status.

[0213] In some embodiments, the smart TV 103c picks up the voice data corresponding to the voice command spoken by the user while continuing to play the audio through the speaker.

[0214] It can be understood that the smart TV 103c, as an answering device and a sound pickup device, can perform noise reduction processing on the voice data corresponding to the subsequent voice commands picked up after it finds out that it is an internal noise device.

[0215] Step 907: The smart TV 103c performs noise reduction processing on the second voice data obtained by picking up the sound according to the noise reduction information.

[0216] Step 908: The smart TV 103c recognizes the second voice data after the noise reduction processing.

[0217] Step 909: The smart TV 103c responds to the user's voice command or controls other electronic devices to respond to the user's voice command based on the recognition result.

[0218] In addition, the above steps 908 and 909 are similar to the above steps 406 to 408, with the only difference being that the voice data recognized by the answering device (i.e., the smart TV 103c) is voice data that has been subjected to noise reduction processing using the above noise reduction information. Specifically, steps 906 and 907 are newly added, i.e., the smart TV 103c, as the answering device, queries the internal noise device, specifically, queries the smart TV 103c that is playing audio externally and provides noise reduction information as an internal noise device. The noise reduction information supports the sound pickup device to perform noise reduction processing on the voice data corresponding to the subsequently picked up voice. Obviously, the sound pickup device and the internal noise device are the same in this scenario.

[0219] It is understood that electronic devices with internal noise reduction capabilities (i.e., noise reduction capabilities) or AEC-enabled devices can introduce audio data from externally played audio into the noise reduction process to mitigate the interference of internal noise generated by the audio being played. That is, the internal noise device obtains noise reduction information based on the audio data of the externally played audio, such as the audio data itself (i.e., internal noise information), or the Voice Activity Detection (VAD) information (or silence suppression information) corresponding to the audio.

[0220] Electronic devices (such as smart TV 103c) can provide noise reduction information for external audio and use its noise reduction information to reduce internal noise, thereby eliminating the impact of the internal noise on other voice data (such as voice data picked up by the user) to improve the quality of the picked up voice data.

[0221] For example, in multi-device scenario 13, if the user directly receives the second voice data after saying the wake-up word, "What's the weather like in Beijing tomorrow?", the second voice data may be directly recognized as "What's the weather like in Beijing tomorrow?" due to the influence of the external audio, thus failing to accurately recognize the user's actual voice command, "What's the weather like in Beijing tomorrow?". In this case, the smart TV 103c can use its internal noise reduction information to eliminate the influence of the external audio, resulting in higher quality of the second voice data after noise reduction processing, and subsequently obtaining an accurate recognition result of the second voice data, "What's the weather like in Beijing tomorrow?"

[0222] Step 910 is similar to the above step 409 and will not be repeated here.

[0223] In addition, in some other embodiments, the answering device can also obtain the noise reduction information of the external audio with internal noise information, and obtain the voice to be recognized by directly picking up the sound by the sound pickup device (that is, the voice to be recognized without eliminating the internal noise of the external audio), and then perform the step of noise reduction processing on the obtained voice to be recognized based on the obtained noise reduction information.

[0224] It can be understood that by using the noise reduction information of the electronic device that plays audio externally, such as the audio data of the external audio itself and / or the VAD information corresponding to the audio, to perform noise reduction processing on the sound pickup process of the pickup device, it is possible to alleviate the impact of the internal noise of the electronic device that plays audio externally on the sound pickup effect of the voice assistant in multi-device scenarios, ensure the sound pickup effect of the voice assistant based on multiple devices, and thus help ensure the voice recognition accuracy of the voice assistant. In turn, it improves the user experience during the voice recognition process and improves the environmental robustness of voice recognition in multi-device scenarios.

[0225] Example 5

[0226] When there is internal noise in a multi-device scenario, in order to avoid the influence of the internal noise on the sound pickup effect of the collaborative sound pickup in the multi-device scenario, not only can the electronic device that plays the audio externally be used as a sound pickup device, but the electronic device that plays the audio externally can also share the noise reduction information of the internal noise with other electronic devices that serve as sound pickup devices, so that the sound pickup devices can eliminate the influence of the internal noise on the sound pickup effect according to the noise reduction information across devices.

[0227] Figure 10Another scenario of multi-device-based voice processing under internal noise interference is shown. The mobile phone 101d and tablet computer 102d in this multi-device scenario (recorded as multi-device scenario 14) are interconnected through a wireless network and deployed at positions 0.3m and 0.6m away from the user, respectively. At this time, the mobile phone 101d is held by the user, and the tablet computer 102d is idle on the desktop. Among them, the tablet computer 102d is in the state of external audio, and has internal noise reduction (i.e., noise reduction capability) or AEC capability. Therefore, in this scenario, the impact of the internal noise of the tablet computer 102d on the pickup effect of collaborative sound pickup in the multi-device scenario can be mainly considered.

[0228] Figure 11 yes Figure 10 The process of the method for collaborative voice processing in a multi-device scenario shown includes:

[0229] Step 1101 - step 1102 are similar to the above-mentioned step 401 - step 402, and the similarities are not repeated here.

[0230] The difference is that in the multi-device scenario 14, the electronic devices interconnected through the wireless network are changed from mobile phone 101c, tablet computer 102c, and smart TV 103c to mobile phone 101d and tablet computer 102d.

[0231] Step 1103: The mobile phone 101d and the tablet computer 102d select the mobile phone 101d as the answering device and the voice pickup device.

[0232] Step 1104: The mobile phone 101d picks up the second voice data corresponding to the voice command spoken by the user.

[0233] Steps 1103-1104 are similar to steps 403-404, except that, after the collaborative answering device is selected in step 1103, the answering device can be directly selected as the pickup device, without executing the steps of selecting a pickup device according to the pickup election strategy in the above embodiment. That is, the answering device and the pickup device are identical, for example, both are mobile phone 101d.

[0234] Step 1105: The mobile phone 101d searches the mobile phone 101d and the tablet computer 102d to find the tablet computer 102d that is playing audio externally, and shares the noise reduction information as an internal noise device.

[0235] Step 1105 is similar to step 906, except that the internal noise device (tablet 102d) and the sound pickup device (mobile phone 101d) that the answering device queries for sharing noise reduction information in step 1105 are different. Therefore, in this embodiment, step 1106 is added to enable the internal noise device to share noise reduction information with the sound pickup device (i.e., mobile phone 101d).

[0236] In addition, it can be understood that in some embodiments, after the tablet computer 102d as the answering device inquires that the mobile phone 101d is an internal noise device, it can send a noise reduction instruction to the mobile phone 101d, so that the mobile phone 101d shares the noise reduction information with the tablet computer 102d as the sound pickup device according to the noise reduction instruction.

[0237] Step 1106: The tablet computer 102d sends the noise reduction information of the tablet computer 102d to the mobile phone 101d.

[0238] It can be understood that by sharing noise reduction information with the mobile phone 101d through the tablet computer 102d, the audio data of the external audio and / or the VAD information corresponding to the audio can be shared across devices, effectively aggregating the peripheral resources of multiple electronic devices equipped with microphone modules and voice assistants.

[0239] Specifically, the tablet computer 102d can send the noise reduction information of the tablet computer 102d to the mobile phone 101d via the wireless network between the tablet computer 102d and the mobile phone 101d.

[0240] Step 1107: The mobile phone 101d performs noise reduction processing on the second voice data obtained by picking up the sound according to the noise reduction information of the tablet computer 102d.

[0241] Step 1108: The mobile phone 101d recognizes the second voice data after the noise reduction processing.

[0242] Step 1109: The mobile phone 101d responds to the user's voice command or controls other electronic devices to respond to the user's voice command based on the recognition result.

[0243] Among them, steps 1107 to 1109 are similar to the above-mentioned steps 907 to 909, with the difference that in step 1107, the sound pickup device (i.e., mobile phone 101d) uses the noise reduction information of other devices (i.e., tablet computer 102d) to perform noise reduction processing on the voice data corresponding to the voice picked up by itself, thereby realizing cross-device noise reduction processing.

[0244] For example, in multi-device scenario 14, when the user speaks the voice command "What's the weather like in Beijing tomorrow?" after saying the wake-up word, and the mobile phone 101d directly picks up the second voice data corresponding to the voice command, the quality of the second voice data is poor due to the influence of the audio played by the tablet computer 102d. As a result, the second voice data may be recognized as "What's the weather like in Beijing tomorrow?", which is different from the user's actual voice command "What's the weather like in Beijing tomorrow?". In other words, the poor quality of the second voice data picked up by the mobile phone 101d causes the subsequent recognition result of the second voice data to be inaccurate. At this time, because the mobile phone 101d can perform noise reduction processing on the picked up second voice data using the noise reduction information shared by the tablet computer 102d, the influence of the audio played by the tablet computer 102d on the sound pickup effect of the mobile phone 101d is eliminated. As a result, the quality of the second voice data after noise reduction processing is higher, and the second voice data is subsequently accurately recognized as "What's the weather like in Beijing tomorrow?"

[0245] It can be understood that by sharing the audio data of the external audio and / or the VAD information corresponding to the audio across devices, the auxiliary pickup device performs noise reduction processing during the pickup process, which can effectively aggregate the peripheral resources of multiple electronic devices equipped with microphone modules and voice assistants, and further improve the accuracy of voice recognition in multi-device scenarios.

[0246] In this way, in the embodiment of the present application, the selected pickup device from multiple devices can have one or more favorable factors such as being closest to the user, furthest from external noise sources, and having internal noise reduction capabilities. This can alleviate the impact of electronic device deployment locations, internal noise interference, or external noise interference on the voice assistant's pickup effect and voice recognition accuracy in multi-device scenarios, thereby improving the user interaction experience and the environmental robustness of voice recognition in multi-device scenarios.

[0247] Figure 12 A schematic structural diagram of the electronic device 100 is shown.

[0248] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0249] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0250] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0251] For example, the processor 110 can be used to detect whether the electronic device 100 has picked up the voice data corresponding to the wake-up word or voice command spoken by the user, and to obtain the sound information, device status information, microphone module information, etc. of the voice data. In addition, based on the information of each electronic device (such as sound pickup election information or answer election information, etc.), the above-mentioned answer device election, sound pickup device election, or internal noise device query can be performed.

[0252] The NPU is a neural-network (NN) computing processor that rapidly processes input information by drawing on biological neural network structures, such as the transmission patterns between neurons in the human brain, and can also continuously self-learn. The NPU can enable applications such as intelligent cognition in the electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension. For example, the NPU can support the electronic device 100 in recognizing voice data picked up by a voice assistant.

[0253] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0254] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0255] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0256] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0257] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0258] For example, the above-mentioned antenna 1, antenna 2, mobile communication module 150, wireless communication module 160 and other modules can be used to support the electronic device 100 to send sound information, device status information, etc. of voice data to other electronic devices in a multi-device scenario, specifically to send the above-mentioned response selection information, sound pickup selection information, noise reduction information, etc.

[0259] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0260] The display screen 194 is used to display images, videos, etc. For example, the display screen 194 can be used to support the electronic device 100 to display a response interface in response to the user's voice command, and the response interface can include information such as a response text.

[0261] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage. For example, files such as music and videos can be stored on the external memory card. For example, the external memory card can be used to support the electronic device 100 in storing the aforementioned sound pickup selection information, response selection information, and noise reduction information.

[0262] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor. For example, an external memory card can be used to support the electronic device 100 in storing the above-mentioned sound pickup election information, answer election information, and noise reduction information, etc.

[0263] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0264] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert analog audio input into a digital audio signal, such as converting the user voice received by the electronic device 100 into a digital audio signal (i.e., voice data corresponding to the user voice), or converting the audio generated by the voice assistant using TTS into a response voice. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0265] Speaker 170A, also known as a "horn," is used to convert audio electrical signals into sound signals. Electronic device 100 can use speaker 170A to listen to music or make hands-free calls, or play a voice response to a user's voice command based on a voice assistant, such as "I'm here" in response to a wake-up word or "It will be sunny in Beijing tomorrow" in response to a voice command like "What's the weather like in Beijing tomorrow?"

[0266] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.

[0267] Microphone (i.e., microphone module) 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals, such as converting the wake-up word or voice command spoken by the user into an electrical signal (i.e., the corresponding voice data). When making a call or sending a voice message, the user can speak by approaching the microphone 170C with their mouth to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction functions. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording functions, etc.

[0268] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0269] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0270] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0271] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed over a network or through other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to floppy disks, optical disks, optical discs, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Therefore, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0272] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.

[0273] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.

[0274] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.

[0275] Although the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the present application.

Claims

1. A multi-device based voice processing method, characterized in that: The method comprises: A first electronic device among the plurality of electronic devices picks up a sound to obtain a first speech to be recognized; The first electronic device receives audio information related to the audio played externally by the second electronic device from the second electronic device among the plurality of electronic devices; the audio information includes at least one of the following: audio data of the played audio, and voice activation detection (VAD) information corresponding to the audio; The first electronic device performs noise reduction processing on the first to-be-recognized voice obtained by picking up the voice according to the received audio information to obtain a second to-be-recognized voice; The first electronic device is an electronic device for picking up sound selected by a third electronic device for voice recognition from among the plurality of electronic devices based on sound picking selection information of the plurality of electronic devices; The sound pickup selection information is used to determine the sound pickup effect of the electronic device. The sound pickup selection information includes device status information and at least one of the following: echo cancellation AEC capability information, microphone module information, voice information corresponding to the wake-up word obtained by the sound pickup, and voice information corresponding to the voice command obtained by the sound pickup; The device status information includes at least one of the following: network connection status information, headphone connection status information, microphone occupancy status information, and scenario mode information; The first electronic device is an electronic device selected by the third electronic device from the plurality of electronic devices based on the device status information; or, The first electronic device is selected by the third electronic device after removing electronic devices not used for sound pickup from the multiple electronic devices based on the device status information, and then based on at least one of the acoustic echo cancellation (AEC) capability information, the microphone module information, and the voice information corresponding to the wake-up word obtained by the sound pickup; Among them, the third electronic device is selected from the multiple electronic devices based on the response election information of the multiple electronic devices, and the response election information includes at least one of the following: voice information corresponding to the wake-up word obtained by picking up sound, device usage record information, microphone module information, and public device indication information.

2. The method according to claim 1, characterized in that The method further comprises: The first electronic device sends the second speech to be recognized to the third electronic device among the plurality of electronic devices for recognizing speech; or The first electronic device recognizes the second voice to be recognized.

3. The method according to claim 2, characterized in that Before a first electronic device among the plurality of electronic devices picks up a sound to obtain a first speech to be recognized, the method further includes: The first electronic device sends sound pickup selection information of the first electronic device to the third electronic device, where the sound pickup selection information of the first electronic device is used to indicate a sound pickup situation of the first electronic device.

4. The method according to claim 3, characterized in that The method further comprises: The first electronic device receives a sound pickup instruction sent by the third electronic device, wherein the sound pickup instruction is used to instruct the first electronic device to pick up sound and send the to-be-recognized speech after noise reduction processing to the third electronic device.

5. The method according to claim 3 or 4, characterized in that The voice command is obtained by picking up the sound after picking up the wake-up word.

6. A multi-device based voice processing method, characterized in that: The method comprises: A second electronic device among the plurality of electronic devices plays audio externally; The second electronic device sends audio information related to the audio to a first electronic device among the multiple electronic devices for picking up sound, wherein the audio information includes at least one of the following: audio data of the audio, and voice activation detection (VAD) information corresponding to the audio; The audio information can be used by the first electronic device to perform noise reduction processing on the audio to be recognized picked up by the first electronic device; The first electronic device is an electronic device for picking up sound selected by a third electronic device for voice recognition from among the plurality of electronic devices based on the acquired sound picking selection information of the plurality of electronic devices; The sound pickup selection information is used to determine the sound pickup effect of the electronic device. The sound pickup selection information includes device status information and at least one of the following: echo cancellation AEC capability information, microphone module information, voice information corresponding to the wake-up word obtained by the sound pickup, and voice information corresponding to the voice command obtained by the sound pickup; The device status information includes at least one of the following: network connection status information, headphone connection status information, microphone occupancy status information, and scenario mode information; The first electronic device is an electronic device selected by the third electronic device from the plurality of electronic devices based on the device status information; or, The first electronic device is selected by the third electronic device after removing electronic devices not used for sound pickup from the multiple electronic devices based on the device status information, and then based on at least one of the acoustic echo cancellation (AEC) capability information, the microphone module information, and the voice information corresponding to the wake-up word obtained by the sound pickup; Among them, the third electronic device is selected from the multiple electronic devices based on the response election information of the multiple electronic devices, and the response election information includes at least one of the following: voice information corresponding to the wake-up word obtained by picking up sound, device usage record information, microphone module information, and public device indication information.

7. The method according to claim 6, characterized in that The method further comprises: The second electronic device receives a sharing instruction from the third electronic device for voice recognition among the plurality of electronic devices; or The second electronic device receives a sharing instruction from the first electronic device; The sharing instruction is used to instruct the second electronic device to send the audio information to the first electronic device.

8. The method according to claim 7, characterized in that Before the second electronic device sends audio information related to the audio to a first electronic device among the plurality of electronic devices for picking up sound, the method further includes: The second electronic device sends the sound pickup selection information of the second electronic device to the third electronic device, where the sound pickup selection information of the second electronic device is used to indicate the sound pickup situation of the second electronic device.

9. A multi-device based voice processing method, characterized in that: The method comprises: A third electronic device among the plurality of electronic devices detects that a second electronic device among the plurality of electronic devices is playing audio; In a case where the second electronic device is different from the third electronic device, the third electronic device sends a sharing instruction to the second electronic device, wherein the sharing instruction is used to instruct the second electronic device to send audio information related to the audio played by the second electronic device to a first electronic device among the multiple electronic devices for sound pickup, wherein the audio information includes at least one of the following: audio data of the audio, and voice activation detection (VAD) information corresponding to the audio; In a case where the second electronic device is the same as the third electronic device, the third electronic device sends the audio information to the first electronic device; The audio information can be used by the first electronic device to perform noise reduction processing on the first speech to be recognized picked up by the first electronic device to obtain the second speech to be recognized; The first electronic device is an electronic device for sound pickup selected by the third electronic device from the plurality of electronic devices based on the sound pickup selection information of the plurality of electronic devices; The sound pickup selection information is used to determine the sound pickup effect of the electronic device. The sound pickup selection information includes device status information and at least one of the following: echo cancellation AEC capability information, microphone module information, voice information corresponding to the wake-up word obtained by the sound pickup, and voice information corresponding to the voice command obtained by the sound pickup; The device status information includes at least one of the following: network connection status information, headphone connection status information, microphone occupancy status information, and scenario mode information; The first electronic device is an electronic device selected by the third electronic device from the plurality of electronic devices based on the device status information; or, The first electronic device is selected by the third electronic device after removing electronic devices not used for sound pickup from the multiple electronic devices based on the device status information, and then based on at least one of the acoustic echo cancellation (AEC) capability information, the microphone module information, and the voice information corresponding to the wake-up word obtained by the sound pickup; Among them, the third electronic device is selected from the multiple electronic devices based on the response election information of the multiple electronic devices, and the response election information includes at least one of the following: voice information corresponding to the wake-up word obtained by picking up sound, device usage record information, microphone module information, and public device indication information.

10. The method according to claim 9, characterized in that The first electronic device is different from the third electronic device, and the method further includes: The third electronic device acquires the second to-be-recognized voice picked up by the first electronic device from the first electronic device; The first electronic device recognizes the second voice to be recognized.

11. The method according to claim 10, characterized in that Before the third electronic device sends the sharing instruction to the second electronic device, the method further includes: The third electronic device acquires the sound pickup selection information of the multiple electronic devices, wherein the sound pickup selection information of the multiple electronic devices is used to represent the sound pickup conditions of the multiple electronic devices; The third electronic device selects an electronic device as the first electronic device from the plurality of electronic devices based on the sound pickup selection information of the plurality of devices.

12. The method according to claim 11, characterized in that The method further comprises: The third electronic device sends a sound pickup instruction to the first electronic device, wherein the sound pickup instruction is used to instruct the first electronic device to pick up sound and send the second to-be-recognized speech obtained by picking up sound to the third electronic device.

13. The method according to claim 11 or 12, characterized in that The voice command is obtained by picking up the sound after picking up the wake-up word.

14. The method according to claim 13, characterized in that The third electronic device selects at least one electronic device from the plurality of electronic devices as the first electronic device based on the sound pickup selection information of the plurality of electronic devices, including at least one of the following: When the third electronic device is in a preset network state, the third electronic device determines the third electronic device as the first electronic device; In a case where the third electronic device is connected to headphones, the third electronic device determines the third electronic device as the first electronic device; The third electronic device determines at least one of the electronic devices in a preset scenario mode among the plurality of electronic devices as the first electronic device.

15. The method according to claim 14, characterized in that The third electronic device selects at least one electronic device from the plurality of electronic devices as the first electronic device based on the sound pickup selection information of the plurality of electronic devices, including at least one of the following: The third electronic device uses at least one of the electronic devices with AEC enabled among the plurality of electronic devices as the first electronic device; The third electronic device selects at least one of the plurality of electronic devices whose noise reduction capability is greater than that of the electronic device that meets the predetermined noise reduction condition as the first electronic device; The third electronic device uses at least one of the plurality of electronic devices, the distance between which the user is less than a first predetermined distance, as the first electronic device; The third electronic device uses at least one of the electronic devices, which is at a distance greater than a second predetermined distance from the external noise source, as the first electronic device.

16. The method according to claim 14, characterized in that The preset network status includes at least one of the following: a network with a communication rate less than or equal to a predetermined rate, and a network line frequency greater than or equal to a predetermined frequency; the preset scenario mode includes at least one of the following: subway mode, flight mode, driving mode, and travel mode.

17. The method according to any one of claims 9, 10, 11, 12, 14, 15 and 16, characterized in that The third electronic device selects the first electronic device from the multiple electronic devices using a neural network algorithm or a decision tree algorithm.

18. A speech processing system, characterized in that: The system includes: a first electronic device and a second electronic device; Wherein, when the second electronic device plays audio externally, it sends audio information related to the audio to the first electronic device for picking up the audio; the audio information includes at least one of the following: audio data of the played audio, voice activation detection (VAD) information corresponding to the audio; The first electronic device is used to pick up a first speech to be recognized, and perform noise reduction processing on the first speech to be recognized obtained by picking up the sound according to the audio information received from the second electronic device to obtain a second speech to be recognized; The system further includes: a third electronic device for speech recognition; The third electronic device is configured to obtain sound pickup selection information of a plurality of electronic devices, wherein the sound pickup selection information of the plurality of electronic devices is used to indicate sound pickup conditions of the plurality of electronic devices; and based on the sound pickup selection information of the plurality of electronic devices, select at least one electronic device from the plurality of electronic devices as the first electronic device for sound pickup; The sound pickup selection information is used to determine the sound pickup effect of the electronic device. The sound pickup selection information includes device status information and at least one of the following: echo cancellation AEC capability information, microphone module information, voice information corresponding to the wake-up word obtained by the sound pickup, and voice information corresponding to the voice command obtained by the sound pickup; The device status information includes at least one of the following: network connection status information, headphone connection status information, microphone occupancy status information, and scenario mode information; The first electronic device is an electronic device selected by the third electronic device from the plurality of electronic devices based on the device status information; or, The first electronic device is selected by the third electronic device after removing electronic devices not used for sound pickup from the multiple electronic devices based on the device status information, and then based on at least one of the acoustic echo cancellation (AEC) capability information, the microphone module information, and the voice information corresponding to the wake-up word obtained by the sound pickup; Among them, the third electronic device is selected from the multiple electronic devices based on the response election information of the multiple electronic devices, and the response election information includes at least one of the following: voice information corresponding to the wake-up word obtained by picking up sound, device usage record information, microphone module information, and public device indication information.

19. The system according to claim 18, wherein The first electronic device, the second electronic device, and the third electronic device are all electronic devices among the plurality of electronic devices, and the third electronic device is the same as or different from the first electronic device; The first electronic device is further configured to send the second speech to be recognized to the third electronic device; and The third electronic device is further configured to recognize the second speech to be recognized obtained from the first electronic device.

20. A computer-readable storage medium, characterized in that The storage medium stores instructions, which, when executed on a computer, enable the computer to execute the multi-device-based speech processing method according to any one of claims 1 to 17.

21. An electronic device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the multi-device based voice processing method according to any one of claims 1 to 17.

22. An electronic device, characterized in that: include: at least one processor, memory, communication interface, and communication bus; The memory is used to store at least one instruction, and the at least one processor, the memory and the communication interface are connected via the communication bus. When the at least one processor executes the at least one instruction stored in the memory, the electronic device executes the multi-device based voice processing method described in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Voice control method and voice control system

    CN105280184A

  • Voice wakeup method and electronic device

    CN110364151A