A wake-up recognition method, an audio device, and an audio device group
By using two audio devices to work together for wake-up recognition and exchanging information using a microphone array to determine the direction of the wake-up source, the problem of inaccurate wake-up recognition in noisy environments by speaker combinations is solved, and the accuracy of wake-up and voice command recognition is improved.
Patent Information
- Application Number
- CN202011556351.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-31
- Filing Date
- 2020-12-23
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2040-12-23
AI Technical Summary
When the speaker system is playing music, environmental noise interference can cause the user's voice wake-up words or voice commands to be inaccurately recognized, affecting the accuracy of wake-up recognition.
Wake-up recognition is achieved through the collaboration of two audio devices. The microphone arrays of the first and second audio devices receive sound signals, perform wake-up recognition, and exchange recognition results. By combining their respective recognition probabilities or angle information, the direction of the wake-up source and whether to wake up the audio device group are determined.
It improves the accuracy of wake-up recognition and voice command recognition, reduces the interference of environmental noise on wake-up recognition, and ensures that the speaker combination can accurately recognize the user's wake-up words and commands even in noisy environments.
Smart Images

Figure CN114121024B_ABST
Abstract
Description
[0001] This application claims priority from the Chinese Patent Application No. 202010893958.3 filed on August 31, 2020, and entitled "A wake-up recognition method and an audio device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the technical field of terminal, in particular to a wake-up recognition method, an audio device and an audio device group. BACKGROUND
[0003] In order to improve the audio playing effect, the combination of two or more audio devices has gradually become a trend. Taking the combination of sound boxes as an example, a user uses the combination of sound boxes to play music in a house (such as a living room). Compared with playing music by a single sound box, playing music by the combination of sound boxes can bring a more extreme auditory experience to the user.
[0004] However, in the process of using the combination of sound boxes, there are some poor experience scenarios: for example, in the process of playing music by the combination of sound boxes, the sound in the environment is large, at this time, if the user utters a voice wake-up word or a voice instruction, the sound information collected by the combination of sound boxes includes all the sound information in the environment. For example, the sound information collected by the combination of sound boxes includes not only the sound uttered by the user, but also the sound generated by the combination of sound boxes. In this way, due to the interference of other sound information, the voice wake-up word or the voice instruction uttered by the user cannot be accurately recognized by the combination of sound boxes. SUMMARY
[0005] The purpose of the present application is to provide a wake-up recognition method, an audio device and an audio device group, which helps to improve the accuracy of wake-up recognition.
[0006] In a first aspect, a wake-up recognition method is provided. The method is applicable to an audio device group, the audio device group including a first audio device and a second audio device, the first audio device being capable of communicating with the second audio device, the first audio device including a first microphone array, and the second audio device including a second microphone array. The method includes: the first audio device receiving a first sound signal, the first sound signal including a sound signal uttered by a wake-up source; the first audio device performing wake-up recognition on the first sound signal to obtain a first recognition result; the second audio device receiving a second sound signal, the second sound signal including a sound signal uttered by the wake-up source; the second audio device performing wake-up recognition on the second sound signal to obtain a second recognition result; the second audio device sending the second recognition result to the first audio device; and the first audio device determining whether to wake up the audio device group based on the first recognition result and the second recognition result. The wake-up recognition method provided by the present application is a wake-up recognition performed by two audio devices in cooperation, which improves the accuracy of wake-up recognition.
[0007] In a possible design, the first recognition result includes a first probability that the first sound signal includes the wake-up information, and the second recognition result includes a second probability that the second sound signal includes the wake-up information; and the first audio device determines whether to wake up the audio device group based on the first recognition result and the second recognition result, including: if the first probability is greater than a first threshold, the second probability is greater than a second threshold, and / or an average or weighted average of the first probability and the second probability is greater than a third threshold, it is determined that the audio device group is woken up. That is, the wake-up recognition is performed by comprehensively considering the wake-up recognition results of the two audio devices, so that the accuracy of the wake-up recognition is improved.
[0008] In a possible design, before the first audio device performs the wake-up recognition on the first sound signal to obtain the first recognition result, the method further includes: the first audio device suppresses the sound signal from the second audio device in the first sound signal; and the first audio device performs the wake-up recognition on the first sound signal to obtain the first recognition result, including: the first audio device performs the wake-up recognition on the first sound signal after the suppression to obtain the first recognition result. That is, the first audio device can suppress the sound from the direction in which the non-wake-up source (such as the second audio device) is located, so as to highlight the sound from the wake-up source, and improve the accuracy of the wake-up recognition.
[0009] In a possible design, before the second audio device performs the wake-up recognition on the second sound signal to obtain the second recognition result, the method further includes: the second audio device suppresses the sound signal from the first audio device in the second sound signal; and the second audio device performs the wake-up recognition on the second sound signal to obtain the second recognition result, including: the second audio device performs the wake-up recognition on the second sound signal after the suppression to obtain the second recognition result. That is, the second audio device can suppress the sound from the direction in which the non-wake-up source (such as the first audio device) is located, so as to highlight the sound from the wake-up source, and improve the accuracy of the wake-up recognition.
[0010] In a possible design, before the first audio device suppresses the sound signal from the second audio device in the first sound signal, the method further includes: the first audio device determines that the second audio device and the wake-up source are in different directions; and before the second audio device suppresses the sound signal from the first audio device in the second sound signal, the method further includes: the second audio device determines that the first audio device and the wake-up source are in different directions. That is, before the first audio device suppresses the sound from the second audio device, it is determined whether the second audio device and the wake-up source are in the same direction, and if not, the sound from the second audio device is suppressed, so as to avoid the case that the sound from the wake-up source is also suppressed when the sound signal in the direction in which the second audio device is located is suppressed in the case that the second audio device and the wake-up source are in the same direction.
[0011] In a possible design, the first recognition result further includes a first angle, the first angle being a direction of the wake-up source relative to the first audio device; the second recognition result further includes a second angle, the second angle being a direction of the wake-up source relative to the second audio device; and the method further includes: determining, by the first audio device, the direction of the wake-up source based on the first angle and the second angle. The wake-up recognition method provided in this application can determine the direction of the wake-up source more accurately by cooperation of the two audio devices.
[0012] In a possible design, the first audio device determines the direction of the wake-up source based on the first angle and the second angle, including: when the first probability is greater than the second probability, determining that the first angle is the direction of the wake-up source; and when the second probability is greater than the first probability, determining that the second angle is the direction of the wake-up source. That is, the two audio devices can determine the direction of the wake-up source more accurately by cooperation.
[0013] In a possible design, the second recognition result further includes a second distance, the second distance being a distance of the wake-up source relative to the second audio device; and the first audio device determines the direction of the wake-up source based on the first angle and the second angle, including: predicting, by the first audio device, a third angle of the wake-up source relative to the first audio device by using the first distance, the second angle and the second distance, where the first distance is a distance between the first audio device and the second audio device; and determining, by the first audio device, the direction of the wake-up source according to the third angle and the first angle. That is, the two audio devices can determine the direction of the wake-up source more accurately by cooperation.
[0014] In a possible design, the third angle satisfies:
[0015]
[0016] where β is the second angle; S c-s is the second distance, D is the first distance, and α' is the third angle.
[0017] In a possible design, the first audio device determines the direction of the wake-up source according to the third angle and the first angle, including: determining an average or a weighted average of the third angle and the first angle as the direction of the wake-up source; or
[0018] The direction of the wake-up source satisfies: α ± (ω i × Δ), ω i is a preset value, Δ = |α-α'|, α is the first angle, and α' is the third angle.
[0019] In a possible design, after the first audio device determines to wake up the audio device group based on the first recognition result and the second recognition result, the method further includes: receiving, by the first audio device, a third sound signal, the third sound signal including a voice instruction; and identifying, by the first audio device, the voice instruction from the sound signal in the third sound signal located in the direction of the wake-up source, the voice instruction being used to control the first audio device. Since the two audio devices cooperate to determine the direction of the wake-up source, the determined direction of the wake-up source is more accurate, and therefore, when performing voice instruction recognition, the voice instruction in the direction of the wake-up source can be recognized, thereby improving the accuracy of voice instruction recognition.
[0020] In a possible design, before the first audio device identifies the voice instruction from the sound signal in the third sound signal located in the direction of the wake-up source, the method further includes: suppressing, by the first audio device, the sound signal in the third sound signal in the direction of the non-wake-up source, the direction of the non-wake-up source being a direction other than the direction of the wake-up source; and identifying, by the first audio device, the voice instruction from the sound signal in the third sound signal located in the direction of the wake-up source, including: identifying, by the first audio device, the voice instruction from the sound signal in the third sound signal located in the direction of the wake-up source after the suppression. That is, when performing voice instruction recognition, the first audio device can suppress the sound in the direction of the non-wake-up source (such as the sound from the second audio device), highlight the sound of the wake-up source, and improve the accuracy of voice instruction recognition.
[0021] In a second aspect, a wake-up recognition method is provided. The method is applicable to a first audio device, and includes: receiving, by the first audio device, a first sound signal, the first sound signal including a sound signal emitted by a wake-up source; performing, by the first audio device, wake-up recognition on the first sound signal to obtain a first recognition result; receiving, by the first audio device, a second recognition result from a second audio device, the second recognition result being a recognition result obtained by performing, by the second audio device, wake-up recognition on a received second sound signal; and determining, by the first audio device, whether to wake up an audio device group based on the first recognition result and the second recognition result.
[0022] In a possible design, the first recognition result includes a first probability, the first probability being a probability that the first sound signal includes wake-up information; and the second recognition result includes a second probability, the second probability being a probability that the second sound signal includes wake-up information.
[0023] Determining, by the first audio device, whether to wake up the audio device group based on the first recognition result and the second recognition result includes: if the first probability is greater than a first threshold value, the second probability is greater than a second threshold value, and / or an average or weighted average of the first probability and the second probability is greater than a third threshold value, determining to wake up the audio device group.
[0024] In a possible design, before the first audio device performs the wake-up recognition on the first sound signal to obtain the first recognition result, the method further includes: the first audio device performs suppression on the sound signal from the second audio device in the first sound signal; and the first audio device performs the wake-up recognition on the first sound signal to obtain the first recognition result, including: the first audio device performs the wake-up recognition on the suppressed first sound signal to obtain the first recognition result.
[0025] In a possible design, before the first audio device performs the suppression on the sound signal from the second audio device in the first sound signal, the method further includes: the first audio device determines that the second audio device and the wake-up source are in different directions.
[0026] In a possible design, the first recognition result further includes a first angle, and the first angle is a direction of the wake-up source relative to the first audio device; the second recognition result further includes a second angle, and the second angle is a direction of the wake-up source relative to the second audio device; and the method further includes:
[0027] The first audio device determines the direction of the wake-up source based on the first angle and the second angle.
[0028] In a possible design, the first audio device determines the direction of the wake-up source based on the first angle and the second angle, including: when the first probability is greater than the second probability, determining that the first angle is the direction of the wake-up source; and when the second probability is greater than the first probability, determining that the second angle is the direction of the wake-up source.
[0029] In a possible design, the second recognition result further includes a second distance, and the second distance is a distance of the wake-up source relative to the second audio device; and the first audio device determines the direction of the wake-up source based on the first angle and the second angle, including: the first audio device predicts a third angle of the wake-up source relative to the first audio device by using the first distance, the second angle and the second distance; and the first audio device determines the direction of the wake-up source according to the third angle and the first angle, wherein the first distance is a distance between the first audio device and the second audio device.
[0030] In a possible design, the third angle satisfies:
[0031]
[0032] wherein β is the second angle; S c-s is the second distance, D is the first distance, and α' is the third angle.
[0033] In a possible design, the first audio device determines the direction of the wake-up source according to the third angle and the first angle, including:
[0034] determining an average or a weighted average of the third angle and the first angle as the direction of the wake-up source;
[0035] or
[0036] the direction of the wake-up source satisfies: α±(ω i ×Δ), ω i is a preset value, Δ=|α-α'|, α is the first angle, and α' is the third angle.
[0037] In a possible design, after the first audio device determines the group of wake-up audio devices based on the first recognition result and the second recognition result, the method further includes:
[0038] The first audio device receives a third sound signal, and the third sound signal includes a voice instruction.
[0039] The first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, and the voice instruction is used to control the first audio device.
[0040] In a possible design, before the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, the method further includes: the first audio device suppresses a sound signal in the third sound signal located in a direction other than the direction of the wake-up source; and the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, including: the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source after the suppression to obtain the voice instruction.
[0041] The third aspect further provides a wake-up recognition method. The method is applicable to a second audio device, and includes: receiving, by the second audio device, a second sound signal, the second sound signal including a sound signal emitted by a wake-up source; performing, by the second audio device, wake-up recognition on the second sound signal to obtain a second recognition result; and sending, by the second audio device, the second recognition result to a first audio device, so that the first audio device performs wake-up judgment based on a first recognition result and the second recognition result, the first recognition result being a recognition result obtained by the first audio device performing wake-up recognition on a received first sound signal.
[0042] In a possible design, before the second audio device performs wake-up recognition on the second sound signal to obtain the second recognition result, the method further includes: suppressing, by the second audio device, a sound signal in the second sound signal from the first audio device; and performing, by the second audio device, wake-up recognition on the second sound signal to obtain the second recognition result, including: performing, by the second audio device, wake-up recognition on the second sound signal after the suppression to obtain the second recognition result.
[0043] In a possible design, before the second audio device suppresses the sound signal from the first audio device in the second sound signal, the method further includes: determining, by the second audio device, that the first audio device and the wake-up source are in different directions.
[0044] In a possible design, the second identification result further includes a second angle, and the second angle is a direction of the wake-up source relative to the second audio device, and / or the second identification result further includes a second distance, and the second distance is a distance of the wake-up source relative to the second audio device.
[0045] In a possible design, before the second audio device sends the second identification result to the first audio device, the method further includes: receiving, by the second audio device, a query request from the first audio device, where the query request is used to request to query the identification result of the second audio device.
[0046] In a fourth aspect, a wake-up identification method is provided. The method is applicable to a first audio device, and the method includes: receiving, by the first audio device, a first sound signal, where the first sound signal includes a sound signal emitted by a wake-up source; performing, by the first audio device, wake-up identification on the first sound signal to obtain a first identification result; suppressing, by the first audio device, sound from a second audio device in the first sound signal; performing, by the first audio device, wake-up identification on the first sound signal after the suppression to obtain a third identification result; and determining, by the first audio device, whether to wake up an audio device group based on the first identification result and the third identification result.
[0047] In a possible design, before the first audio device suppresses the sound from the second audio device in the first sound signal, the method further includes: determining, by the first audio device, a direction of the second audio device relative to the first audio device; and suppressing, by the first audio device, the sound from the second audio device in the first sound signal includes: suppressing, by the first audio device, sound in the direction in the first sound signal.
[0048] In a possible design, the first identification result includes a first probability, the first probability is used to describe a probability that the first sound signal includes wake-up information, the third identification result includes a third probability, the third probability is used to describe a probability that the first sound signal after the suppression includes wake-up information, and the first audio device determines whether to wake up the audio device group based on the first identification result and the third identification result includes: when the first probability is greater than a first threshold value, and / or the third probability is greater than a fourth threshold value, and / or an average or weighted average of the first probability and the third probability is greater than a fifth threshold value, it is determined that the audio device group is woken up.
[0049] In a fifth aspect, a method for locating a wake-up source is provided. The method is applicable to a group of audio devices, the group of audio devices including a first audio device and a second audio device, the first audio device including a first microphone array, the second audio device including a second microphone array, and the first audio device being capable of communicating with the second audio device. The method includes: receiving, by the first audio device, a first sound signal; calculating, by the first audio device, a first angle based on the first sound signal, the first angle being a direction of the wake-up source relative to the first audio device;
[0050] receiving, by the second audio device, a second sound signal; calculating, by the second audio device, a second angle based on the second sound signal, the second angle being a direction of the wake-up source relative to the second audio device; sending, by the second audio device, the second angle to the first audio device; and determining, by the first audio device, a direction of the wake-up source based on the first angle and the second angle. The method provided in the present application is a method for determining the direction of the wake-up source by two audio devices in cooperation, and can determine a more accurate direction of the wake-up source.
[0051] In a possible design, the determining, by the first audio device, of the direction of the wake-up source based on the first angle and the second angle includes: when a first probability is greater than a second probability, determining that the first angle is the direction of the wake-up source; and when the second probability is greater than the first probability, determining that the second angle is the direction of the wake-up source; wherein the first probability is a probability that the first sound signal includes wake-up information; and the second probability is a probability that the second sound signal includes wake-up information.
[0052] In a possible design, before the determining, by the first audio device, of the direction of the wake-up source based on the first angle and the second angle, the method further includes: calculating, by the second audio device, a second distance based on the second sound signal, the second distance being a distance of the wake-up source relative to the second audio device; and the determining, by the first audio device, of the direction of the wake-up source based on the first angle and the second angle includes: predicting, by the first audio device, a third angle of the wake-up source relative to the first audio device by using a first distance, the second angle and the second distance; wherein the first distance is a distance between the first audio device and the second audio device; and the determining, by the first audio device, of the direction of the wake-up source based on the third angle and the first angle.
[0053] In a possible design, the third angle satisfies:
[0054]
[0055] wherein β is the second angle; S c-s is the second distance, D is the first distance, and α' is the third angle.
[0056] In a possible design, the first audio device determines the direction of the wake-up source according to the third angle and the first angle, including: determining an average or a weighted average of the third angle and the first angle as the direction of the wake-up source; or the direction of the wake-up source satisfies: α ± (ω i × Δ), ω i is a preset value, and Δ = |α - α'|, where α is the first angle and α' is the third angle.
[0057] In a possible design, the method further includes: the first audio device receiving a third sound signal, the third sound signal including a voice instruction; and the first audio device identifying a sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, the voice instruction being used to control the first audio device.
[0058] In a possible design, before the first audio device identifies the sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, the method further includes: the first audio device suppressing a sound signal in the third sound signal located in a direction other than the direction of the wake-up source; and the first audio device identifying the sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, including: the first audio device identifying the sound signal in the third sound signal located in the direction of the wake-up source after the suppression.
[0059] In a sixth aspect, a method for positioning a wake-up source is provided. The method is applicable to a first audio device, the first audio device including a first microphone array, a second audio device including a second microphone array, and the first audio device being capable of communicating with the second audio device. The method includes: the first audio device receiving a first sound signal; the first audio device calculating a first angle based on the first sound signal, the first angle being a direction of the wake-up source relative to the first audio device; the first audio device receiving a second angle from the second audio device, the second angle being calculated by the second audio device based on a second sound signal received by the second audio device, the second angle being a direction of the wake-up source relative to the second audio device; and the first audio device determining a direction of the wake-up source based on the first angle and the second angle.
[0060] In a possible design, the first audio device determines the direction of the wake-up source based on the first angle and the second angle, including: when a first probability is greater than a second probability, determining that the first angle is the direction of the wake-up source; when the second probability is greater than the first probability, determining that the second angle is the direction of the wake-up source; where the first probability is a probability that the first sound signal includes wake-up information, and the second probability is a probability that the second sound signal includes wake-up information.
[0061] In a possible design, before the first audio device determines the direction of the wake-up source based on the first angle and the second angle, the method further includes: the second audio device calculates a second distance according to the second sound signal, the second distance being a distance of the wake-up source relative to the second audio device; and the first audio device determines the direction of the wake-up source based on the first angle and the second angle, including: the first audio device predicts a third angle of the wake-up source relative to the first audio device by using the first distance, the second angle and the second distance; and the first audio device determines the direction of the wake-up source according to the third angle and the first angle.
[0062] In a possible design, the third angle satisfies:
[0063]
[0064] wherein β is the second angle; S c-s is the second distance, D is the first distance, and α' is the third angle.
[0065] In a possible design, the first audio device determines the direction of the wake-up source according to the third angle and the first angle, including: determining an average or a weighted average of the third angle and the first angle as the direction of the wake-up source; or the direction of the wake-up source satisfies: α ± (ω i × Δ), ω i is a preset value, Δ = |α-α'|, α is the first angle, and α' is the third angle.
[0066] In a possible design, the method further includes: the first audio device receives a third sound signal, the third sound signal including a voice instruction; and the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, the voice instruction being used to control the first audio device.
[0067] In a possible design, before the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, the method further includes: the first audio device suppresses a sound signal in the third sound signal located in a direction other than the direction of the wake-up source; and the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain the voice instruction, including: the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source after the suppression to obtain the voice instruction.
[0068] In a seventh aspect, a method for locating a wake-up source is provided. The method is applicable to a second audio device, the second audio device comprising a second microphone array, and the second audio device being capable of communicating with a first audio device. The method comprises: receiving, by the second audio device, a second sound signal; calculating, by the second audio device, a second angle based on the second sound signal, the second angle being a direction of the wake-up source relative to the second audio device; and sending, by the second audio device, the second angle to the first audio device, so that the first audio device determines a direction of the wake-up source based on the first angle and the second angle.
[0069] In a possible design, before the second audio device sends the second angle to the first audio device, the method further includes: receiving, by the second audio device, a query request from the first audio device, the query request being used to query the second angle of the second audio device.
[0070] In a possible design, the method further includes: performing, by the second audio device, wake-up recognition on the second sound signal to obtain a second recognition result, the second recognition result comprising a second probability, the second probability being used to describe a probability that the second sound signal comprises the wake-up information, and / or the second recognition result further comprising a second distance, the second distance being the direction of the second audio device relative to the second audio device.
[0071] In an eighth aspect, an audio device group is provided, comprising: a first audio device and a second audio device.
[0072] The first audio device comprises: a processor; a memory; and a first microphone array. The memory stores a computer program, and the computer program comprises instructions which, when executed by the processor, cause the first audio device to perform the steps of the method provided in the first aspect or the fifth aspect.
[0073] The second audio device comprises: a processor; a memory; and a second microphone array. The memory stores a computer program, and the computer program comprises instructions which, when executed by the processor, cause the second audio device to perform the steps of the method provided in the first aspect or the fifth aspect.
[0074] In a ninth aspect, a first audio device is provided, comprising: a processor; a memory; and a first microphone array. The memory stores a computer program, and the computer program comprises instructions which, when executed by the processor, cause the first audio device to perform the steps of the method provided in the second aspect or the fourth aspect or the sixth aspect.
[0075] In a tenth aspect, a first audio device is provided, which includes modules / units for performing the method of any possible design of the second aspect or the fourth aspect or the sixth aspect; these modules / units can be implemented by hardware, or by hardware executing corresponding software.
[0076] In an eleventh aspect, a second audio device is provided, which includes a processor, a memory, and a second microphone array; the memory stores a computer program including instructions, which, when executed by the processor, cause the second audio device to perform the method steps of the third aspect or the seventh aspect.
[0077] In a twelfth aspect, a second audio device is provided, which includes modules / units for performing the method of any possible design of the third aspect or the seventh aspect; these modules / units can be implemented by hardware, or by hardware executing corresponding software.
[0078] In a thirteenth aspect, a chip is provided, which is coupled with a memory in an electronic device, for invoking a computer program stored in the memory and performing the method of any one of the first aspect to the seventh aspect; in embodiments of the present application, "coupled with" means that two components are directly or indirectly combined with each other.
[0079] In a fourteenth aspect, a computer readable storage medium is provided, which includes a computer program, which, when running on an electronic device, causes the electronic device to perform the method of any one of the first aspect to the seventh aspect.
[0080] In a fifteenth aspect, a computer program product is provided, which includes instructions, which, when running on a computer, cause the computer to perform the method of any one of the first aspect to the seventh aspect.
[0081] The beneficial effects of the second aspect to the fifteenth aspect are the same as those of the first aspect, and are not repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0082] Figure 1 A flowchart of a sound recognition process using an algorithm is provided for embodiments of the present application;
[0083] Figure 2 A schematic diagram of a principle for an audio device to determine a direction in which another audio device is located is provided for embodiments of the present application;
[0084] Figure 3 A schematic diagram of an application scenario is provided for embodiments of the present application;
[0085] Figure 4A schematic diagram of one example of an application scenario provided by the embodiments of the present application;
[0086] Figure 5A A flowchart of a wake-up recognition method provided by the embodiments of the present application;
[0087] Figure 5B A flowchart of another wake-up recognition method provided by the embodiments of the present application;
[0088] Figure 5C A schematic diagram of another example of an application scenario provided by the embodiments of the present application;
[0089] Figure 5D A flowchart of another wake-up recognition method provided by the embodiments of the present application;
[0090] Figure 6A A flowchart of another wake-up recognition method provided by the embodiments of the present application;
[0091] Figure 6B 、 Figure 6C and Figure 6D A schematic diagram of a principle of calculating a third angle provided by the embodiments of the present application;
[0092] Figure 7 A schematic diagram of interaction between an audio device combination and a cloud provided by the embodiments of the present application;
[0093] Figure 8 A structural schematic diagram of an audio device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0094] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, the terms used in the following embodiments are only for the purpose of describing the specific embodiments of the present application, and are not intended to be limiting to the present application. As used in the specification and the appended claims of the present application, the singular forms "a," "an," and "the" are intended to include the plural forms, such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" refer to one or more than two (including two). The term "and / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships; for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects.
[0095] Reference in the specification to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places in the specification are not necessarily all referring to the same embodiment, although it can. The terms "including," "comprising," "having" and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. The terms "connected" and "coupled" are not restricted to direct connections or couplings but include indirect connections or couplings through another component or intervening components. The terms "first," "second," and "third" are used merely as labels, and are not intended to impose numerical or sequential order unless it is expressly so defined.
[0096] In the present application, the words "exemplary" and "for example" are used to illustrate aspects of the present application. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Rather, the use of the words "exemplary" or "for example" is intended to present concepts in a concrete manner.
[0097] Before the present application is described in detail, the related terms of the present application are explained.
[0098] (1) Audio device
[0099] The audio device is a device for playing a sound signal. The audio device can be a sound box, a mobile phone, a notebook computer, a television, a smart bracelet, a watch, etc. The audio device can also be a logical device, which can be understood as a logical unit / module as long as it can play a sound signal, without limiting the type and performance of the hardware device. For example, it can be one or more logical units / modules in one or more hardware devices.
[0100] The audio device can also be used for voice recognition. Generally, the voice recognition process includes two processes: wake-up recognition and voice instruction recognition. Among them, the wake-up recognition can be understood as recognizing a wake-up statement, such as "Xiao Ai Xiao Ai"; the voice instruction recognition can be understood as recognizing a voice instruction (or command), such as "play a song XXX", "switch to the next song", etc. Among them, the module responsible for wake-up recognition in the audio device is called a wake-up recognition module, and the module responsible for voice instruction recognition is called a voice instruction recognition module. The wake-up recognition module and the voice instruction recognition module are logical functional divisions, and the corresponding physical devices of the two can be the same or different.
[0101] In order to save power consumption, the voice instruction recognition module can not be in an enabled state all the time. For example, when the wake-up recognition module detects a wake-up sentence, the voice instruction recognition module is enabled, and the voice instruction recognition module performs voice instruction recognition. An example scenario is that a wake-up source (such as a user) emits a sound signal, the sound signal includes a wake-up sentence, the sound signal is collected by the wake-up recognition module in the audio device, and the wake-up recognition module detects the wake-up sentence in the sound signal, and enables the voice instruction recognition module to perform voice recognition through the voice instruction recognition module. For example, when the wake-up source (such as a user) emits a sound signal (which contains a voice instruction) again, the sound signal is collected by the voice recognition module, the voice recognition module recognizes the voice instruction in the sound signal, and performs a corresponding operation in response to the voice instruction.
[0102] (2) Algorithm
[0103] The algorithm involved in the present application can include a wake-up recognition algorithm and a voice instruction recognition algorithm. The wake-up recognition algorithm is used for wake-up recognition. For example, whether the collected sound signal includes wake-up information (such as Xiao Ai Xiao Ai) is recognized. The voice instruction recognition algorithm is used for voice instruction recognition. For example, whether the collected sound signal includes a voice instruction (such as playing a song XXX) is recognized.
[0104] For example, taking the voice instruction recognition algorithm as an example, Figure 1 (a) is a flowchart of a voice instruction recognition algorithm. As shown in Figure 1 (a), the flow of the voice instruction recognition algorithm includes steps 1 to 5.
[0105] Step 1, receiving a sound signal, for example, receiving a sound signal emitted by a wake-up source (such as a user).
[0106] Step 2, feature extraction, which can be understood as extracting the components with recognition degree in the sound signal. In order to improve accuracy, the features of each frame of sound signal can be extracted. The feature extraction can use a mel-frequency cepstral coefficient (MFCC) algorithm for extraction, and the embodiments of the present application will not be described in detail.
[0107] Step 3, obtaining phonemes based on features. The pronunciation of a word is composed of phonemes, and Chinese generally uses initial, final, etc. as a phoneme set. This process can be achieved through an acoustic model such as a hidden markov model (HMM), and the embodiments of the present application will not be described in detail.
[0108] Step 4, obtaining a word based on a phoneme; for example, matching a word in a phonetic dictionary through a phoneme.
[0109] Step 5, obtaining a sentence based on the word. Assuming that the obtained sentence is playing song XXX, the audio device responds to the voice instruction and plays the song XXX.
[0110] For example, taking the wake-up recognition algorithm as an example. Figure 1 (b) is a flowchart of a wake-up recognition algorithm. As shown in (b), the flow of the wake-up recognition algorithm includes steps 1 to 7. Among them, steps 1 to 5 are the same as steps 1 to 5 in the voice instruction recognition process, and will not be described again. Steps 6 to 7 are introduced below. Figure 1
[0111] Step 6, comparing the recognized sentence with the preset sentence to obtain a similarity. Among them, the preset sentence is a wake-up sentence set in advance.
[0112] Step 7, outputting the similarity. For example, the similarity is 80%, 60%, etc.
[0113] Alternatively, the similarity can also be converted into a probability value. For example, the similarity is 80%, and the corresponding probability is 0.8; the similarity is 20%, and the corresponding probability is 0.2. In this case, step 7 can also output the probability value.
[0114] That is, the wake-up recognition algorithm can calculate the similarity or the probability value. For the convenience of description, the similarity and the probability value are collectively referred to as "confidence".
[0115] It should be noted that Figure 1 (a) of the application illustrates a voice instruction recognition algorithm, but the application is not limited to this. Other voice instruction recognition algorithms are also possible. Similarly, Figure 1 (b) illustrates a wake-up recognition algorithm, but other wake-up recognition algorithms are also possible.
[0116] (3) Confidence
[0117] The confidence refers to the similarity or the probability value calculated by the wake-up recognition algorithm. Taking the confidence as a probability value as an example, it is used to indicate the probability that the sound signal includes wake-up information. For example, the probability value can be 0.1, 0.5, 0.9, etc.
[0118] (4) Direction
[0119] The direction involved in the application includes the direction of one audio device in the audio device combination relative to another audio device, or the direction of the wake-up source relative to the audio device, etc. The direction can be described using an angle, for example, the direction of the wake-up source relative to the audio device can mean that the wake-up source is at 30 degrees west of north, 60 degrees east of north, etc. relative to the audio device.
[0120] The following section uses microphone array positioning technology as an example to explain the principle by which the first audio device determines the direction of the second audio device relative to the first audio device.
[0121] The first audio device includes a microphone array. A microphone array can be understood as multiple microphones arranged according to a specific rule (such as three rows of three columns, five rows of five columns, etc.).
[0122] Taking the example that the sound wave emitted by the second audio device is a parallel wave, see [reference needed]. Figure 2 As shown, microphone 1 on the first audio device receives sound wave 1 at time t1, and microphone 2 receives sound wave 2 at time t2. Therefore, the time difference between sound wave 1 and sound wave 2 is t2-t1, and this time difference can be found in [reference needed]. Figure 2 As shown.
[0123] Because the sound waves emitted by the second audio device are parallel waves, the time it takes for parallel sound waves to reach the vertical plane (the plane perpendicular to the sound waves) should be the same. (See also: [link to vertical plane]). Figure 2 As shown. Therefore, the distance difference r between the parallel waves reaching microphone 1 and microphone 2 is (t2-t1)*c. Here, t1 and t2 are known quantities, the speed of sound c is a known quantity, and the distance D between microphone 1 and microphone 2 (e.g., the distance between the centroids of microphone 1 and microphone 2) is a known quantity (this distance D can be the default value stored at the factory). Therefore, as... Figure 2 As shown, the included angle θ can be determined using the known quantities t1, t2, c, and D mentioned above. For example, the included angle θ satisfies the following formula:
[0124]
[0125] The value of the included angle θ can be obtained from the formula above. The included angle θ is used to indicate the direction of the second audio device relative to the first audio device.
[0126] Figure 2 This example uses a microphone array with two microphones. It's understandable that a microphone array can have more microphones. For instance, if the microphone array is in a 3x3 grid, including nine microphones, then multiple angles can be obtained. Any one of these angles can be used as the direction of the second audio device relative to the first audio device, or the average of these angles can be used as the direction of the second audio device relative to the first audio device.
[0127] The above describes the microphone array positioning technology as an example. It can be understood that, in addition to the microphone array positioning technology described above, other positioning technologies can also be used, such as a steered-beamformer method, a high-resolution spectral analysis-based directional method, and a sound time-delay estimation (TDE)-based directional method, and the like, and the embodiments of the present application are not limited thereto.
[0128] It should be noted that the above describes the first audio device determining the direction of the second audio device relative to the first audio device as an example. It can be understood that the principle of the second audio device determining the direction of the first audio device relative to the second audio device is similar to the above principle, and is not described again. In addition, the first audio device or the second audio device can also determine the direction of the wake-up source (such as a user) based on the principle as shown in Figure 2 .
[0129] (5) Distance
[0130] The distance involved in the present application includes the distance between the first audio device and the second audio device in the combination of audio devices, or the distance between the wake-up source and the audio device.
[0131] The following mainly describes the distance between the first audio device and the second audio device as an example.
[0132] As an example, one calculation method of the distance is to determine the distance by the formula X = (2L2 / γ). Wherein, X is the distance of the second audio device relative to the first audio device, γ is the wavelength of the sound wave, and L is the length of the microphone array (a value stored in advance). Assuming that the microphone array is two microphones in Figure 2 , then L is equal to D; assuming that the number of microphones in the microphone array is more, such as a three-by-three or three-by-four matrix, then L is equal to the length of the matrix, such as the distance between the centroid of the microphone in the first row and the first column and the centroid of the microphone in the third row and the first column.
[0133] The above only exemplifies one distance calculation method, and other methods capable of calculating the distance between the first audio device and the second audio device are also possible. In addition, the first audio device can also calculate the distance between the sound source and the first audio device based on the above distance calculation method, and of course, the second audio device can also calculate the distance between the wake-up source and the second audio device based on the above distance calculation method.
[0134] (6) In the embodiments of the present application, "multiple" refers to two or more (including two). Therefore, in the embodiments of the present application, "multiple" can also be understood as "at least two". "At least one" can be understood as one or more, for example, one, two or more. For example, "including at least one" means including one or more, and does not limit which ones are included. For example, including at least one of A, B and C means that A, B, C, A and B, A and C, B and C, or A and B and C can be included. Similarly, the understanding of "at least one" and the like is also similar. "And / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. In addition, the character " / ", unless otherwise specified, generally represents an "or" relationship between the associated objects before and after it.
[0135] Unless otherwise stated, the ordinal numbers "first", "second" and the like mentioned in the embodiments of the present application are used to distinguish a plurality of objects. For example, the first audio device and the second audio device are only used to distinguish the two audio devices, and are not used to limit the order, time sequence, priority or importance of the two audio devices.
[0136] Figure 3 The application scenario provided by the embodiments of the present application is shown in the schematic diagram. It includes N audio devices, and N is a positive integer greater than or equal to 2. Figure 3 In the embodiments of the present application, N=2 is simplified as an example. Among them, the two audio devices can be respectively arranged at different geographical positions. Each of the two audio devices can play a sound signal, and the two audio devices can communicate to realize data transmission. Among them, the two audio devices can be the same type of audio device, such as two audio devices are both sound boxes; or, both are mobile phones, both are tablet computers, etc.; or, the two audio devices can include different types of audio devices, such as two audio devices are a combination of a sound box and a mobile phone; or, a combination of a mobile phone and a tablet computer, etc. Compared with a single audio device playing music, two audio devices playing music at the same time can provide a better auditory experience for users.
[0137] In the embodiments of the present application, Figure 3The application scenario shown is an example. It is assumed that the first audio device is a primary audio device, and the second audio device is a secondary audio device. Generally, the wake-up recognition process includes: after the primary audio device receives a sound signal, it identifies whether the sound signal includes a wake-up statement; if so, it wakes up the voice instruction recognition module in the primary audio device to perform voice instruction recognition. In this process, the secondary audio device does not participate and the primary audio device mainly performs wake-up recognition. When the primary audio device performs wake-up recognition, it locates the direction of the wake-up source (such as a user) relative to the primary audio device. Therefore, the voice instruction recognition process generally includes: after the primary audio device determines the direction of the wake-up source, it performs voice instruction recognition based on the direction. For example, it is assumed that the wake-up source is determined to be at angle 1, the primary audio device receives a sound signal from the wake-up source, and considers that the sound at angle 1 in the sound signal is the sound of the wake-up source, and performs voice instruction recognition on the sound at angle 1. In this process, the secondary audio device still does not participate. When the primary audio device detects a voice instruction, it executes the voice instruction and notifies the secondary audio device to execute the voice instruction (such as playing a song, etc.), and the secondary audio device participates in use. In short, only the primary audio device performs wake-up recognition, determination of the direction of the wake-up source, and voice instruction recognition, and the secondary audio device does not participate in wake-up recognition, determination of the direction of the wake-up source, and voice instruction recognition.
[0138] However, there are some possible scenarios: for example, the noise in the environment is large, the sound signal collected by the primary audio device includes a lot of noise, which reduces the accuracy of wake-up recognition, the accuracy of the determination of the direction of the wake-up source, and further reduces the accuracy of voice instruction recognition.
[0139] An example of the scenario is described below.
[0140] Figure 4 An example of an application scenario is provided for the embodiments of the present application.
[0141] Figure 4 A schematic diagram of a living room in a home is shown. Two audio devices (such as audio devices on the left and right sides of a TV cabinet, such as a first audio device and a second audio device) are arranged in the living room. It is assumed that the first audio device is a primary audio device, and the second audio device is a secondary audio device. For convenience of description, the first audio device is referred to as a primary audio device, and the second audio device is referred to as a secondary audio device. Figure 4 The scenario (or environment) shown is an example in which only two audio devices and a wake-up source (such as a user) emit sound. For example, the two audio devices are playing music, and the wake-up source emits a wake-up statement. In this case, the sound signal received by the primary audio device includes not only the sound signal emitted by the wake-up source (i.e., the user), but also the sound signal emitted by the noise source (i.e., the secondary audio device). In this case, the following problems may occur:
[0142] 1. Due to the interference of noise, the accuracy of the main audio device wake-up recognition is low. For example, the wake-up source (such as a user) utters a wake-up statement, but due to the interference of the sound of the secondary audio device, the main audio device cannot recognize the wake-up statement uttered by the wake-up source. For example, the user may utter the wake-up statement multiple times, but the audio device cannot be awakened, and the user experience is poor.
[0143] 2. Due to the interference of noise, the main audio device cannot accurately determine the direction of the wake-up source. For example, it is assumed that the user is actually located at angle A, but due to the interference of the sound signal of the secondary audio device, the main audio device recognizes that the user is at angle B, which is obviously inaccurate.
[0144] 3. Since the wake-up source is recognized at angle B, in the subsequent voice command recognition process, the main audio device performs voice command recognition on the sound signal in angle B in the received sound signal. Obviously, the sound signal corresponding to angle B is not the sound signal uttered by the user, so the accuracy of the voice command is low.
[0145] In view of this, the embodiments of the present application provide a wake-up recognition method. The method does not simply rely on the main audio device for wake-up recognition, determination of the direction of the wake-up source, and recognition of voice commands, but rather multiple audio devices cooperate to improve the accuracy of wake-up recognition and voice recognition.
[0146] In order to more clearly show the technical solutions provided by the present application, the technical solutions provided by the present application are described in the following embodiments.
[0147] Embodiment one
[0148] This embodiment one introduces multiple audio devices cooperating to perform wake-up recognition, which does not simply rely on one audio device for wake-up device, but improves the accuracy of wake-up recognition.
[0149] This embodiment one takes the application scenario shown in Figure 4 as an example for introduction, that is, there are only two audio devices and a wake-up source (such as a user) in the scene (or environment) to utter a sound signal, for example, the two audio devices are playing music, and the wake-up source utters a wake-up statement.
[0150] First mode
[0151] In the first mode, the first audio device and the second audio device can both perform wake-up recognition, and obtain respective recognition results, and the first audio device comprehensively makes a further wake-up judgment based on the two recognition results. For example, referring to Figure 5A , the flow of the first mode includes the following steps:
[0152] S501, the first audio device receives a first sound signal. The first sound signal includes a wake-up sentence from a wake-up source, and also includes a sound signal from the second audio device (the process is not shown in the figure).
[0153] S502, the second audio device receives a second sound signal. The second sound signal includes a wake-up sentence from a wake-up source, and also includes a sound signal from the first audio device (the process is not shown in the figure).
[0154] The execution order of S501 and S502 is not limited in the present application.
[0155] S503, the first audio device performs wake-up recognition on the first sound signal to obtain a first recognition result. The first recognition result includes a first confidence. The first confidence can be a first probability that the first audio device calculates by using a wake-up recognition algorithm that the wake-up sentence is included in the first sound signal. The wake-up recognition algorithm is described in the previous term explanation section and will not be repeated here.
[0156] S504, the second audio device performs wake-up recognition on the second sound signal to obtain a second recognition result. The second recognition result includes a second confidence. The second confidence can be a second probability that the second audio device calculates by using a wake-up recognition algorithm that the wake-up sentence is included in the second sound signal.
[0157] Optionally, the first audio device and / or the second audio device can start the wake-up recognition algorithm to calculate the confidence every time a sound signal is received. Alternatively, in order to save power consumption, the first audio device and / or the second audio device can monitor the receiving time of the sound signal of the wake-up source in real time. When the time interval between the receiving time is less than a preset value, it indicates that the wake-up source in the environment is continuously sending voice, which may be chatting. At this time, the wake-up recognition algorithm does not need to be started, and the monitoring state is maintained. When it is monitored that the time interval is greater than the preset value, the wake-up recognition algorithm is started.
[0158] S505, the second audio device sends the second recognition result to the first audio device.
[0159] Optionally, the second audio device can actively send the first audio device the second confidence level each time it calculates the second confidence level; or the first audio device can also send the second audio device query information for querying the second confidence level (such as before S505), and the second audio device sends the first audio device the second confidence level after receiving the query information. Alternatively, the second audio device can also pre-judge the second confidence level before sending the first audio device the second confidence level. For example, it is judged whether the second confidence level is greater than a preset threshold; if yes, the second confidence level is sent to the first audio device; otherwise, the second confidence level is not sent to the first audio device. Because when the second confidence level is very low, it means that the second audio device can determine that the received second sound signal does not include the wake-up word, so it is not necessary to send the second confidence level to the first audio device for further wake-up recognition. Of course, the second audio device can also not pre-judge the second confidence level, but directly send the first audio device the second confidence level.
[0160] S506, the first audio device judges whether to wake up based on the first recognition result and the second recognition result. If yes, S507 is executed; otherwise, it can not respond.
[0161] The first recognition result includes the first confidence level, and the second recognition result includes the second confidence level. The first audio device judges whether to wake up based on the first recognition result and the second recognition result, specifically including: judging whether to wake up based on the first confidence level and the second confidence level. For example, the first audio device determines that it needs to wake up when it judges that the following conditions are met, including at least one of the following conditions:
[0162] 1. The first confidence level is greater than a first threshold, and / or the second confidence level is greater than a second threshold. The first threshold and the second threshold can be the same or different; for example, the first confidence level is 0.95, the first threshold is 0.9, the second confidence level is 0.85, and the second threshold is 0.8, then the first audio device wakes up. The first threshold and the second threshold can be preset. Alternatively, the first threshold and the second threshold can be set or adjusted according to user needs.
[0163] 2. The average value or weighted average value of the first confidence level and the second confidence level is greater than a third threshold. The weighted average value can be represented as P = first confidence level * A + second confidence level * B, where P is the weighted average value, A and B are weights, A + B = 1, and the values of A and B can be set in advance. The third threshold can be preset. Alternatively, the third threshold can be set or adjusted according to user needs.
[0164] It can be understood that the first audio device includes the wake-up recognition module and the voice instruction recognition module, and the voice instruction recognition module does not need to be in an enabled state all the time. Therefore, the first audio device determines whether to wake up based on the first recognition result and the second recognition result, which can be understood as determining whether to wake up the voice instruction recognition module in the first audio device based on the first recognition result and the second recognition result.
[0165] S507, the first audio device sends a wake-up response.
[0166] For example, when the first audio device determines to wake up, an audio response of "I am here" can be sent to notify the user to wake up the first audio device.
[0167] Optionally, S507 is not a necessary step. That is, in an embodiment, the flow can only include S501-S506, and does not include S507.
[0168] Optionally, the first audio device can also send a wake-up instruction to the second audio device to wake up the second audio device. For example, the voice instruction recognition module in the second audio device is woken up. Of course, if there is no need for the second device to execute the voice instruction recognition in the future, the first audio device can not need to send a wake-up instruction to the second audio device; or the second audio device can not be provided with a voice instruction recognition module.
[0169] It should be noted that, Figure 5A is an example in which the first audio device determines whether to wake up based on the first recognition result and the second recognition result (S506). It can be understood that this step can also be performed by the second audio device. For example, the first audio device sends the first recognition result to the second audio device, and the second audio device determines whether to wake up based on the first recognition result and the second recognition result.
[0170] Therefore, in the first mode, the first audio device determines whether to wake up by using the second recognition result (including the second confidence) calculated by the second audio device, instead of simply relying on the first recognition result (including the first confidence) calculated by the first audio device itself, which to some extent improves the accuracy of wake-up recognition.
[0171] The second mode
[0172] In the first mode above, the two audio devices each perform wake-up recognition, and two recognition results are obtained, and then the two recognition results are combined to determine whether to wake up. In the second mode, the first audio device and / or the second audio device can suppress the sound from the other party in the received sound signal before performing wake-up recognition. For example, the first audio device suppresses the sound from the second audio device in the received first sound signal, and the second audio device suppresses the sound from the first audio device in the received second sound signal. The suppressed sound signal highlights the wake-up source sound, and wake-up recognition performed on the suppressed sound signal further improves the accuracy of wake-up recognition.
[0173] Specifically, please refer to Figure 5B for a flowchart of the second mode. The flowchart includes:
[0174] S601, the first audio device receives a third sound signal from the second audio device.
[0175] S602, the first audio device determines the direction of the second audio device relative to the first audio device according to the third sound signal.
[0176] The calculation process of the direction is described above in the glossary section, and will not be repeated here. To improve accuracy, the direction can be calculated in real time. For example, the two audio devices calculate the direction each time they start playing music, or calculate the direction once every certain period of time, or calculate the direction when the first audio device or the second audio device detects a change in position (for example, a sensor on the audio device detects a change in position).
[0177] Optionally, S602 can also be performed by the second audio device. For example, the second audio device calculates the direction of the first audio device relative to the second audio device, and sends the direction to the first audio device. The opposite direction of the direction is the direction of the second audio device relative to the first audio device.
[0178] That is, the direction of the second audio device relative to the first audio device can be calculated in advance, that is, before the wake-up source sends the wake-up statement (i.e., the first sound signal), the direction is calculated so as to be used in the following process to suppress the sound from the second audio device.
[0179] Optionally, S601 and S602 can not be performed. For example, the first audio device can also obtain the direction of the second audio device relative to the first audio device through other ways. For example, manual input by the user, etc., so S601 and S602 in the figure are represented by dashed lines.
[0180] S603, the first audio device receives the first sound signal. The first sound signal includes the wake-up sentence uttered by the wake-up source, and of course, also includes the sound from the second audio device.
[0181] S604, the second audio device receives the second sound signal. The second sound signal includes the wake-up sentence uttered by the wake-up source, and of course, also includes the sound from the first audio device.
[0182] S605, the first audio device performs wake-up recognition on the first sound signal to obtain a first recognition result, which includes a first confidence.
[0183] S606, the second audio device performs wake-up recognition on the second sound signal to obtain a second recognition result, which includes a second confidence.
[0184] S607, the second audio device sends the second recognition result to the first audio device.
[0185] S603 to S607 have the same implementation principle as S501 to S505 in Figure 5A S501 to S505, which will not be described here.
[0186] S608, the first audio device suppresses the sound in the first sound signal located in the direction.
[0187] In short, the first sound signal includes the sound signal of the wake-up source, and also includes the sound signal from the noise source (i.e., the second audio device) that is relatively strong. The sound signal after suppression includes the sound signal of the wake-up source, and may also include the sound signal from the noise source that is relatively weak, highlighting the sound of the wake-up source relative to the original sound signal (i.e., the first sound signal).
[0188] Optionally, assuming that the direction is 30 degrees west of north, the first audio device can suppress the sound in the first sound signal located in 30 degrees west of north. Alternatively, the first audio device can also determine an angle range based on 30 degrees west of north. For example, 30 degrees minus a threshold 1 as the minimum value min of the angle range, and 30 degrees plus a threshold 2 as the maximum value max of the angle range, so the angle range is the interval of (min, max). For example, the threshold 1 is 5 degrees, so the minimum value min is 30 degrees-5 degrees = 25 degrees; the threshold 2 is 10 degrees, so the maximum value max is 30 degrees+10 degrees = 40 degrees; therefore, the angle range is the interval (30 degrees west of north, 40 degrees west of north). The first audio device suppresses the sound in the first sound signal located in this angle range.
[0189] The suppression principle can be that the first sound signal is a superposition of sounds from multiple directions. For example, the first sound signal satisfies A * sound signal 1 + B * other sound signals. A is a first weight, B is a second weight, and A + B = 1. Assume that sound signal 1 is a sound signal from the direction, and the other sound signals include sound signals of the wake-up source. The suppressed sound signal satisfies C * sound signal 1 + D * other sound signals. C is a third weight, D is a fourth weight, C + D = 1, and C is less than A. That is, the weight of the sound signal from the second audio device in the suppressed sound signal is reduced, and the sound signal emitted by the wake-up source is highlighted.
[0190] Optionally, the above suppression process can be performed in real time. For example, after determining the direction, the first audio device suppresses each received sound signal. Alternatively, in order to save power consumption, the first audio device can listen to the receiving time of the sound signal of the wake-up source in real time. When the time interval between the receiving times is less than a preset value, the suppression is not required (there is a conversation in the environment), and the listening state is maintained. When the time interval is greater than the preset value, the collected sound signal is suppressed.
[0191] S609, the first audio device performs wake-up identification on the suppressed first sound signal to obtain a third identification result, and the third identification result includes a third confidence.
[0192] It should be noted that the execution order of S603 to S609 is not limited in the present application.
[0193] S610, the first audio device determines whether to wake up according to the first identification result, the second identification result, and the third identification result.
[0194] The first identification result includes a first confidence, the second identification result includes a second confidence, and the third identification result includes a third confidence. The first audio device determines whether to wake up based on the first confidence, the second confidence, and the third confidence. If yes, S611 is executed; otherwise, it can not respond. For example, the first audio device determines that it needs to wake up when the following conditions are met, and the conditions include at least one of the following:
[0195] 1. The third confidence is greater than a fourth threshold, and / or the first confidence is greater than a first threshold, and / or the second confidence is greater than a second threshold.
[0196] 2. an average or a weighted average of the first confidence level and the second confidence level is greater than a third threshold value, and / or, an average or a weighted average of the third confidence level and the first confidence level is greater than a fifth threshold value, and / or, an average or a weighted average of the third confidence level and the second confidence level is greater than a sixth threshold value, and / or, an average or a weighted average of the third confidence level, the first confidence level and the second confidence level is greater than a seventh threshold value.
[0197] The first threshold value to the seventh threshold value can be preset, or set or adjusted according to user needs.
[0198] S611, the first audio device sends a wake-up response.
[0199] Optionally, S611 is not a necessary step; that is, in an embodiment, the flow only includes S601-S610, and does not include S611.
[0200] It should be noted that, Figure 5B S605 can not be performed, so the figure uses a dashed line. If S605 is not performed, i.e., the first audio device does not need to perform wake-up identification on the first sound signal to obtain the first identification result, then in S610, it can be determined whether to wake up based only on the second identification result and the third identification result. The second identification result includes the second confidence level, and the third identification result includes the third confidence level. For example, when the third confidence level is greater than the fourth threshold value, and / or, the second confidence level is greater than the second threshold value, and / or, an average or a weighted average of the third confidence level and the second confidence level is greater than the sixth threshold value, the device is woken up.
[0201] It should be noted that, Figure 5B S606 and S607 can not be performed, so the figure uses a dashed line. If S606 and S607 are not performed, i.e., the second audio device does not need to perform confidence level calculation, the requirement for the second audio device is lower. In this case, if S605 is performed, then in S610, it can be determined whether to wake up based on the first identification result and the third identification result. The first identification result includes the first confidence level, and the third identification result includes the third confidence level. For example, when the third confidence level is greater than the fourth threshold value, and / or, the first confidence level is greater than the first threshold value, and / or, an average or a weighted average of the third confidence level and the first confidence level is greater than the fifth threshold value, the device is woken up. If S605 is not performed, then in S610, it can be determined whether to wake up based only on the third identification result. For example, when the third confidence level is greater than the fourth threshold value, the device is woken up.
[0202] Figure 5BThe embodiments of the present application are described by taking the first audio device suppressing the sound from the second audio device as an example. It can be understood that the second audio device can also suppress the sound from the first audio device, and then perform the wake-up recognition on the suppressed sound and send the recognition result to the first audio device for comprehensive judgment. The principle of the second audio device suppressing the sound from the first audio device is the same as that of the first audio device suppressing the sound from the second audio device, and thus is not described herein again.
[0203] In some embodiments, the first audio device can use the first mode or the second mode by default, or a switching button is provided on the first audio device to switch between the first mode and the second mode.
[0204] The third mode
[0205] The first mode above does not require the first audio device to suppress the first sound signal, and in the second mode, the first audio device needs to suppress the sound signal in the direction of the second audio device from the received first sound information. Considering a possible scenario that the wake-up source and the second audio device are in the same direction for the first audio device. For example, as shown in FIG. 6, for the first audio device, the second audio device and the wake-up source are in the same direction. In this scenario, if the sound in the direction of the second audio device is suppressed, the sound of the wake-up source will also be suppressed. Therefore, in order to avoid suppressing the sound of the wake-up source, the third mode can be used. Specifically, refer to FIG. 7, which is a flowchart of the third mode. Figure 5C Figure 5D The third mode includes the following steps:
[0206] S801, the first audio device receives a third sound signal from the second audio device.
[0207] S802, the first audio device determines the direction of the second audio device relative to the first audio device according to the third sound signal.
[0208] S803, the first audio device receives a first sound signal. The first sound signal includes a wake-up sentence issued by the wake-up source, and of course also includes the sound from the second audio device.
[0209] S804, the first audio device calculates the direction of the wake-up source relative to the first audio device according to the first sound signal.
[0210] Here, since the first sound signal includes not only the sound of the wake-up source but also the sound of the second audio device, the first audio device can calculate the direction of the wake-up source based on the sound of the wake-up source in the first sound signal. The calculation process can refer to the previous part of the specification.
[0211] S805, the first audio device determines whether the wake-up source and the second audio device are in the same direction.
[0212] The first audio device determines whether the wake-up source and the second audio device are in the same direction. If yes, S806 is executed. Otherwise, S807 is executed.
[0213] S806, the first audio device uses the first way to perform the wake-up recognition.
[0214] S807, the first audio device uses the first way or the second way to perform the wake-up recognition.
[0215] It should be noted that the first way does not need to suppress the sound from the direction where the second audio device is located. Therefore, when the wake-up source and the second audio device are in the same direction, the first way can be used for processing. When the wake-up source and the second audio device are not in the same direction, the second way or the first way can be used for processing.
[0216] Optionally, the first audio device can use the first way, the second way, or the third way by default, or a switching button is provided on the first audio device to switch between the three ways.
[0217] In summary, in the first embodiment, when the first audio device performs the wake-up recognition, the first audio device and the second audio device cooperate to perform the wake-up recognition, which improves the accuracy of the wake-up recognition.
[0218] In some embodiments, the first audio device can include multiple wake-up strategies. For example, the first wake-up strategy and the second wake-up strategy. The first wake-up strategy is the existing wake-up strategy, i.e., the first audio device only performs the wake-up recognition based on the information detected by itself, without referring to the information of the second audio device. The second wake-up strategy refers to the cooperation of the first audio device and the second audio device in the above-mentioned first embodiment. In addition, the first audio device can provide a wake-up strategy switching button to switch between the first wake-up strategy and the second wake-up strategy. In the first wake-up strategy, the accuracy of the first audio device to identify the wake-up source is low. For example, in a noisy environment, the user issues a wake-up instruction, but the device cannot be woken up for a long time. In the second wake-up strategy, the accuracy of the first audio device to identify the wake-up source is high.
[0219] Optionally, the first audio device and / or the second audio device may include a non-wake-up state (or sleep state), a pre-wake-up stage, and a wake-up stage. Taking the first audio device as an example, in the non-wake-up state, the voice command recognition module in the first audio device is in a turned-off state, but the sound acquisition module (such as a microphone) is in an enabled state and can acquire sound. The pre-wake-up stage can be a stage before entering the wake-up stage, which can be considered as a stage of initially determining that wake-up is needed. The wake-up stage can be a state in which the voice recognition module in the first audio device has been awakened and can perform voice command recognition.
[0220] The following describes the three-stage switching process.
[0221] Optionally, when the first audio device is in a non-wake-up state, it can enter the pre-wake-up phase when certain conditions are met. For example, with Figure 5A Taking the illustrated process as an example, in the non-wake-up state, the first audio device receives a first sound signal, recognizes the first sound signal to obtain a first recognition result, which includes a first confidence level. If the first confidence level is greater than a first threshold, the first audio device enters the pre-wake-up stage from the non-wake-up state. This can be understood as the first audio device initially determining that it needs to be woken up, but further judgment is needed based on information from the second audio device; therefore, the first audio device can first enter the pre-wake-up stage. After entering the pre-wake-up stage, the first audio device can perform some preparatory work for entering the wake-up stage. For example, it can power on the voice command recognition module to prepare for its startup. If, based on information from the second audio device, it is determined that wake-up is required, then the first audio device enters the wake-up stage from the pre-wake-up stage. For example, it can start the voice command recognition module. If, based on information from the second audio device, it is determined that wake-up is not required, then the first audio device returns to the non-wake-up stage from the pre-wake-up stage. At this time, powering on the voice command recognition module can be stopped.
[0222] Similarly, in the non-wake-up state, the second audio device receives the second sound signal and identifies it to obtain a second identification result. The second identification result includes a second confidence level. When the second confidence level is determined to be greater than a second threshold, the second audio device enters the pre-wake-up stage from the non-wake-up state. That is, the second audio device initially determines that it needs to be woken up, but further judgment is needed based on the information from the first audio device, so the second audio device first enters the pre-wake-up stage. If, after further judgment based on the information from the first audio device, it is indeed determined that wake-up is needed, the second audio device enters the wake-up stage from the pre-wake-up stage; otherwise, the second audio device exits the pre-wake-up stage.
[0223] Optionally, the duration of the pre-wake-up phase for the first or second audio device can be a first preset duration (such as a pre-set duration, which can be a user setting or a default setting before the device leaves the factory). If it is not determined whether to wake up within the first preset duration, the pre-wake-up phase is exited. The duration of the wake-up phase can be a second preset duration. If no voice command is recognized within the second preset duration, the wake-up phase is exited.
[0224] The above description uses the example of the first audio device and / or the second audio device including three stages (non-wake-up state, pre-wake-up state and wake-up state). Optionally, it may also include only two stages, such as non-wake-up state and wake-up state. This application embodiment does not limit this.
[0225] Example 2
[0226] In Embodiment 1, the first identification result includes a first confidence level, and the second identification result includes a second confidence level. Therefore, the first audio device can make a wake-up judgment based on the first and second confidence levels. In this Embodiment 2, the first identification result obtained by the first audio device may include a first angle in addition to the first confidence level. The first angle describes the direction of the wake-up source relative to the first audio device. The second identification result obtained by the second audio device may include a second angle in addition to the second confidence level. The second angle describes the direction of the wake-up source relative to the second audio device. Therefore, in this Embodiment 2, the first audio device can not only determine whether wake-up is needed based on the first confidence level in the first identification result and the second confidence level in the second identification result, but also determine the direction of the wake-up source based on the first angle in the first identification result and the second angle in the second identification result. Simply put, the two audio devices cooperate to determine the direction of the wake-up source. Compared to the prior art where the first audio device performs wake-up positioning alone without the participation of the second audio device, this method can determine a more accurate direction of the wake-up source. Once the direction of the wake-up source is determined, the voice command can be recognized based on the sound signal from the direction of the wake-up source during the speech recognition stage. This eliminates the need to recognize voice commands from sound information in all directions, which not only improves the efficiency of speech recognition but also enhances the accuracy of voice command recognition.
[0227] This second embodiment continues with... Figure 4 Taking the application scenario shown as an example, where there are only two audio devices and a wake-up source emitting sound, and assuming the two audio devices are located on a horizontal line (e.g., Figure 4 For example, a horizontal dashed line.
[0228] Please see Figure 6A The diagram shown is a flowchart of the wake-up recognition method provided in Embodiment 2. The process includes:
[0229] S901, the first audio device receives a third sound signal from the second audio device.
[0230] S902, the first audio device determines a distance between the two audio devices according to the third sound signal (for the convenience of description, the distance is referred to as the device distance), such as the distance D.
[0231] It should be noted that S901 and S902 can not be executed, for example, the first audio device can also use other ways to measure the distance, such as laser ranging, user manual input distance, etc., so S901 and S902 in the figure are represented by dashed lines.
[0232] S903, the first audio device receives a first sound signal. The first sound signal includes a wake-up statement issued by a wake-up source, and of course also includes sound from the second audio device.
[0233] S904, the second audio device receives a second sound signal. The second sound signal includes a wake-up statement issued by a wake-up source, and of course also includes sound from the first audio device.
[0234] S905, the first audio device performs wake-up recognition on the first sound signal to obtain a first recognition result. The first recognition result includes a first confidence and a first angle, the first confidence is a first probability that the first sound signal includes a wake-up statement, and the first angle describes the direction of the wake-up source relative to the first audio device. Wherein, the principle of calculating the first angle of the wake-up source relative to the first audio device according to the first sound signal is described in the previous glossary.
[0235] S906, the second audio device performs wake-up recognition on the second sound signal to obtain a second recognition result. The second recognition result includes a second confidence and a second angle, the second confidence is a second probability that the second sound signal includes a wake-up statement, and the second angle describes the direction of the wake-up source relative to the second audio device. Wherein, the principle of calculating the second angle of the wake-up source relative to the second audio device according to the second sound signal is described in the previous glossary.
[0236] S907, the second audio device sends the second recognition result to the first audio device.
[0237] S908, the first audio device determines whether to wake up according to the first confidence and the second confidence. If so, go to S909.
[0238] S909, the first audio device determines the direction of the wake-up source according to the first angle and the second angle.
[0239] The first mode is that if the first confidence is greater than the second confidence, the first angle is determined as the direction of the wake-up source; if the second confidence is greater than the first confidence, the second angle is determined as the direction of the wake-up source.
[0240] The second mode is that the second recognition result further includes a second distance, and the second distance is the distance from the wake-up source to the second audio device. Then, the third angle of the wake-up source relative to the first audio device is predicted by using the device distance D, the second angle included in the second recognition result and the second distance, and then the final angle of the wake-up source is determined according to the predicted third angle of the wake-up source relative to the first audio device and the actually detected first angle of the wake-up source relative to the first audio device (i.e., S905).
[0241] The process of predicting the third angle of the wake-up source relative to the first audio device is described below.
[0242] Referring to Figure 6B , the second angle and the second distance are known, and the device distance between the first audio device and the second audio device is also known, so a triangle can be constructed by using the distance D, the second angle and the second distance (the side-angle-side principle) (this process does not need to use the first angle). Based on the trigonometric relationship, the third angle in the triangle can be determined. The constructed triangle satisfies the following trigonometric relationship:
[0243] sinβ×Sc-s=h
[0244] D-(cosβ×Sc-s)=P
[0245] α′=arctan(h / P)
[0246] Wherein, β is the second angle; S c-s is the second distance, h is the distance from the wake-up source to the line connecting the centers of the two audio devices. D is the distance between the first audio device and the second audio device (device distance), P is the difference between D and the distance from the second audio device to the line where h is located, and α' is the third angle.
[0247] Therefore, the third angle α' satisfies:
[0248]
[0249] It should be noted that Figure 6B is an example in which the second angle β is an acute angle. Of course, the second angle β can also be a right angle or an obtuse angle.
[0250] Taking an obtuse angle as an example, please refer to Figure 6C , which satisfies the following trigonometric function:
[0251] sin(180-β)×Sc-s=h
[0252] D + cos(180 - β) × Sc - s) = P
[0253] α′=arctan(h / P)
[0254] Therefore, the third angle α′ can be obtained by transforming the above formula:
[0255]
[0256] Taking the second angle β as a right angle as an example, please refer to [link / reference]. Figure 6D As shown, the third angle α ′ satisfy:
[0257]
[0258] Therefore, the first audio device can select an appropriate calculation method based on the size of the second angle. For example, when the second angle is an acute angle, it can use... Figure 6B The calculation is performed as shown, when the second angle is an obtuse angle, using... Figure 6C The calculation is performed as shown, when the second angle is a right angle, using... Figure 6D Calculate the angle shown.
[0259] The above describes the process by which two audio devices cooperate to predict the third angle of the wake-up source relative to the first audio device.
[0260] After predicting the third angle of the wake-up source relative to the first audio device, the final angle of the wake-up source can be determined based on the predicted third angle and the actually detected first angle (S905). For example, the first audio device corrects the first angle based on the third angle, and the corrected angle is the final angle of the wake-up source. The corrected angle can be determined in at least one of the following ways:
[0261] Method 1: Take the average of the third angle α′ and the first angle α. This average value is the corrected angle.
[0262] Method 2: Determine the absolute value of the difference between the third angle α′ and the first angle α: Δ=|α-α′|; The corrected angle satisfies: α±(ω i ×Δ), ω i It can be a preset value.
[0263] In addition to the above-mentioned manner 1 and manner 2, in some embodiments, the first audio device can determine the absolute value of the difference between the third angle a' and the first angle a: Δ = |a-a'|; and determine whether the first angle a needs to be corrected according to Δ. For example, refer to the following formula:
[0264]
[0265] When Δ is small (for example, less than the first threshold θ1), that is, the absolute value of the difference between the third angle a ′ and the first angle a is small, at this time, the first angle a can not be corrected, that is, the corrected a modify = a. Of course, the included angle a can also be corrected, for example, using the above-mentioned manner 1 or manner 2.
[0266] When Δ is large (for example, greater than the second threshold θ2), that is, the absolute value of the difference between the third angle a ′ and the first angle a is large, at this time, the first angle a can not be corrected, that is, the corrected a modify = a. Of course, the included angle a can also be corrected, for example, using the above-mentioned manner 1 or manner 2.
[0267] When Δ is in the range of the first threshold θ1 to the second threshold θ2, the first angle a needs to be corrected, and the specific correction manner can be the above-mentioned manner 1 or manner 2.
[0268] Wherein, the first threshold θ1 and the second threshold θ2 can be preset values, and the values of the two are not limited by the embodiments of the present application. For example, θ1 is 2, 3, 5 degrees, etc.; θ2 is 2, 3, 5 degrees, etc. The two can be equal or not equal.
[0269] After the direction of the wake-up source is determined, in the voice instruction recognition stage, the voice instruction recognition can be performed on the sound signal in the direction of the wake-up source, without performing voice instruction recognition on all direction sound signals, improving the efficiency. Specifically, the voice instruction recognition process refers to S910 to S914 in Figure 6A .
[0270] S910, the first audio device receives a fourth sound signal. The fourth sound signal includes a voice instruction.
[0271] S911, the first audio device suppresses the sound in the fourth sound signal in the direction other than the direction of the wake-up source.
[0272] The non-wakeup source direction can be a direction other than the wakeup source direction. Assuming that the wakeup source direction is north by west 30 degrees, the fourth sound signal can be suppressed at angles other than north by west 30 degrees. Alternatively, the first audio device can determine an angle range based on the wakeup source direction (e.g., north by west 30 degrees) and suppress sound information in the fourth sound signal within the angle range.
[0273] It should be noted that S911 is an optional step. The flow of the second embodiment can contain S911 or not.
[0274] S912, the first audio device performs speech command recognition on the sound in the fourth sound signal in the direction of the wakeup source.
[0275] Assuming that the wakeup source direction is north by west 30 degrees, the first audio device can perform speech command recognition on the sound signal in the fourth sound signal at north by west 30 degrees. Alternatively, the first audio device can determine an angle range based on the wakeup source direction (e.g., north by west 30 degrees). For example, 30 degrees minus a threshold 1 is the minimum value min of the angle range, and 30 degrees plus a threshold 2 is the maximum value max of the angle range. Then the angle range is the interval of (min, max). For example, the threshold 1 is 5 degrees, so the minimum value min is 30 degrees-5 degrees=25 degrees; the threshold 2 is 10 degrees, so the maximum value max is 30 degrees+10 degrees=40 degrees, so the angle range is the interval (north by west 25 degrees, north by west 40 degrees). The first audio device performs speech command recognition on the sound within this angle range in the fourth sound signal. In this way, the first audio device does not need to perform speech command recognition on all angles, saving power consumption.
[0276] The speech command recognition process is described in the foregoing definition section and will not be repeated here.
[0277] S913, the first audio device executes the speech command.
[0278] For example, if the speech command is to switch to the next song, the first audio device switches to the next song.
[0279] S914, the first audio device sends the speech command to the second audio device to control the second audio device to execute the speech command.
[0280] The above embodiments are introduced with the first audio device and the second audio device as the execution subject. Alternatively, some steps in the above embodiments can also be executed by the cloud. Figure 5A For example, S506 can be executed by the cloud. For example, see Figure 7The first audio device sends a pre-wakeup event to the second audio device, the pre-wakeup event being the second recognition result, and the first audio device reports the first recognition result and the second recognition result to the cloud. The cloud determines whether the first audio device needs to be woken up according to the first recognition result and the second recognition result. If so, the cloud sends a wakeup instruction to the first audio device to wake up the first audio device, and then the first audio device sends the wakeup instruction to the second audio device. In this way, the computing task of the first audio device can be reduced.
[0281] Optionally, before the cloud sends the wakeup instruction to the first audio device, the cloud can also send prompt information to the mobile terminal of the user to prompt the user whether to confirm to wake up the first audio device. When a confirmation operation is detected on the mobile terminal, a confirmation instruction is sent to the cloud to confirm to wake up the first audio device. After the first audio device receives the confirmation instruction, the first audio device sends the wakeup instruction to the first audio device, which can avoid false wake-up.
[0282] In some embodiments, the first audio device can be a primary audio device, and the second audio device can be a secondary audio device; or the second audio device can be a primary audio device, and the first audio device can be a secondary audio device. The primary audio device can be determined by at least one of the following ways.
[0283] The primary audio device can be one or more audio devices that are pre-set. The remaining audio devices in the N audio devices are the secondary audio devices. For example, the primary audio device is pre-set before the device is shipped. The secondary audio device is the remaining audio device in the N audio devices.
[0284] Alternatively, the primary audio device can be specified by the user. For example, each of the N audio devices detects an input operation through a touch display screen, and the input operation is used to select whether the each audio device is a primary audio device or a secondary audio device.
[0285] Alternatively, the primary audio device can be the audio device with the strongest performance in the N audio devices. The audio device with the strongest performance can include an audio device with the strongest processor performance, such as the fastest processor operation speed.
[0286] The primary audio device and the secondary audio device can have the same or different structures. For example, the primary audio device has a display screen, and the secondary audio device does not have a display screen. Alternatively, the primary audio device and the secondary audio device can be responsible for different functions. For example, the primary audio device plays left channel sound information of an audio file, and the secondary audio device plays right channel sound information of the audio file.
[0287] The above embodiments are introduced by taking N=2 as an example. It can be understood that more audio devices can be included in the audio device combination. When N is other values, similar principles can be adopted, which are not repeated here.
[0288] Therefore, the cooperation among the N audio devices in the embodiments of the present application achieves the following beneficial effects:
[0289] 1. The two audio devices cooperate to perform wake-up recognition, improving the accuracy of wake-up recognition.
[0290] 2. The two audio devices cooperate to determine the direction of the wake-up source, improving the accuracy of determining the direction of the wake-up source.
[0291] 3. Since the direction of the wake-up source is accurate, when performing subsequent voice command recognition, the voice command recognition can be performed on the sound signal located in the direction of the wake-up source, improving the accuracy of voice command recognition.
[0292] The above embodiments provided by the present application are introduced from the perspective of the audio device as the execution subject. In order to realize each function in the method provided by the above embodiments of the present application, the terminal device can include a hardware structure and / or a software module, and the above functions are realized in the form of hardware structure, software module, or hardware structure plus software module. Whether a certain function in the above functions is executed in the form of hardware structure, software module, or hardware structure plus software module depends on the specific application of the technical solution and the design constraint conditions.
[0293] As shown in Figure 8 , some other embodiments of the present application disclose an audio device. The audio device is, for example, an electronic device such as a sound box, a mobile phone, a tablet computer, a notebook computer, a desktop computer, etc. As shown in Figure 8 , the audio device can include one or more processors 801, a microphone 802, a speaker 803, a memory 804, wherein the memory 804 includes one or more computer programs 805. The above devices can be connected through one or more communication buses 806. The microphone 802 can be a microphone array, which is used to receive a sound signal and can also be used to determine the direction, distance, etc. (see the previous glossary section). The speaker 803 is used to play a sound signal.
[0294] The one or more computer programs 805 are stored in the above memory 804 and are configured to be executed by the one or more processors 801. The one or more computer programs 805 include instructions, which can be used to execute each step in the above embodiments and corresponding embodiments. Figures 5A to 6A
[0295] Figure 8 The audio device shown can be the first audio device or the second audio device in the above. When Figure 8 The audio device shown can be used to perform the relevant steps of the first audio device in the above when the audio device is the first audio device. Figure 8 The audio device shown can be used to perform the relevant steps of the second audio device in the above when the audio device is the second audio device.
[0296] In the above embodiments, the term "when" or "after" can be interpreted as meaning "if" or "after" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "upon determining" or "if detecting (the stated condition or event)" can be interpreted as meaning "if determining" or "in response to determining" or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)" according to the context. In addition, in the above embodiments, relational terms such as first, second and the like are used to distinguish one entity from another entity, and do not limit any actual relationship and sequence between the entities.
[0297] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (for example, floppy disk, hard disk, magnetic tape), optical media (for example, DVD), or semiconductor media (for example, solid state disk (SSD)) and the like.
[0298] Those skilled in the art should be aware that, in the above one or more examples, the functions described in the embodiments of the present application can be implemented in hardware, software, firmware or any combination thereof. When implemented in software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or code on a computer readable medium. The computer readable medium includes computer storage medium and communication medium, and the communication medium includes any medium that facilitates transfer of a computer program from one place to another. The storage medium can be any available medium accessible by a general purpose or special purpose computer.
[0299] The above detailed description of the specific implementation of the present application has further detailed the purpose, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above description is only a specific implementation of the embodiments of the present application and is not used to limit the protection scope of the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the embodiments of the present application should be included in the protection scope of the embodiments of the present application. The above description of the present application specification can enable any person skilled in the art to utilize or implement the embodiments of the present application. Any modification based on the disclosed content should be considered as obvious in the art, and the basic principles described in the embodiments of the present application can be applied to other variations without deviating from the essence and scope of the present application. Therefore, the content disclosed in the embodiments of the present application is not limited to the described embodiments and designs, but can be extended to the maximum scope consistent with the principles and new features disclosed in the present application.
[0300] Although the present application is described in conjunction with specific features and embodiments thereof, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of the embodiments of the present application. Accordingly, the present specification and drawings are merely illustrative of the exemplary embodiments of the present application and are to be regarded as covering any and all modifications, variations, combinations or equivalents that fall within the scope of the present application. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the embodiments of the present application are intended to include these modifications and variations.
Claims
1. A wake-up recognition method, characterized by, The method is suitable for an audio device group, the audio device group comprising a first audio device and a second audio device, the first audio device comprising a first microphone array, the second audio device comprising a second microphone array, the method comprising: The first audio device receives a first sound signal, the first sound signal comprising a sound signal emitted by a wake-up source; The first audio device determines, in a case where the second audio device and the wake-up source are in different directions, to suppress a sound signal from the second audio device in the first sound signal; The first audio device performs wake-up identification on the suppressed first sound signal to obtain a first identification result; The second audio device receives a second sound signal, the second sound signal comprising a sound signal emitted by the wake-up source; The second audio device determines, in a case where the first audio device and the wake-up source are in different directions, to suppress a sound signal from the first audio device in the second sound signal; The second audio device performs wake-up identification on the suppressed second sound signal to obtain a second identification result; The second audio device sends the second identification result to the first audio device; The first audio device determines, based on the first identification result and the second identification result, whether to wake up the audio device group.
2. The method of claim 1, wherein, The first identification result comprises a first probability, the first probability being a probability that the first sound signal comprises wake-up information; the second identification result comprises a second probability, the second probability being a probability that the second sound signal comprises wake-up information; The first audio device determines, based on the first identification result and the second identification result, whether to wake up the audio device group, comprising: If the first probability is greater than a first threshold value, the second probability is greater than a second threshold value, and / or an average or weighted average of the first probability and the second probability is greater than a third threshold value, the audio device group is woken up. The first identification result further comprises a first angle, the first angle being used to indicate a direction of the wake-up source relative to the first audio device; the second identification result further comprises a second angle, the second angle being used to indicate a direction of the wake-up source relative to the second audio device; the method further comprises that the first audio device determines, based on the first angle and the second angle, a direction in which the wake-up source is located.
3. The method of claim 2, wherein, The first audio device determines, based on the first angle and the second angle, a direction in which the wake-up source is located, comprising:
4. The method of claim 3, wherein, When the first probability is greater than the second probability, it is determined that the first angle indicates the direction of the wake-up source relative to the first audio device; When the second probability is greater than the first probability, it is determined that the second angle indicates the direction of the wake-up source relative to the second audio device. The second identification result further comprises a second distance, the second distance being a distance of the wake-up source relative to the second audio device; the first audio device determines, based on the first angle and the second angle, a direction in which the wake-up source is located, comprising:
5. The method of claim 3, wherein, The first audio device predicts a third angle of the wake-up source relative to the first audio device according to the first distance, the second angle and the second distance, wherein the first distance is a distance between the first audio device and the second audio device; The first audio device determines a direction of the wake-up source relative to the first audio device according to the third angle and the first angle.
6. The method of claim 5, wherein, The third angle satisfies: where β is the second angle; S c-s D is the first distance, and α ′ is the third angle.
7. The method of claim 5, wherein, The first audio device determines a direction of the wake-up source according to the third angle and the first angle, including: determining an average or a weighted average of the third angle and the first angle, indicating the direction of the wake-up source relative to the first audio device; Or, An angle of the wake-up source relative to the first audio device is α ± (ω i × Δ), ω i is a preset value, and Δ = |α - α ′ |, α is the first angle, and α ′ is the third angle.
8. The method according to any one of claims 3 to 7, characterized in that, After the first audio device determines to wake up the audio device group based on the first identification result and the second identification result, the method further includes: The first audio device receives a third sound signal, and the third sound signal includes a voice instruction; The first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain a voice instruction, and the voice instruction is used to control the first audio device.
9. The method of claim 8, wherein, Before the first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain a voice instruction, the method further includes: The first audio device suppresses a sound signal in the third sound signal located in a direction of a non-wake-up source, and the direction of the non-wake-up source is a direction other than the direction of the wake-up source; The first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source to obtain a voice instruction, including: The first audio device identifies a sound signal in the third sound signal located in the direction of the wake-up source after the suppression to obtain a voice instruction.
10. A wake-up recognition method, characterized by, The method is applicable to a first audio device including a first microphone array, and the method includes: The first audio device receives a first sound signal, and the first sound signal includes a sound signal emitted by a wake-up source; The first audio device determines to suppress a sound signal in the first sound signal from a second audio device in a case where the second audio device and the wake-up source are in different directions; The first audio device performs wake-up identification on the suppressed first sound signal to obtain a first identification result; The first audio device receives a second identification result from the second audio device, and the second identification result is an identification result obtained by the second audio device in a case where the second audio device determines to suppress a sound signal in a received second sound signal from the first audio device and performs wake-up identification on the suppressed second sound signal; The first audio device determines whether to wake up the first audio device based on the first identification result and the second identification result.
11. The method of claim 10, wherein, The first recognition result comprises a first probability that the first sound signal comprises the wake-up information; and the second recognition result comprises a second probability that the second sound signal comprises the wake-up information. The first audio device determines whether to wake up the first audio device based on the first recognition result and the second recognition result, comprising: If the first probability is greater than a first threshold value, and the second probability is greater than a second threshold value; and / or, an average or weighted average of the first probability and the second probability is greater than a third threshold value, it is determined that the first audio device is woken up. Comprising:
12. A set of audio devices, characterized in that The first audio device and the second audio device; The first audio device comprises a processor, a memory, and a first microphone array; the memory stores a computer program, and when the computer program is executed by the processor, the first audio device executes the steps performed by the first audio device in the method of any one of claims 1-9; The second audio device comprises a processor, a memory, and a second microphone array; the memory stores a computer program, and when the computer program is executed by the processor, the second audio device executes the steps performed by the second audio device in the method of any one of claims 1-9. The computer readable storage medium comprises a computer program, and when the computer program runs on a computer, the computer executes the method of any one of claims 1-9, or the method of any one of claims 10-11.
13. A computer-readable storage medium, characterized in that, The computer program, when running on a computer, causes the computer to execute the method of any one of claims 1-9, or the method of any one of claims 10-11.
14. A computer program product, characterised in that,
Citation Information
Patent Citations
Voice control method of home appliance system and home appliance control system
CN107622652A
Voice recognizing method, device and facility, and storage medium
CN110010126A
Audio processing method and device
CN110660407A
Sound box control method, sound box and sound box system
CN110677801A
Voice awakening method and electronic equipment
CN111369988A