Speech recognition methods, devices, electronic equipment, media and software products

By obtaining the identification error and spectral ratio of the secondary path in the audio device, the user's voice and the interference voice can be distinguished, which solves the problem of misidentification by electronic devices in the presence of interference voices and improves the accuracy of voice recognition.

CN119479660BActive Publication Date: 2026-03-10VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

When there are interfering human voice signals around the user, electronic devices may misidentify the interfering human voices as the user's voice, resulting in poor recognition accuracy.

Method used

By acquiring the identification error and spectral ratio corresponding to the secondary path of the audio device, it is determined whether the signal collected by the microphone is the sound emitted by the user wearing the audio device. By utilizing the energy enhancement characteristic of the bone conduction signal collected by the in-ear microphone, it is determined whether the spectral ratio of the signal is greater than the threshold, thereby distinguishing between user speech and interference speech.

Benefits of technology

It effectively avoids misidentification and improves the accuracy of user voice recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479660B_ABST
    Figure CN119479660B_ABST
Patent Text Reader

Abstract

This application discloses a speech recognition method, apparatus, electronic device, medium, and program product, belonging to the field of audio technology. The method includes: acquiring an identification error corresponding to a secondary path of an audio device, the identification error being used to characterize whether the signals collected by at least two microphones in the audio device are human voice signals; if the identification error is greater than or equal to a first threshold, acquiring a first spectral ratio, the first spectral ratio being the ratio of the spectra of the signals collected by at least two microphones; if the first spectral ratio is greater than or equal to a second threshold, determining that the signals collected by at least two microphones were emitted by a user wearing the audio device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of audio technology, specifically relating to a speech recognition method, device, electronic device, medium, and program product. Background Technology

[0002] Currently, electronic devices can perform real-time voice detection. When a user's voice is detected, the device can perform the corresponding operation based on the semantic information of the voice, such as pausing music or adjusting the volume, thus enabling voice control of electronic devices.

[0003] In related technologies, electronic devices can collect sound signals through microphones, and then use a Voice Activity Detection (VAD) algorithm to determine that the sound signals collected by the microphone are human voice signals; then, based on the semantic information corresponding to the human voice signals, the operation corresponding to that semantic information is executed.

[0004] However, in the above method, when there are other users around the user, the microphone in the electronic device may pick up the interfering voice signals emitted by other users. As a result, the electronic device may identify the interfering voice signals as the original voice signals and perform corresponding operations based on the interfering voice signals, leading to misidentification by the electronic device. Thus, the accuracy of the electronic device in recognizing the user's voice is poor. Summary of the Invention

[0005] The purpose of this application is to provide a speech recognition method, device, electronic device, medium, and program product that can improve the accuracy of electronic devices in recognizing user speech in the presence of interfering human voice signals in the surrounding environment.

[0006] In a first aspect, embodiments of this application provide a speech recognition method executed by an audio device. The speech recognition method includes: acquiring an identification error corresponding to a secondary path of the audio device, the identification error being used to characterize whether the signals collected by at least two microphones in the audio device are human voice signals; if the identification error is greater than or equal to a first threshold, acquiring a first spectrum ratio, the first spectrum ratio being the ratio of the spectra of the signals collected by at least two microphones; and if the first spectrum ratio is greater than or equal to a second threshold, determining that the signals collected by at least two microphones were emitted by a user wearing the audio device.

[0007] Secondly, embodiments of this application provide a speech recognition device, comprising: an acquisition module and a recognition module. The acquisition module is configured to acquire an identification error corresponding to a secondary path of an audio device, the identification error being used to characterize whether the signals collected by at least two microphones in the audio device are human voice signals; and, if the identification error is greater than or equal to a first threshold, acquire a first spectrum ratio, the first spectrum ratio being the ratio of the spectra of the signals collected by the at least two microphones. The recognition module is configured to determine, if the first spectrum ratio acquired by the acquisition module is greater than or equal to a second threshold, that the signals collected by the at least two microphones are emitted by a user wearing the audio device.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment of the application, the identification error corresponding to the secondary path of the audio device can be obtained. The identification error is used to characterize whether the signal collected by at least two microphones in the audio device is a human voice signal. Then, if the identification error is greater than or equal to a first threshold, a first spectrum ratio is obtained. The first spectrum ratio is the ratio of the spectra of the signals collected by at least two microphones. Finally, if the first spectrum ratio is greater than or equal to a second threshold, it is determined that the signal collected by at least two microphones is emitted by the user wearing the audio device. In this solution, since the identification error corresponding to the secondary path can be used to determine whether the signal collected by at least two microphones is a human voice signal, when it is determined that the signal collected by at least two microphones is a human voice signal, the ratio of the spectra of the signals collected by at least two microphones can be used to determine whether the signal collected by at least two microphones was emitted by a user wearing an audio device. It can be understood that when a user wearing an audio device speaks, since the in-ear microphone of at least two microphones collects the sound signal emitted through bone conduction, the energy of the signal output by the in-ear microphone will be enhanced. At this time, the spectrum of the signal output by at least two microphones will change significantly, and thus the ratio of the spectra will also change. Therefore, if the audio device determines that the first spectrum ratio is greater than or equal to the second threshold, it can determine that the collected sound was emitted by a user wearing an audio device. In this way, false recognition is avoided and the accuracy of recognizing user speech is improved. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of a speech recognition system provided in an embodiment of this application;

[0014] Figure 2 This is one of the flowcharts of a speech recognition method provided in the embodiments of this application;

[0015] Figure 3 This is a second flowchart of a speech recognition method provided in an embodiment of this application;

[0016] Figure 4 This is one of the schematic diagrams of the active noise reduction system signal model provided in the embodiments of this application;

[0017] Figure 5 This is a second schematic diagram of the signal model of the active noise reduction system provided in the embodiments of this application;

[0018] Figure 6 This is the third flowchart of a speech recognition method provided in the embodiments of this application;

[0019] Figure 7 This is the fourth flowchart of a speech recognition method provided in the embodiments of this application;

[0020] Figure 8This is a schematic diagram of the microphone spectrum distribution inside and outside the ear canal when no self-talk occurs, provided in an embodiment of this application.

[0021] Figure 9 This is a schematic diagram of the microphone spectrum distribution inside and outside the ear canal during self-talk provided in an embodiment of this application;

[0022] Figure 10 This is a schematic diagram of the structure of a speech recognition device provided in an embodiment of this application;

[0023] Figure 11 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;

[0024] Figure 12 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0027] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, "at least one of a, b, and c" can mean "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0028] The following explains the technical terms involved in the speech recognition methods, devices, electronic devices, media, and program products provided in the embodiments of this application.

[0029] Own voice detection, also known as self-talk detection, is a technical solution for identifying whether a user wearing a smart audio device is speaking. It can also be considered a sub-direction of voice activity detection (VAD).

[0030] VAD: also known as voice endpoint detection or voice boundary detection; the purpose of VAD is to identify and eliminate long silence periods from the audio signal stream in order to save voice channel resources without reducing service quality.

[0031] Active Noise Cancellation (ANC) refers to the use of a noise cancellation system to generate a reverse sound wave that is equal to or opposite to the external noise, thereby neutralizing the noise and achieving a noise reduction effect.

[0032] Main path: Acoustic response from reference microphone to error microphone, where the reference microphone is the external microphone in the audio device and the error microphone is the internal microphone in the audio device.

[0033] Secondary Path: Noise cancellation in ANC is acoustic cancellation performed within the ear canal. It requires a speaker to play the inverse noise for cancellation and an error microphone to collect the residual noise. This path is called the secondary path.

[0034] Fast Fourier Transform (FFT): A general term for efficient and fast computation methods that utilize computers to calculate the Discrete Fourier Transform. The basic idea of ​​FFT is to decompose the original N-point sequence into a series of shorter sequences. By fully utilizing the symmetry and periodicity properties of the exponential factors in the Discrete Fourier Transform formula, the corresponding DFTs of these shorter sequences can be calculated and appropriately combined to eliminate redundant calculations, reduce multiplication operations, and simplify the structure.

[0035] The speech recognition method, device, electronic device, medium, and program product provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0036] The speech recognition method, device, electronic device, medium, and program products provided in the embodiments of this application can be applied to the following scenarios 1 and 2.

[0037] Scenario 1: Users control electronic devices, such as mobile phones, through audio devices, and perform voice command interactions, such as voice wake-up, pausing or playing music, adjusting volume, etc., and use a large language model for question and answer.

[0038] Scenario 2: A scenario where the active noise cancellation system of an audio device is controlled to update its parameters.

[0039] In Scenario 1 above, taking headphones as the audio device and a mobile phone as the electronic device, when the two microphones of the headphones collect sound signals, the headphones can obtain the identification error corresponding to the secondary path, and then determine whether the signal detected by the headphones is a human voice signal based on the identification error. Then, if the identification error is greater than or equal to a first threshold, the headphones can determine that the signal detected by the headphones is a human voice signal. Next, the headphones can obtain the ratio of the spectrum of the signals output by the two microphones of the headphones, and if the first spectrum ratio is greater than or equal to a second threshold, determine that the signal detected by the headphones is emitted by the user wearing the headphones. Finally, the semantic information of the signal detected by the headphones is identified to determine the user's intention, such as pausing music, and sending a control command containing pausing music to the mobile phone to control the mobile phone to perform the pausing music operation.

[0040] In scenario 2 above, with headphones as the audio device, when the two microphones of the headphones pick up sound signals, the headphones can obtain the identification error corresponding to the secondary path, and then determine whether the signal detected by the headphones is a human voice signal based on the identification error. Then, if the identification error is greater than or equal to a first threshold, the headphones can determine that the signal detected by the headphones is a human voice signal. Next, the headphones can obtain the ratio of the spectrum of the signals output by the two microphones of the headphones, and if the first spectrum ratio is greater than or equal to a second threshold, determine that the signal detected by the headphones is emitted by the user wearing the headphones. At this time, the headphones can update the ANC noise reduction parameters to pause ANC noise reduction.

[0041] In the speech recognition methods, devices, electronic devices, media, and program products provided in this application, since the identification error corresponding to the secondary path can be used to determine whether the signals collected by the two microphones are human voice signals, when it is determined that the signals collected by at least two microphones are human voice signals, the ratio of the spectra of the signals collected by at least two microphones can be used to determine whether the signals collected by at least two microphones were emitted by a user wearing an audio device. It can be understood that when a user wearing an audio device speaks, since the in-ear microphone among the at least two microphones collects the sound signal emitted through bone conduction, the energy of the signal output by the in-ear microphone will be enhanced. At this time, the spectra of the signals output by at least two microphones will change significantly, and thus the ratio of the spectra will also change. Therefore, if the audio device determines that the first spectrum ratio is greater than or equal to the second threshold, it can determine that the collected sound was emitted by a user wearing an audio device. This avoids misidentification and improves the accuracy of user speech recognition.

[0042] This application provides a speech recognition system, such as... Figure 1As shown, the speech recognition system 10 includes: a system identification module 11, a VAD detection module 12, and an internal and external spectral information analysis module 13.

[0043] The system can calculate the identification error using signals from the external ear microphone, the internal ear microphone, and the speaker playback signal. It can then determine whether the relative identification error of the system is greater than or equal to a first threshold. If it is less than the first threshold, it is determined that no self-talk has occurred, and the active noise reduction algorithm is updated in real time.

[0044] If the identification error is greater than or equal to the first threshold, the VAD is used to detect whether the signals collected by the two microphones are speech signals. If the output of the VAD detection is not a speech signal, then the signals collected by the two microphones are not speech signals, and the active noise reduction algorithm is updated in real time. If the error is greater than the set threshold, it is determined that the current signal is a speech signal with strong energy, and it is further determined whether the current speech signal is the user speaking or someone next to the user speaking loudly.

[0045] The output of VAD detection is a speech signal. The system can further analyze the ratio of the signal spectrum inside and outside the ear canal. If the ratio is greater than the set second threshold, it is determined that the current strong speech signal energy comes from the user's self-speaking, triggering the system's self-speaking state. At this time, the active noise reduction algorithm is stopped, and the system identifies the signals collected by the external and internal microphones, as well as obtains the semantic information of the signal. Based on the semantic information, the system controls the electronic device to perform the operation corresponding to the semantic information.

[0046] The speech recognition method provided in this application can be executed by a speech recognition device, which can be an electronic device or a functional module within an electronic device. The following description uses an electronic device as an example to illustrate the technical solution provided in this application.

[0047] This application provides a speech recognition method. Figure 2 A flowchart of a speech recognition method provided in an embodiment of this application is shown, which can be applied to audio devices. Figure 2 As shown, the speech recognition method provided in this application embodiment may include the following steps 201 to 203.

[0048] Step 201: The audio device obtains the identification error corresponding to the secondary path of the audio device.

[0049] In this embodiment of the application, the above-mentioned identification error is used to characterize whether the signals collected by at least two microphones in the audio device are human voice signals.

[0050] In this embodiment, the audio device can be a pair of headphones with active noise cancellation.

[0051] Optionally, in the embodiments of this application, the above-mentioned headphones can be any of the following: in-ear headphones, over-ear headphones, ear-hook headphones, or bone conduction headphones.

[0052] For example, the aforementioned in-ear headphones can be True Wireless Stereo (TWS) headphones.

[0053] In this embodiment of the application, the above-mentioned at least two microphones can be an external ear microphone and an internal ear microphone, respectively.

[0054] In one example, where the above-mentioned at least two microphones are two microphones, the two microphones can be an external ear microphone and an internal ear microphone, respectively.

[0055] In another example, when there are at least two microphones and a total of three microphones, the three microphones can be two external-ear microphones and one internal-ear microphone; or, the three microphones can be one external-ear microphone and two internal-ear microphones. The specific configuration can be determined according to requirements, and this application does not impose any limitations.

[0056] In another example, when there are at least two microphones and a total of four microphones, the four microphones can be three external ear microphones and one internal ear microphone; or, the four microphones can be two external ear microphones and two internal ear microphones. The specific configuration can be determined according to requirements, and this application does not impose any limitations.

[0057] Optionally, in the embodiments of this application, the microphone can be any of the following: an electrodynamic microphone, a condenser microphone, a piezoelectric microphone, or a semiconductor microphone, etc. The specific type can be determined according to actual usage requirements, and the embodiments of this application do not impose any restrictions.

[0058] In this embodiment, the secondary path is the sound propagation path between the player and the in-ear microphone in the audio device.

[0059] For example, the player described above can be a speaker.

[0060] For example, the aforementioned speaker can be any of the following: a super tweeter, a tweeter, a midrange speaker, a mid-bass speaker, a woofer, or a subwoofer, etc. The specific type can be determined based on actual usage requirements, and this application does not impose any limitations.

[0061] Step 202: If the identification error is greater than or equal to the first threshold, the audio device acquires the first spectrum ratio.

[0062] In this embodiment of the application, the first spectrum ratio is the ratio of the spectra of the signals collected by the at least two microphones.

[0063] In this embodiment of the application, the first spectrum ratio can effectively distinguish whether the signal is emitted by a user wearing an audio device or by other users around that user.

[0064] Optionally, in this embodiment, the first threshold can be a user-defined threshold. For example, the first threshold can be any of the following: 0.01, 0.02, 0.04, 0.08, 0.1, or 0.2, etc. The specific threshold can be determined according to actual usage requirements, and this embodiment does not impose any limitations.

[0065] In this embodiment of the application, after the audio device obtains the identification error, it can compare the identification error with a first threshold to determine whether the identification error is greater than or equal to the first threshold.

[0066] For example, if the identification error obtained by the audio device is 0.5 and the first threshold is 0.01, the audio device compares the identification error of 0.5 with the first threshold of 0.01 and can determine that the identification error is greater than the first threshold, so that the audio device can continue to obtain the first spectrum ratio.

[0067] Optionally, in this embodiment of the application, if the identification error obtained by the audio device is less than the first threshold, the audio device may continue to perform the active noise reduction function.

[0068] Step 203: If the first spectrum ratio is greater than or equal to the second threshold, the audio device determines that the signals collected by at least two microphones are emitted by the user wearing the audio device.

[0069] Optionally, in this embodiment, the second threshold can be a user-defined threshold. For example, the second threshold can be any of the following: -6, -4, -2, -1, 0, 1, 2, etc. The specific threshold can be determined according to actual usage requirements, and this embodiment does not impose any limitations.

[0070] In this embodiment of the application, after obtaining the first spectrum ratio, the audio device can compare the first spectrum ratio with the second threshold to determine whether the first spectrum ratio is greater than or equal to the first threshold.

[0071] For example, when the first spectrum ratio obtained by the audio device is 1 and the second threshold is 0, the audio device compares the first spectrum ratio 1 with the second threshold 0 and can determine that the above identification error is greater than the first threshold. At this time, the audio device can determine that the signals collected by at least two microphones are emitted by the user wearing the audio device.

[0072] Optionally, in this embodiment of the application, after determining that the signals collected by at least two microphones are emitted by the user wearing the audio device, the audio device can generate control commands based on the signals to control the electronic devices connected to the audio device to perform the operations corresponding to the control commands.

[0073] Optionally, in this embodiment of the application, the audio device can identify the signals collected by the two microphones through a first algorithm to obtain the semantic information corresponding to the signals, and then generate control commands based on the semantic information, and control the electronic device to perform the operation corresponding to the control commands based on the control commands.

[0074] Optionally, in this embodiment of the application, the audio device can be connected to the electronic device via Bluetooth; or the audio device can be connected to the electronic device via a Wireless Fidelity (WiFi) network.

[0075] Optionally, in this embodiment, the first algorithm can be any of the following: Artificial Intelligence (AI) algorithm, neural network algorithm, Automatic Speech Recognition (ASR) algorithm, etc. The specific algorithm can be determined according to actual usage requirements, and this embodiment does not impose any limitations.

[0076] For example, the audio device can input the above signal into the ASR algorithm. The ASR algorithm splits the signal into audio frames and transforms the split audio frames into multi-dimensional vector information according to human ear characteristics. Then, the multi-dimensional vector information is combined to form phonemes, and finally, the phonemes are combined into words and strung together into sentences to obtain the semantic information corresponding to the signal.

[0077] Optionally, in this embodiment of the application, after the audio device generates a control command based on semantic information, it can send the control command to the electronic device to control the electronic device to perform the operation corresponding to the control command.

[0078] For example, taking headphones as an audio device and a mobile phone as an electronic device, and taking the voice corresponding to the above signal as "pause music playback" as an example, after receiving the voice message "pause music playback", the headphones can obtain semantic information through the above ASR algorithm, that is, perform a "pause playback" operation on the "music application"; then, the headphones can directly generate a control command "pause playback of the song playing in the music application" based on the semantic information, and send the control command to the mobile phone via Bluetooth, so that the mobile phone can pause playback of the song playing in the music application after receiving the control command.

[0079] For example, taking the voice message corresponding to the above signal as "find the registration time for qualification certificate A", after receiving the voice message "find the registration time for qualification certificate A", the earphone can obtain semantic information through the above ASR algorithm, that is, perform the operation of "find the registration time for qualification certificate A". Then, the earphone can directly generate the control command "find the registration time for qualification certificate A" based on the semantic information, and send the control command to the mobile phone via Bluetooth, so that after receiving the control command, the mobile phone can find the registration time for qualification certificate A through the browser.

[0080] Optionally, in this embodiment of the application, after the audio device determines that the signals collected by at least two microphones are emitted by the user wearing the audio device, it can update the ANC noise reduction parameters to pause ANC noise reduction so that the user can hear their own voice clearly.

[0081] In the speech recognition method provided in this application embodiment, the identification error corresponding to the secondary path of the audio device can be obtained. The identification error is used to characterize whether the signal collected by at least two microphones in the audio device is a human voice signal. Then, if the identification error is greater than or equal to a first threshold, a first spectrum ratio is obtained. The first spectrum ratio is the ratio of the spectra of the signals collected by at least two microphones. Finally, if the first spectrum ratio is greater than or equal to a second threshold, it is determined that the signal collected by at least two microphones is emitted by the user wearing the audio device. In this solution, since the identification error corresponding to the secondary path can be used to determine whether the signals collected by at least two microphones are human voice signals, when it is determined that the signals collected by at least two microphones are human voice signals, the ratio of the spectra of the signals collected by at least two microphones can be used to determine whether the signals collected by at least two microphones were emitted by a user wearing an audio device. It can be understood that when a user wearing an audio device speaks, because the in-ear microphones of at least two microphones collect the sound signals emitted through bone conduction, the energy of the signal output by the in-ear microphones will be enhanced. At this time, the spectra of the signals output by at least two microphones will change significantly, and thus the ratio of the spectra will also change. Therefore, if the audio device determines that the first spectrum ratio is greater than or equal to the second threshold, it can determine that the collected sound was emitted by a user wearing an audio device. This avoids false recognition and improves the accuracy of user voice recognition.

[0082] Optionally, in the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, before step 201 above, the speech recognition method provided in this application embodiment further includes step 301 below, and step 201 above can be specifically implemented by steps 201a and 201b below.

[0083] Step 301: The audio device acquires at least two first signals and performs a fusion process on the at least two first signals to obtain a second signal.

[0084] In this embodiment of the application, the above-mentioned at least two first signals are signals collected by at least two microphones, and the at least two first signals correspond one-to-one with the at least two microphones.

[0085] In this embodiment of the application, the audio device can acquire the first signal through each of at least two microphones, thereby obtaining the aforementioned at least two first signals.

[0086] Optionally, in this embodiment of the application, one of the at least two microphones can be an external ear microphone or an internal ear microphone.

[0087] Optionally, in this embodiment, the other microphone among the at least two microphones can be an external ear microphone or an internal ear microphone.

[0088] For example, when the above-mentioned at least two microphones are two microphones, if one of the two microphones is an external microphone, the other microphone is an internal microphone; or, if one of the above-mentioned two microphones is an internal microphone, the other microphone is an external microphone.

[0089] Optionally, in this embodiment of the application, after the audio device obtains at least two signals collected by at least two microphones, it can filter each of the at least two signals to obtain at least two first signals. That is, each of the at least two first signals can be a filtered signal.

[0090] For example, after receiving at least two signals, the audio device can filter each of the at least two signals separately using a filter to obtain at least two first signals.

[0091] Optionally, in this embodiment of the application, the audio device may directly superimpose at least two first signals to obtain a second signal; or, the audio device may convolve at least two first signals to obtain a second signal.

[0092] Step 201a: The audio device inputs the second signal into the secondary path, processes the second signal through the secondary path, and outputs the third signal.

[0093] In this embodiment of the application, after the audio device inputs the second signal into the secondary path, the secondary path can perform noise reduction processing on the second signal to obtain the noise-reduced second signal, namely the aforementioned third signal.

[0094] It should be noted that the aforementioned secondary paths may include unconstrained frequency domain adaptive algorithms.

[0095] For example, after the audio device inputs the second signal into the unconstrained frequency domain adaptive algorithm, the inverse signal corresponding to the second signal can be obtained through the unconstrained frequency domain adaptive algorithm. Then, the second signal is fused through the inverse signal to obtain the aforementioned third signal.

[0096] It should be noted that the aforementioned inverse signal is a signal with the opposite phase to the second signal.

[0097] Step 201b: The audio device determines the identification error based on the spectrum of the second signal and the spectrum of the third signal.

[0098] In this embodiment of the application, the audio device can perform FFT processing on the second signal and the third signal respectively to obtain the second signal and the third signal in the frequency domain, and then obtain the spectrum of the second signal and the spectrum of the third signal respectively according to the signal energy of the second signal and the third signal in the frequency domain.

[0099] In this embodiment of the application, the audio device can determine the identification error based on the ratio between the signal energy contained in the spectrum of the second signal and the signal energy contained in the spectrum of the third signal.

[0100] In this embodiment, the audio device can determine whether a self-talk state has occurred by the identification error of the secondary path. When self-talk occurs, bone conduction sound is picked up by the in-ear microphone through the ear canal, and the identification error of the secondary path increases, which can effectively distinguish whether the audio device has collected a sound signal.

[0101] It should be noted that the above-mentioned self-talk state refers to the state in which the microphone in the audio device picks up the sound signal.

[0102] For example, the following explanation uses the principle that the identification error of the secondary path by the active noise cancellation system signal model increases when self-talk occurs, which can effectively distinguish whether the audio device has collected a sound signal.

[0103] like Figure 4 As shown, when a user wearing an audio device fails to speak, the external ear microphone picks up the first sound signal. Figure 4 Let x(n) represent the first sound signal, and input the first sound signal into the primary transmission path. Figure 4 In this context, P(Z) represents the secondary path system identification module. Figure 4 In this context, it is represented by sys_iden, and it is a feedforward ANC controlled filter. Figure 4Let Kf(Z) represent this; then, the feedforward ANC control filter can filter the first sound signal to obtain the filtered first sound signal; then, the second sound signal collected by the in-ear microphone... Figure 4 In this context, e(n) represents the inputs to the secondary path system identification module and the feedback ANC control filter, respectively. Figure 4 In the diagram, Kb(Z) represents the filtering process of the second audio signal through a feedback ANC control filter. The filtered second audio signal is then fused with the first audio signal and the second audio signal, and input to the secondary path system identification module and the secondary transmission path, respectively. Figure 4 S(Z) represents the secondary path system identification module, which is used to calculate the above identification error. Finally, the audio device can fuse the first sound signal output from the primary transmission path and the sound signal output from the secondary transmission path to obtain the fused signal.

[0104] According to the above Figure 4 As can be seen, the output of the secondary path can be represented by the following formula 1.

[0105] o(n)=e(n)-x(n)*P(n)+v(n) (1)

[0106] Where 0(n) is the secondary path output, e(n) is the sound signal collected by the in-ear microphone, P(n) is the signal output through the primary transmission path P(Z), and v(n) is the system measurement noise.

[0107] Combination Figure 4 ,like Figure 5 As shown, when a user wearing an audio device speaks to themselves, the external microphone picks up the sound signal emitted by the user. Figure 5 In this context, x(n) represents the sound signal emitted by the user's inner ear bone structure via an in-ear microphone. Figure 5 Let s(n) represent this. Then, the sound emitted by the bone structure in the inner ear will travel through the air. Figure 5 In this context, Air_tf(z) is represented and collected by the external microphone. At this point, the audio device collects a signal that is x(n) superimposed on s(n). Then, the audio device inputs the superimposed signal into the primary transmission path. Figure 5 In this context, P(Z) represents the secondary path system identification module. Figure 5 In this context, it is represented by sys_iden, and includes a feedforward ANC control filter. Figure 5In this context, Kf(Z) represents the signal. Next, the feedforward ANC control filter filters the superimposed signal, hereinafter referred to as the first audio signal, and outputs the filtered first audio signal. Then, the audio device inputs the user's audio signal, hereinafter referred to as the second audio signal, collected by the in-ear microphone, into the feedback ANC control filter. Figure 5 In the secondary path system identification module, represented by Kb(Z), a filtered second audio signal is obtained through a feedback ANC control filter. Then, the audio device can fuse the filtered first and second audio signals and input them respectively into the secondary path system identification module and the secondary transmission path. Figure 5 S(Z) represents the secondary path; the identification error corresponding to the secondary path can be obtained through the secondary path system identification module. Then, the sound signal output from the secondary transmission path, the sound signal output from the primary transmission path, and the bone conduction module in the audio device are compared. Figure 5 The output is represented by Bone_tf(Z), which is the sound signal collected by the in-ear microphone from the user's in-ear bone and then fused.

[0108] According to the above Figure 5 It is understandable that when a user speaks to themselves while wearing the headset, the signal transmission path of the entire recognition system changes. The occurrence of self-speaking introduces two acoustic propagation paths into the system: a. The self-speaking sound source is transmitted through bone conduction to the ear canal and picked up by the internal microphone; b. The self-speaking sound source is transmitted through the air to the outside of the ear canal and picked up by the external microphone of the headset. The output signal after the secondary path can then be represented by the following formula 2.

[0109] o(n)=e(n)-x(n)*P(n)+s(n)*Bone_tf(n)+v(n) (2)

[0110] Where 0(n) is the secondary path output, e(n) is the sound signal collected by the in-ear microphone, P(n) is the signal output through the primary transmission path P(Z), v(n) is the system measurement noise, s(n) is the signal output through the secondary transmission path S(Z), and Bone_tf(n) is the signal output through Bone_tf(Z).

[0111] Comparing Formula 1 and Formula 2 above, it can be seen that before and after the self-talk occurs, the occurrence of self-talk leads to a larger measurement error in the system identification. Therefore, by setting an appropriate threshold, the system identification error can be used to effectively distinguish whether the current state has experienced self-talk.

[0112] Optionally, in the embodiments of this application, step 201b above can be implemented by steps 401 to 405 below.

[0113] Step 401: The audio device performs a fast Fourier transform on the second signal to obtain the transformed second signal.

[0114] In this embodiment of the application, the audio device can convert the second signal in the time domain into the second signal in the frequency domain through FFT.

[0115] For example, an audio device can use FFT to convert the waveform of a second signal in the time domain into a sine wave, which is a description of the frequency domain, thereby obtaining a second signal in the frequency domain.

[0116] Step 402: The audio device obtains the fourth signal in the first frequency band from the transformed second signal.

[0117] Optionally, in this embodiment of the application, the first frequency band is a frequency band preset by the user.

[0118] It should be noted that the first frequency band mentioned above is the frequency band that best reflects the difference between the fourth and fifth signals.

[0119] For example, the frequency bands included in the above-mentioned assumed spectrum are [0, 1000], and the first frequency band can be [200, 300], [150, 400], or [300, 500], etc. The specific frequency band can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions.

[0120] In this embodiment of the application, the audio device can sample the fourth signal from the spectrum of the transformed second signal based on the first frequency band.

[0121] Optionally, in this embodiment of the application, after the audio device obtains the fourth signal, it can perform conjugation processing on the fourth signal to obtain the conjugated fourth signal.

[0122] In this embodiment, the audio device performs a phase reversal process on the fourth signal, namely the aforementioned conjugation process, to obtain a conjugated fourth signal. In other words, the phases of the conjugated fourth signal are opposite to those of the original fourth signal.

[0123] It should be noted that since the fourth signal after FFT transformation is in complex form, the audio device can obtain the signal energy value of the fourth signal in the first frequency band by performing conjugation processing on the fourth signal, which is the conjugated fourth signal mentioned above.

[0124] Step 403: The audio device performs a fast Fourier transform on the third signal to obtain the transformed third signal.

[0125] In this embodiment of the application, the audio device can convert the third signal in the time domain into a third signal in the frequency domain through FFT.

[0126] For example, an audio device can use FFT to convert the waveform of a third signal in the time domain into a sine wave, which is a description of the frequency domain, thereby obtaining the third signal in the frequency domain.

[0127] Step 404: The audio device obtains the fifth signal in the first frequency band from the transformed third signal.

[0128] In this embodiment of the application, the audio device can sample the aforementioned fifth signal from the spectrum of the transformed third signal based on the first frequency band.

[0129] Optionally, in this embodiment of the application, after the audio device obtains the fifth signal, it can perform conjugation processing on the fifth signal to obtain the conjugated fifth signal.

[0130] In this embodiment, the audio device performs a phase reversal process on the fifth signal, namely the aforementioned conjugation process, to obtain a conjugated fifth signal. In other words, the phases of the conjugated fifth signal are opposite to those of the original fifth signal.

[0131] It should be noted that the execution timing of steps 403 and 404 can be parallel; or, step 403 can be executed first, followed by step 404; or step 404 can be executed first, followed by step 403. The specific timing can be determined according to the usage requirements, and this application embodiment does not impose any limitations. For example, as shown... Figure 4 As shown, the audio device can first perform step 403 above, and then perform step 404 above.

[0132] Step 405: The audio device calculates the identification error based on the fourth signal, the fifth signal, and the number of frequency points contained in the first frequency band.

[0133] For example, the audio device can divide the spectral size of the fourth signal by the spectral size of the fifth signal, and then divide by the number of frequency points contained in the first frequency band to calculate the identification error.

[0134] Optionally, in this embodiment of the application, the audio device can calculate the identification error based on the fourth signal, the conjugate fourth signal, the fifth signal, the conjugate fifth signal, and the number of frequency points included in the first frequency band.

[0135] In this embodiment, the audio device can multiply the fourth signal with its conjugate to obtain a first energy value, and multiply the fifth signal with its conjugate to obtain a second energy value. Then, it can divide the first energy value by the second energy value to obtain a third energy value. Finally, it can divide the third energy value by the number of frequency points included in the first frequency band to obtain the aforementioned identification error. Specifically, this can be achieved using the following formula 3.

[0136]

[0137] Where re_err is the identification error, l1, l m The first frequency band, E(k) is the fourth signal, E * X(k) is the fourth signal after conjugation, and X(k) is the fifth signal. * (k) is the fifth signal after conjugation, and m is the number of frequency points included in the first frequency band.

[0138] In this embodiment, the audio device calculates the identification error using the fourth signal, the conjugate fourth signal, the fifth signal, the conjugate fifth signal, and the number of frequency points included in the first frequency band. It can be understood that a larger identification error indicates that the microphone has collected a sound signal. Therefore, the audio device can determine whether the microphone has collected a sound signal by the magnitude of the identification error, thereby improving the accuracy of the audio device in detecting sound signals.

[0139] Optionally, in the embodiments of this application, combined with Figure 2 ,like Figure 6 As shown, before "obtaining the first spectrum ratio" in step 202 above, the speech recognition method provided in this application embodiment further includes the following step 501, and the "obtaining the first spectrum ratio" in step 202 above can be specifically implemented through the following step 202a.

[0140] Step 501: The audio device performs signal recognition on at least two first signals collected from at least two microphones respectively to obtain at least two feature information.

[0141] In this embodiment of the application, the above-mentioned at least two feature information correspond one-to-one with at least two first signals.

[0142] In this embodiment of the application, the audio device can input at least two first signals into the VAD algorithm to obtain at least two feature information.

[0143] For example, taking at least two first signals as an example, the audio device can input the first first signal from the at least two first signals into the VAD algorithm. The VAD algorithm can extract the feature information corresponding to the first first signal from the first first signal through convolution. Then, the second first signal from the at least two first signals can be input into the VAD algorithm, and the feature information corresponding to the second first signal can be extracted from the second first signal through convolution.

[0144] Optionally, in the embodiments of this application, the above-mentioned feature information may include at least one of the following: signal energy, spectrum, cepstrum and harmonics, etc.

[0145] Step 202a: If at least two feature information matches the preset feature information, obtain the first spectrum ratio.

[0146] Optionally, in the embodiments of this application, when at least two feature information are signal energy, the audio device can further determine that the first signal is a human voice signal when the signal energy matches a preset energy threshold.

[0147] For example, the signal energy matching the preset energy threshold can be that the signal energy is the same as or greater than the preset energy threshold.

[0148] Optionally, in the embodiments of this application, when at least two feature information are spectra, and when the spectrum of the first signal matches the spectrum of a preset human voice signal, the audio device can further determine that the first signal is a human voice signal.

[0149] For example, the spectrum of the first signal and the spectrum of the preset human voice signal can be matched if the difference between the spectrum of the first signal and the spectrum of the preset human voice signal is less than 5%. In this case, the spectrum of the first signal and the spectrum of the preset human voice signal can be considered to match.

[0150] In this embodiment of the application, the audio device can further detect at least two first signals collected by at least two microphones through the VAD algorithm, so as to increase the probability that the collected signals are human voice signals, improve the probability that at least two first signals are human voice signals, and improve the accuracy of the audio device in determining human voice signals.

[0151] Optionally, in the embodiments of this application, combined with Figure 2 ,like Figure 7 As shown, step 202 above can be specifically implemented through steps 202b to 202f below.

[0152] Step 202b: If the identification error is greater than or equal to the first threshold, the audio device performs a fast Fourier transform on the sixth signal collected by the first microphone of at least two microphones to obtain the transformed sixth signal.

[0153] Optionally, in this embodiment, the first microphone can be one or more. It can be determined according to actual usage requirements, and this embodiment does not impose any limitations.

[0154] It is understandable that when there are multiple first microphones, the sixth signal can be the sum of signals collected by multiple first microphones.

[0155] In this embodiment of the application, the audio device can convert the sixth signal in the time domain into the sixth signal in the frequency domain through FFT.

[0156] For example, an audio device can use FFT to convert the waveform of the sixth signal in the time domain into a sine wave, which is a description of the frequency domain, thereby obtaining the sixth signal in the frequency domain.

[0157] Step 202c: The audio device obtains the seventh signal in the first frequency band from the transformed sixth signal.

[0158] In this embodiment of the application, the audio device can sample the aforementioned seventh signal from the spectrum of the transformed sixth signal based on the first frequency band.

[0159] Optionally, in this embodiment of the application, after the audio device obtains the seventh signal, it can perform conjugation processing on the seventh signal to obtain the conjugated seventh signal.

[0160] In this embodiment, the audio device performs a phase reversal process on the seventh signal, namely the aforementioned conjugation process, to obtain a conjugated seventh signal. In other words, the phases of the conjugated seventh signal are opposite to those of the original seventh signal.

[0161] Step 202d: The audio device performs a fast Fourier transform on the eighth signal acquired by the second microphone of at least two microphones to obtain the transformed eighth signal.

[0162] Optionally, in this embodiment, the second microphone can be one or more. The specific type can be determined based on actual usage requirements, and this embodiment does not impose any limitations.

[0163] It is understandable that when there are multiple second microphones, the eighth signal can be the sum of signals collected by multiple second microphones.

[0164] In this embodiment of the application, the audio device can convert the eighth signal in the time domain into the eighth signal in the frequency domain through FFT.

[0165] For example, an audio device can use FFT to convert the waveform of the eighth signal in the time domain into a sine wave, which is a description of the frequency domain, thereby obtaining the eighth signal in the frequency domain.

[0166] Step 202e: The audio device obtains the ninth signal in the first frequency band from the transformed eighth signal.

[0167] In this embodiment of the application, the audio device can sample the aforementioned ninth signal from the spectrum of the transformed eighth signal based on the first frequency band.

[0168] Optionally, in this embodiment of the application, after the audio device obtains the ninth signal, it can perform conjugation processing on the ninth signal to obtain the conjugated ninth signal.

[0169] In this embodiment, the audio device can perform phase reversal processing on the ninth signal, i.e., the aforementioned conjugation processing, to obtain the conjugated ninth signal. In other words, the phase of the conjugated ninth signal is opposite to that of the original ninth signal.

[0170] It should be noted that the execution timing of steps 202d and 202e can be parallel; or, step 202d can be executed first, followed by step 202e; or step 202e can be executed first, followed by step 202d. The specific timing can be determined according to the usage requirements, and this application embodiment does not impose any limitations. For example, as shown... Figure 6 As shown, the audio device can first perform step 202d above, and then perform step 202e above.

[0171] Step 202f: The audio device calculates the first spectrum ratio based on the seventh signal, the ninth signal, and the number of frequency points contained in the first frequency band.

[0172] For example, the audio device can divide the amplitude of the seventh signal by the amplitude of the ninth signal, and then divide by the number of frequency points included in the first frequency band to calculate the first spectrum ratio.

[0173] Optionally, in this embodiment of the application, the audio device can calculate the first spectrum ratio based on the seventh signal, the conjugate seventh signal, the ninth signal, the conjugate ninth signal, and the number of frequency points included in the first frequency band.

[0174] In this embodiment, the audio device can multiply the seventh signal with its conjugate to obtain a fourth energy value, and multiply the ninth signal with its conjugate to obtain a fifth energy value. Then, it can divide the fourth energy value by the fifth energy value to obtain a sixth energy value. Finally, it can perform logarithmic processing on the sixth energy value to obtain a seventh energy value, and then divide the seventh energy value by the number of frequency points included in the first frequency band to obtain the aforementioned first spectrum ratio. Specifically, this can be achieved through the following formula...

[0175] Equation 4 is implemented.

[0176]

[0177] Where spec_ratio is the first spectral ratio, FB(K) is the seventh signal, and FB... * (K) is the seventh signal after conjugation, FF(K) is the ninth signal, and FF * (K) represents the ninth signal after conjugation, m represents the number of frequency points included in the first frequency band, and i1 to i m This is the first frequency band.

[0178] Understandably, when self-talk is not occurring, the intensity of the external microphone in the first frequency band is much greater than that of the internal microphone due to the physical isolation of the audio equipment and the active noise cancellation effect. After self-talk occurs, due to the change in the signal path, the energy of the internal microphone in the first frequency band increases, and the energy ratio of the two microphones changes significantly. Figure 8 The spectral distribution of the microphones inside and outside the ear canal is shown when self-talk does not occur. Figure 9 The spectral distribution of the microphone inside and outside the ear canal is shown when self-talk occurs.

[0179] It should be noted that, Figure 8 and Figure 9 In this context, fb-mic is an external ear microphone, and ff-mic is an internal ear microphone.

[0180] In this embodiment, since the in-ear microphone collects the sound signal transmitted via bone conduction, the energy of the signal output by the in-ear microphone is enhanced. This causes a significant change in the spectrum of the signals output by the two microphones, resulting in an increase in the ratio of their spectra. Therefore, if the audio device determines that the collected sound is emitted by the user wearing the audio device when the first spectrum ratio is greater than or equal to a second threshold, it can confirm that the sound was emitted by the user. This avoids misidentification and improves the accuracy of user voice recognition.

[0181] The above-described method embodiments, or various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0182] It should be noted that the speech recognition method provided in this application embodiment can be executed by a speech recognition device. This application embodiment uses a speech recognition device executing the speech recognition method as an example to illustrate the speech recognition device provided in this application embodiment.

[0183] Figure 10 A schematic diagram of a possible structure of the speech recognition device involved in an embodiment of this application is shown. For example... Figure 10 As shown, the voice recognition device 70 may include an acquisition module 71 and a recognition module 72.

[0184] The acquisition module 71 is used to acquire the identification error corresponding to the secondary path of the audio device. This identification error is used to characterize whether the signals collected by at least two microphones in the audio device are human voice signals. If the identification error is greater than or equal to a first threshold, it acquires a first spectral ratio, which is the ratio of the spectra of the signals collected by the at least two microphones. The identification module 72 is used to determine that the signals collected by the at least two microphones are emitted by a user wearing the audio device if the first spectral ratio acquired by the acquisition module is greater than or equal to a second threshold.

[0185] In one possible implementation, the aforementioned speech recognition device further includes an input module and an output module. The acquisition module 71 is further configured to acquire at least two first signals before acquiring the identification error corresponding to the secondary path of the audio device, and to fuse the at least two first signals to obtain a second signal. The at least two first signals are signals collected by the at least two microphones, and each of the at least two first signals corresponds one-to-one with the at least two microphones. The input module is configured to input the second signal acquired by the acquisition module into the secondary path. The output module is configured to process the second signal through the secondary path and output a third signal. Specifically, the acquisition module 71 is configured to determine the identification error based on the spectrum of the second signal and the spectrum of the third signal.

[0186] In one possible implementation, the acquisition module 71 is specifically used to perform a fast Fourier transform on the second signal to obtain the transformed second signal; to obtain a fourth signal in the first frequency band from the transformed second signal; to perform a fast Fourier transform on the third signal to obtain the transformed third signal; to obtain a fifth signal in the first frequency band from the transformed third signal; and to calculate the identification error based on the fourth signal, the fifth signal, and the number of frequency points contained in the first frequency band.

[0187] In one possible implementation, the acquisition module 71 is specifically used to perform a fast Fourier transform on the sixth signal acquired by the first microphone among at least two microphones to obtain the transformed sixth signal; to acquire the seventh signal in the first frequency band from the transformed sixth signal; to perform a fast Fourier transform on the eighth signal acquired by the second microphone among at least two microphones to obtain the transformed eighth signal; to acquire the ninth signal in the first frequency band from the transformed eighth signal; and to calculate the first spectral ratio based on the seventh signal, the ninth signal, and the number of frequency points contained in the first frequency band.

[0188] In one possible implementation, the identification module 72 is further configured to perform signal identification on at least two first signals collected by at least two microphones before the acquisition module 71 acquires the first spectrum ratio, thereby obtaining at least two feature information pieces, wherein the at least two feature information pieces correspond one-to-one with the at least two first signals. Specifically, the acquisition module 71 is configured to acquire the first spectrum ratio when both of the at least two feature information pieces match preset feature information.

[0189] This application provides a speech recognition device. Since it can determine whether the signals collected by at least two microphones are human voice signals based on the recognition error corresponding to the secondary path, when it is determined that the signals collected by at least two microphones are human voice signals, it can determine whether the signals collected by at least two microphones were emitted by a user wearing an audio device by comparing the ratio of the spectra of the signals collected by at least two microphones. It can be understood that when a user wearing a speech recognition device speaks, the energy of the signal output by the in-ear microphone is enhanced because the in-ear microphone among the at least two microphones collects the sound signal transmitted through bone conduction. At this time, the spectra of the signals output by at least two microphones will change significantly, and thus the ratio of the spectra will also change. Therefore, if the speech recognition device determines that the first spectrum ratio is greater than or equal to a second threshold, it can determine that the collected sound was emitted by a user wearing an audio device. This avoids misidentification and improves the accuracy of user speech recognition.

[0190] The voice recognition device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific type of device.

[0191] The voice recognition device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0192] The speech recognition device provided in this application embodiment can implement the various processes implemented in the above embodiments. To avoid repetition, it will not be described again here.

[0193] Optionally, such as Figure 11 As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described speech recognition method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0194] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0195] Figure 12 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0196] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0197] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 12 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0198] The processor 110 is configured to acquire the identification error corresponding to the secondary path of the audio device, the identification error being used to characterize whether the signals collected by at least two microphones in the audio device are human voice signals; if the identification error is greater than or equal to a first threshold, acquire a first spectrum ratio, the first spectrum ratio being the ratio of the spectra of the signals collected by at least two microphones; if the first spectrum ratio is greater than or equal to a second threshold, determine that the signals collected by at least two microphones are emitted by a user wearing the audio device.

[0199] Optionally, in this embodiment of the application, the processor 110 is further configured to acquire at least two first signals before acquiring the identification error corresponding to the secondary path of the audio device, and to perform fusion processing on the at least two first signals to obtain a second signal. The at least two first signals are signals collected by at least two microphones, and the at least two first signals correspond one-to-one with the at least two microphones. Specifically, the processor 110 is configured to input the second signal into the secondary path, process the second signal through the secondary path to output a third signal, and determine the identification error based on the spectrum of the second signal and the spectrum of the third signal.

[0200] Optionally, in this embodiment of the application, the processor 110 is specifically configured to perform a fast Fourier transform on the second signal to obtain a transformed second signal; obtain a fourth signal in the first frequency band from the transformed second signal; perform a fast Fourier transform on the third signal to obtain a transformed third signal; obtain a fifth signal in the first frequency band from the transformed third signal; and calculate the identification error based on the fourth signal, the fifth signal, and the number of frequency points included in the first frequency band.

[0201] Optionally, in this embodiment of the application, the processor 110 is specifically configured to perform a fast Fourier transform on the sixth signal acquired by the first microphone among at least two microphones to obtain the transformed sixth signal; obtain a seventh signal in the first frequency band from the transformed sixth signal; perform a fast Fourier transform on the eighth signal acquired by the second microphone among at least two microphones to obtain the transformed eighth signal; obtain a ninth signal in the first frequency band from the transformed eighth signal; and calculate a first spectral ratio based on the seventh signal, the ninth signal, and the number of frequency points contained in the first frequency band.

[0202] Optionally, in this embodiment of the application, the processor 110 is further configured to perform signal recognition on at least two first signals collected by at least two microphones respectively before obtaining the first spectrum ratio, to obtain at least two feature information, wherein the at least two feature information corresponds one-to-one with the at least two first signals. Specifically, the processor 110 is configured to obtain the first spectrum ratio when at least two feature information matches preset feature information.

[0203] This application provides an electronic device that can determine whether the signals collected by at least two microphones are human voice signals based on the identification error corresponding to the secondary path. When it is determined that the signals collected by the at least two microphones are human voice signals, the ratio of the spectra of the signals collected by the at least two microphones can be used to determine whether the signals were emitted by a user wearing an audio device. It is understood that when a user wearing an audio device speaks, the energy of the signal output by the in-ear microphone is enhanced because it collects the sound signal transmitted through bone conduction. This causes a significant change in the spectra of the signals output by the at least two microphones, and consequently, a change in the ratio of their spectra. Therefore, if the audio device determines that the first spectrum ratio is greater than or equal to a second threshold, it can determine that the collected sound was emitted by a user wearing an audio device. This avoids misidentification and improves the accuracy of user voice recognition.

[0204] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0205] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.

[0206] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0207] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0208] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0209] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0210] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0211] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0212] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0213] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described speech recognition method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0214] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0215] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0216] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A voice recognition method, characterized by, The method is executed by an audio device, and comprises: obtaining a recognition error corresponding to a secondary path of the audio device, the recognition error being used to represent whether signals collected by at least two microphones in the audio device are human voice signals; in a case where the recognition error is greater than or equal to a first threshold, obtaining a first spectral ratio, the first spectral ratio being a ratio of spectrums of the signals collected by the at least two microphones; in a case where the first spectral ratio is greater than or equal to a second threshold, determining that the signals collected by the at least two microphones are uttered by a user wearing the audio device, the at least two microphones comprising at least one extra-ear microphone and at least one intra-ear microphone; before the obtaining of the recognition error corresponding to the secondary path of the audio device, the method further comprises: obtaining at least two first signals and performing fusion processing on the at least two first signals to obtain a second signal, the at least two first signals being signals collected by the at least two microphones, and the at least two first signals corresponding to the at least two microphones in a one-to-one manner; the obtaining of the recognition error corresponding to the secondary path of the audio device comprises: inputting the second signal into the secondary path, processing the second signal through the secondary path, and outputting a third signal; determining the recognition error based on a spectrum of the second signal and a spectrum of the third signal.

2. The method of claim 1, wherein, the determining of the recognition error based on the spectrum of the second signal and the spectrum of the third signal comprises: performing fast Fourier transform on the second signal to obtain a transformed second signal; obtaining a fourth signal in a first frequency band from the transformed second signal; performing fast Fourier transform on the third signal to obtain a transformed third signal; obtaining a fifth signal in the first frequency band from the transformed third signal; calculating the recognition error based on the fourth signal, the fifth signal, and a number of frequency points contained in the first frequency band.

3. The method of claim 1, wherein, the obtaining of the first spectral ratio comprises: performing fast Fourier transform on a sixth signal collected by a first microphone in the at least two microphones to obtain a transformed sixth signal; obtaining a seventh signal in a first frequency band from the transformed sixth signal; performing fast Fourier transform on an eighth signal collected by a second microphone in the at least two microphones to obtain a transformed eighth signal; obtaining a ninth signal in the first frequency band from the transformed eighth signal; calculating the first spectral ratio based on the seventh signal, the ninth signal, and a number of frequency points contained in the first frequency band.

4. The method according to claim 1 or 3, characterized in that, before the obtaining of the first spectral ratio, the method further comprises: performing signal recognition on at least two first signals respectively collected by the at least two microphones to obtain at least two feature information, the at least two feature information corresponding to the at least two first signals in a one-to-one manner; the obtaining of the first spectral ratio comprises: in a case where the at least two feature information all match preset feature information, obtaining the first spectral ratio.

5. A speech recognition apparatus characterized by comprising: The speech recognition device comprises an obtaining module and a recognition module. The acquisition module is configured to acquire a recognition error corresponding to a secondary path of an audio device, the recognition error being used to represent whether signals collected by at least two microphones in the audio device are human voice signals; and acquire a first spectral ratio when the recognition error is greater than or equal to a first threshold, the first spectral ratio being a ratio of spectrums of the signals collected by the at least two microphones. The identification module is configured to determine that the signals collected by the at least two microphones are issued by a user wearing the audio device when the first spectral ratio acquired by the acquisition module is greater than or equal to a second threshold, the at least two microphones including at least one extra-ear microphone and at least one intra-ear microphone. The speech recognition apparatus further includes an input module and an output module. The acquisition module is further configured to acquire at least two first signals, and perform fusion processing on the at least two first signals to obtain a second signal, the at least two first signals being signals collected by the at least two microphones, and the at least two first signals corresponding to the at least two microphones in a one-to-one manner. The input module is configured to input the second signal acquired by the acquisition module to the secondary path. The output module is configured to output a third signal by processing the second signal through the secondary path. The acquisition module is specifically configured to determine the recognition error based on a spectrum of the second signal and a spectrum of the third signal.

6. The apparatus of claim 5, wherein, The acquisition module is specifically configured to: perform fast Fourier transform on a sixth signal collected by a first microphone in the at least two microphones to obtain a transformed sixth signal; acquire a seventh signal in a first frequency band from the transformed sixth signal; perform fast Fourier transform on an eighth signal collected by a second microphone in the at least two microphones to obtain a transformed eighth signal; acquire a ninth signal in the first frequency band from the transformed eighth signal; and calculate the first spectral ratio based on the seventh signal, the ninth signal, and a number of frequency points contained in the first frequency band.

7. An electronic device, comprising: A processor, a memory, and a program or instructions stored on the memory and executable on the processor, the program or instructions being executed by the processor to implement the steps of the speech recognition method in any one of claims 1 to 4.

8. A readable storage medium, characterized by, A program or instructions are stored on the readable storage medium, the program or instructions being executed by the processor to implement the steps of the speech recognition method in any one of claims 1 to 4.

9. A computer program product stored in a storage medium, the computer program product being executed by at least one processor to implement the speech recognition method in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Sequenced adaptation of anti-noise generator response and secondary path response in an adaptive noise canceling system

    CN104272379A

  • Method and device for detecting voice of earphone wearer and storage medium

    CN111933140A