Mode switching method and apparatus for TWS earphone

Through the joint positioning of error microphone, reference microphone, call microphone and speaker, the voice activation detection value and sound source direction value are calculated, and the TWS headphone mode is automatically switched, which solves the problems of poor user experience and high cost of traditional solutions, and realizes the functions of intelligent removal-free and automatic switching.

WO2025091700A1PCT designated stage expired Publication Date: 2025-05-08ZHUHAI JIELI TECH
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072780
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-01
Filing Date
2024-01-17
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Traditional TWS headsets have poor user experience when they need to switch from active noise reduction mode to transmissive mode, and traditional solutions require additional sensors, which increases costs.

Method used

Through the joint positioning of error microphone, reference microphone, call microphone and speaker, the voice activation detection value and sound source direction value are calculated, and the headphone mode is automatically switched to achieve intelligent removal-free switching.

Benefits of technology

No additional sensors are required, which improves user experience, reduces costs, and realizes the intelligent removal-free and automatic switching mode function of TWS headsets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072780_08052025_PF_FP_ABST
    Figure CN2024072780_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a mode switching method for a TWS earphone. The method comprises: calculating a cancellation power spectrum on the basis of a first audio signal and an active-noise-cancellation audio signal, calculating a voice activation detection value and a power value in a human voice bandwidth on the basis of the cancellation power spectrum, and calculating a first probability value of speaking, wherein the first audio signal is a signal collected by an error microphone, and the active-noise-cancellation audio signal is a signal which is output to a loudspeaker for playing; when the first probability value is greater than or equal to a first preset threshold value, calculating power spectra of first and second audio signals, and a correlation coefficient and a first sound source direction value of the first and second audio signals, and calculating a second probability value of speaking on the basis of the power spectra, the correlation coefficient and the first sound source direction value, wherein the second audio signal is a signal collected by a reference microphone; when the second probability value is greater than or equal to a second preset threshold value, calculating a second sound source direction value on the basis of the second audio signal and a third audio signal, and calculating a third probability value of speaking on the basis of the second sound source direction value, wherein the third audio signal is a signal collected by a call microphone; and when the third probability value is greater than or equal to a third preset threshold value, switching an earphone to a non-active noise cancellation mode.
Need to check novelty before this filing date? Find Prior Art

Description

A TWS earphone mode switching method and device Technical Field

[0001] The present invention relates to the field of earphone technology, and in particular to a method and device for switching modes of a TWS earphone. Background Art

[0002] In recent years, with the rapid adoption of new-generation consumer electronics devices, wireless headphones have seen a significant growth trend. Active noise-canceling headphones, as a new research direction, are attracting increasing attention from technical professionals. Wireless headphones with active noise cancellation offer a comfortable experience in noisy environments like buses, subways, and airplanes, and are gaining increasing recognition from the market and consumers.

[0003] Active noise-canceling headphones play a signal of equal amplitude and opposite phase to the noise through the speaker, canceling out the noise inside the ear and creating a noise-canceling zone in the ear canal, effectively reducing external noise. When you need to receive external voice signals, such as when communicating with others, you need to switch the headphones from active noise-canceling mode (ANC mode) to transparent mode (transparency mode).

[0004] Traditional technical solutions either require manually switching to transparent mode to effectively receive external information while wearing headphones; or use additional sensors (such as bone conduction sensors) to recognize the wearer's speech and switch to transparent mode.

[0005] Traditional technical solutions have the following disadvantages:

[0006] 1. Frequent mode switching is required, and the user experience is not good enough;

[0007] 2. The solution of using additional sensors provides an improved experience compared to manual switching, but the additional hardware sensors increase the cost.

[0008] Summary of the Invention

[0009] Based on the above situation, the main purpose of the present invention is to provide a TWS earphone mode switching method and device to realize the intelligent free-to-remove and automatic mode switching functions of TWS active noise reduction earphones.

[0010] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0011] A TWS headset mode switching method is used to switch the headset from an active noise reduction mode to a non-active noise reduction mode, the headset including an error microphone, a reference microphone, a call microphone, and a speaker, and the headset is in the active noise reduction mode. The method includes:

[0012] S100: Calculating a cancellation power spectrum based on a first audio signal and an active noise reduction audio signal, calculating a voice activation detection value and a power value within a human voice bandwidth based on the cancellation power spectrum, and calculating a first probability value of the wearer speaking based on the voice activation detection value and the power value within the human voice bandwidth, wherein the first audio signal is a signal collected by the error microphone, and the active noise reduction audio signal is a signal output to the speaker for playback;

[0013] S200: When the first probability value is greater than or equal to a first preset threshold, calculating the power spectrum of the first audio signal and the power spectrum of the second audio signal, calculating the correlation coefficient between the first audio signal and the second audio signal, and calculating a first sound source direction value based on the first audio signal and the second audio signal, and then calculating a second probability value that the wearer is speaking based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient, and the first sound source direction value, where the second audio signal is a signal collected by the reference microphone;

[0014] S300: When the second probability value is greater than or equal to a second preset threshold, calculating a second sound source direction value based on the second audio signal and the third audio signal, and calculating a third probability value that the wearer is speaking based on the second sound source direction value, where the third audio signal is a signal collected by the call microphone;

[0015] S400: When the third probability value is greater than or equal to a third preset threshold, the headset switches from the active noise reduction mode to the non-active noise reduction mode.

[0016] Preferably, calculating the cancellation power spectrum according to the first audio signal and the active noise reduction audio signal in S100 includes:

[0017] Calculating a power spectrum Xf1(m,k) of each frame of the first audio signal and a power spectrum Xf4(m,k) of each frame of the active noise reduction audio signal;

[0018] By calibration or adaptive means, at each frequency point, the power value of the first audio signal is offset by the power value of the active noise reduction audio signal to obtain the offset power spectrum Xf(m,k): Xf(m,k)=Xf1(m,k)-a(k)Xf4(m,k);

[0019] Where m is the number of frames, k is the frequency, and a(k) is the adaptive or calibrated weight.

[0020] Preferably, calculating the power value within the human voice bandwidth according to the cancellation power spectrum in S100 includes:

[0021] An integration operation is performed on the cancellation power spectrum within the preset frequency range to obtain a power value XF(m) within the human voice bandwidth.

[0022] Preferably, the preset frequency range is 100 Hz to 1 kHz.

[0023] Preferably, calculating the voice activity detection value according to the cancellation power spectrum includes:

[0024] The voice activation detection value is calculated based on the cancellation power spectrum, energy characteristics, endpoint detection or deep learning method.

[0025] Preferably, the first probability value of the wearer speaking is calculated in S100 according to the power value within the human voice bandwidth and the voice activation detection value: result1 (m) = β 11 *p xf (m)+β 12 *p vad (m)

[0026] Among them, β 11 , β 12 Weighted value, β 11 +β 12 =1, p xf (m) is the power probability value within the human voice bandwidth, p vad (m) is the probability value of the voice activation detection value, p vad (m) is equal to the voice activity detection value, and m is the number of frames.

[0027] Preferably, the p xf (m) is:

[0028] When the XF(m) is less than or equal to the preset power reference value, the p xf (m) is the XF(m) divided by the first preset power reference value, otherwise the p xf (m) is 1, and the preset power reference value is a power reference value set according to the sensitivity of the error microphone and the ADC gain.

[0029] Preferably, the correlation coefficient between the first audio signal and the second audio signal calculated in S200 is:

[0030] Wherein, f1[n] is the sampling value of the first audio signal, f2[n+k] is the sampling value of the second audio signal, n is the sampling point, and k is the number of sampling points of the second audio signal delayed relative to the first audio signal.

[0031] Preferably, calculating the first sound source direction value according to the first audio signal and the second audio signal in S200 includes:

[0032] S201: Calculate the frequency response between each frame of the first audio signal and each frame of the second audio signal:

[0033] Wherein, H1(m,k) includes the sound source phase information, m is the frame number, k is the frequency point, Xf1(m,k) is the power spectrum of the first audio signal, and Xf2(m,k) is the power spectrum of the second audio signal;

[0034] S202: Calculate the phase information of each frame using the inverse tangent function according to the frequency response: phase(m,k)=∠H1(m,k)

[0035] S203: Smoothing the phase information: phase(m,k)=(1-α)phase(m-1,k)+α*phase(m,k)

[0036] Among them, α is the smoothing factor;

[0037] S204: Calculate the delay of frequency point k based on the smoothed phase information: t1 = phase(k) / (2*pi*k / L*Fs)

[0038] t1 is the direction value of the first sound source, L is the number of sampling points, and Fs is the sampling rate.

[0039] Preferably, calculating the second probability value of the wearer speaking according to the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value in S200 includes: result (m) = β 21 *p pwr (m)+β 22 *p rec (m)+β 23 *p pha1 (m)

[0040] Among them, β 21 , β 22 , β 23 Weighted value, β 21 +β 22 +β 23 =1, p pwr (m) is a weighted value of the power probability value of the first audio signal and the power probability value of the second audio signal;

[0041] The p rec(m) is obtained by normalizing the correlation coefficient between the first audio signal and the second audio signal;

[0042] The p pha1 (m) is based on the first sound source direction value t1, the preset first delay threshold T ref1 and time parameter δ T1 calculated;

[0043] m is the number of frames.

[0044] Preferably, the p pwr (m) is:

[0045] Calculating a power spectrum Xf1(m,k) of each frame of the first audio signal and a power spectrum Xf2(m,k) of each frame of the second audio signal;

[0046] Performing integration operations on Xf1(m,k) and Xf2(m,k) within a preset frequency range to obtain a first audio signal power XF1(m) and a second audio signal power XF2(m);

[0047] When the XF1(m) is less than or equal to the first preset power reference value, the p xf1 (m) is the XF1(m) divided by the first preset power reference value, otherwise the p xf1 (m) is 1, and the first preset power reference value is a power reference value set according to the error microphone sensitivity and the ADC gain;

[0048] When the XF2(m) is less than or equal to the second preset power reference value, the p xf2 (m) is the XF2(m) divided by the second preset power reference value, otherwise the p xf2 (m) is 1, and the second preset power reference value is a power reference value set according to the reference microphone sensitivity and ADC gain;

[0049] The p pwr For the p xf1 (m) and p xf2 The weighted value of (m).

[0050] Preferably, calculating the second sound source direction value according to the second audio signal and the third audio signal in S300 includes:

[0051] S301: Calculate the frequency response between each frame of the second audio signal and each frame of the third audio signal:

[0052] Wherein, H2(m,k) includes the sound source phase information, m is the frame number, k is the frequency point, Xf2(m,k) is the power spectrum of the second audio signal, and Xf3(m,k) is the power spectrum of the third audio signal;

[0053] S302: Calculate the phase information of each frame using the inverse tangent function according to the frequency response: phase(m,k)=∠H2(m,k)

[0054] S303: Smoothing the phase information: phase(m,k)=(1-α)phase(m-1,k)+α*phase(m,k)

[0055] α is the smoothing factor;

[0056] S304: Calculate the delay of frequency point k based on the smoothed phase information: t2 = phase(k) / (2*pi*k / L*Fs)

[0057] L is the number of sampling points, Fs is the sampling rate, and t2 is the second sound source direction value.

[0058] Preferably, calculating the third probability value of the wearer speaking according to the second sound source direction value in S300 includes: result3 (m) = p pha2 (m)

[0059] The p pha2 (m) is based on the second sound source direction value t2, the preset first delay threshold T ref2 and time parameter δ T2 Calculated.

[0060] Preferably, the method further comprises:

[0061] S500: When the headset is in the non-active noise reduction mode, within the preset time, the first probability value is calculated according to the S100. If the first probability value is less than the first preset threshold, the headset switches to the active noise reduction mode. Otherwise, the second probability value is calculated according to the S200. If the second probability value is less than the second preset threshold, the headset switches to the active noise reduction mode. Otherwise, the third probability value is calculated according to the S300. If the third probability value is less than the third preset threshold, the headset switches to the active noise reduction mode. Otherwise, the headset remains in the non-active noise reduction mode.

[0062] The present invention also discloses a TWS headset mode switching device for switching the headset from an active noise reduction mode to a non-active noise reduction mode. The headset includes an error microphone, a reference microphone, a call microphone, and a speaker. The headset is in the active noise reduction mode. The device includes a first probability value calculation module, a second probability value calculation module, a third probability value calculation module, and a mode switching module:

[0063] The first probability value calculation module is configured to calculate a cancellation power spectrum based on the first audio signal and the active noise reduction audio signal, calculate a voice activation detection value and a power value within a human voice bandwidth based on the cancellation power spectrum, and calculate a first probability value of the wearer speaking based on the voice activation detection value and the power value within the human voice bandwidth, wherein the first audio signal is a signal collected by the error microphone, and the active noise reduction audio signal is a signal output to the speaker for playback;

[0064] The second probability value calculation module is configured to calculate, when the first probability value is greater than or equal to a first preset threshold, a power spectrum of the first audio signal and a power spectrum of the second audio signal, calculate a correlation coefficient between the first audio signal and the second audio signal, and calculate a first sound source direction value based on the first audio signal and the second audio signal, and then calculate a second probability value that the wearer is speaking based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient, and the first sound source direction value, where the second audio signal is a signal collected by the reference microphone;

[0065] The third probability value calculation module is configured to calculate a second sound source direction value based on the second audio signal and the third audio signal when the second probability value is greater than or equal to a second preset threshold, and calculate a third probability value that the wearer is speaking based on the second sound source direction value, wherein the third audio signal is a signal collected by the call microphone;

[0066] The mode switching module is configured to output a mode switching signal to switch the headset from the active noise reduction mode to the non-active noise reduction mode when the third probability value is greater than or equal to a third preset threshold.

[0067] Preferably, the first probability value calculation module includes a cancellation power spectrum calculation unit, a power value calculation unit within the human voice bandwidth, a voice activity detection value calculation unit and a first probability value calculation unit.

[0068] The cancellation power spectrum calculation unit is configured to calculate a power spectrum of each frame of the first audio signal and a power spectrum of each frame of the active noise reduction audio signal, and to cancel the power value of the active noise reduction audio signal by the power value of the first audio signal at each frequency point through calibration or adaptive means to obtain the cancellation power spectrum;

[0069] The power value calculation unit within the human voice bandwidth is used to perform an integration operation on the cancellation power spectrum within a preset frequency range to obtain the power value within the human voice bandwidth;

[0070] The voice activity detection value calculation unit is configured to calculate the voice activity detection value according to the cancellation power spectrum;

[0071] The first probability value calculation unit is used to calculate a first probability value of the wearer speaking according to the power value within the human voice bandwidth and the voice activation detection value.

[0072] Preferably, the second probability value calculation module includes a power spectrum calculation unit, a correlation coefficient calculation unit, a first sound source direction value calculation unit and a second probability value calculation unit.

[0073] The power spectrum calculation unit is used to calculate the power spectrum of the first audio signal and the power spectrum of the second audio signal;

[0074] The correlation coefficient calculation unit is used to calculate the correlation coefficient between the first audio signal and the second audio signal;

[0075] The first sound source direction value calculation unit is used to calculate a first sound source direction value according to the power spectrum of the first audio signal and the power spectrum of the second audio signal;

[0076] The second probability value calculation unit is used to calculate a second probability value that the wearer is speaking based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value.

[0077] Preferably, the first sound source direction value calculation unit includes a first frequency response calculation unit, a first phase calculation unit and a first delay calculation unit.

[0078] The first frequency response calculation unit is used to calculate the frequency response of each frame of the first audio signal and each frame of the second audio signal;

[0079] The first phase calculation unit is used to calculate the phase information of each frame according to the frequency response through an inverse tangent function, and smooth the phase information:

[0080] The first delay calculation unit is used to calculate the delay of frequency point k according to the smoothed phase information.

[0081] Preferably, the third probability value calculation module includes a second sound source direction value calculation unit and a third probability value calculation unit.

[0082] The second sound source direction value calculation unit is used to calculate a second sound source direction value according to the power spectrum of the second audio signal and the power spectrum of the third audio signal;

[0083] The third probability value calculation unit is used to calculate a third probability value of the wearer speaking according to the second sound source direction value.

[0084] Preferably, the second sound source direction value calculation unit includes a second frequency response calculation unit, a second phase calculation unit and a second delay calculation unit.

[0085] The second frequency response calculation unit is used to calculate the frequency response of each frame of the second audio signal and each frame of the third audio signal;

[0086] The second phase calculation unit is used to calculate the phase information of each frame according to the frequency response through an inverse tangent function, and smooth the phase information;

[0087] The second delay calculation unit is used to calculate the delay of frequency point k according to the smoothed phase information.

[0088] Preferably, the mode switching module is further used to calculate the first probability value through the first probability value calculation module within a preset time when the headset is in the non-active noise reduction mode. If the first probability value is less than the first preset threshold, the headset switches to the active noise reduction mode; otherwise, calculate the second probability value through the second probability value calculation module. If the second probability value is less than the second preset threshold, the headset switches to the active noise reduction mode; otherwise, calculate the third probability value through the third probability value calculation module. If the third probability value is less than the third preset threshold, the headset switches to the active noise reduction mode; otherwise, the headset remains in the non-active noise reduction mode.

[0089] The present invention also discloses an earphone mode switching chip capable of executing any TWS earphone mode switching method described in any one of the present inventions.

[0090] The present invention also discloses a TWS headset, comprising the TWS headset mode switching device described in any one of the present invention, or the headset mode switching chip of the present invention.

[0091] The present invention also discloses a storage medium, which stores a program, wherein the program is used to be executed to implement the TWS earphone mode switching method described in any one of the present invention.

[0092] The TWS headphone mode switching method of the present invention calculates three probability values ​​based on power, correlation coefficient, and sound source direction, and combines them to determine whether the current mode is switched from active noise reduction mode to non-active noise reduction mode. The three probabilities are recursively related. Only when the first probability is satisfied can the second probability be calculated. Through software analysis, the intelligent mode switching function of TWS active noise reduction headphones without removal is realized, improving the user experience. Compared with traditional solutions, this solution does not require the use of additional sensors such as bone conduction. The error microphone, reference microphone, call microphone, and speaker are used to jointly locate the sound source, ensuring the accuracy of identifying the wearer's speech and reducing costs.

[0093] Other beneficial effects of the present invention will be explained through the introduction of specific technical features and technical solutions in the specific implementation methods. Those skilled in the art should be able to understand the beneficial technical effects brought about by the introduction of these technical features and technical solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The following describes the preferred embodiments of the TWS headset mode switching method and device of the present invention with reference to the accompanying drawings.

[0095] FIG1 is a flow chart of a method for switching modes of a TWS headset according to a preferred embodiment of the present invention;

[0096] FIG2 is a flow chart of calculating a first sound source direction value according to a preferred embodiment of the present invention;

[0097] FIG3 is a flow chart of calculating a second sound source direction value according to a preferred embodiment of the present invention;

[0098] FIG4 is a flow chart of a method for switching modes of a TWS headset according to another preferred embodiment of the present invention;

[0099] FIG5 is a block diagram of a TWS headset mode switching device according to a preferred embodiment of the present invention;

[0100] FIG6 is a block diagram of a first probability value calculation module according to a preferred embodiment of the present invention;

[0101] FIG7 is a block diagram of a second probability value calculation module according to a preferred embodiment of the present invention;

[0102] FIG8 is a block diagram of a first sound source direction value calculation unit according to a preferred embodiment of the present invention;

[0103] FIG9 is a block diagram of a third probability value calculation module according to a preferred embodiment of the present invention;

[0104] FIG10 is a block diagram of a second sound source direction value calculation unit according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0105] The present invention is described below based on the following embodiments, but the present invention is not limited to these embodiments. In the following detailed description of the present invention, some specific details are described in detail. In order to avoid obscuring the essence of the present invention, well-known methods, processes, procedures, and components are not described in detail.

[0106] Furthermore, persons of ordinary skill in the art will appreciate that the figures provided herein are for illustration purposes only and are not necessarily drawn to scale.

[0107] Unless the context clearly requires otherwise, throughout the specification and claims, the words "include," "comprising," and similar words should be construed in an inclusive sense rather than an exclusive or exhaustive sense; that is, in the sense of "including but not limited to."

[0108] In the description of the present invention, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance. In addition, in the description of the present invention, unless otherwise specified, "plurality" means two or more.

[0109] Figure 1 is a flow chart of a TWS headset mode switching method according to a preferred embodiment of the present invention, which is used to switch the headset from active noise reduction mode to non-active noise reduction mode. The headset of the present invention includes an error microphone, a reference microphone, a call microphone and a speaker. Usually, the error microphone is also called a feedback microphone, and the reference microphone is also called a feedforward microphone. The error microphone, the reference microphone and the call microphone all collect audio information of the environment, such as the wearer's speech, the music sound of the speaker, the speech of an outsider and other current ambient sounds, but because the positions of the microphones are different, the collected signals are also different. As we all know, active noise reduction mode is ANC mode. Non-active noise reduction mode means that the headset does not perform active noise reduction function, which can include normal mode, transparent mode, etc. In order to better hear the ambient sound, it is generally switched to the transparent mode of the headset, also known as the transparent transmission mode.

[0110] The headset is in active noise reduction mode, and the TWS headset mode switching method includes:

[0111] S100: Calculating a cancellation power spectrum based on a first audio signal and the active noise reduction audio signal, calculating a voice activation detection value and a power value within a human voice bandwidth based on the cancellation power spectrum, and calculating a first probability value of the wearer speaking based on the voice activation detection value and the power value within the human voice bandwidth, wherein the first audio signal is a signal collected by the error microphone, and the active noise reduction audio signal is a signal output to the speaker for playback.

[0112] Typically, the error microphone receives ambient sound, the sound played by the speaker, and the sound of people speaking. The signal played by the speaker may cause false triggering. In order to accurately identify the sound of people speaking, the sound played by the speaker can be offset. Since the sound played by the speaker is known, in one embodiment, the power spectrum Xf1(m,k) of each frame of the first audio signal and the power spectrum Xf4(m,k) of each frame of the active noise reduction audio signal can be calculated;

[0113] By calibration or adaptive means, at each frequency point, the power value of the first audio signal is offset by the power value of the active noise reduction audio signal to obtain the offset power spectrum Xf(m,k): Xf(m,k)=Xf1(m,k)-a(k)Xf4(m,k);

[0114] Where m is the number of frames, k is the frequency, and a(k) is the adaptive or calibrated weight.

[0115] In a specific embodiment, in order to accurately calculate the power spectrum of the first audio signal, it is necessary to perform processing such as framing, windowing, and power spectrum calculation on the first audio signal. For example, the sampling rate of the error microphone can be set to 16000 Hz or 8000 Hz, and the frame length is 16 to 32 ms, which corresponds to 256 / 512 sampling points at a sampling rate of 16000 Hz, and the overlap is selected to be 50%. The Hanning window can be used for windowing, and the windowed signal is subjected to a fast Fourier transform to obtain the power spectrum as follows:

[0116] m is the number of frames, k is the frequency, L is the number of sampling points, and x(n) is the sampling signal of the first audio signal.

[0117] Similarly, the power spectrum Xf4(m,k) of the active noise reduction audio signal is calculated.

[0118] In one embodiment, the cancellation power spectrum within a preset frequency range may be integrated to obtain the power value XF(m) within the human voice bandwidth. In a specific embodiment, the preset frequency range is 100 Hz to 1 kHz.

[0119] In one embodiment, the voice activation detection value may be calculated based on energy features, endpoint detection, or deep learning methods.

[0120] In one embodiment, the following formula may be used to calculate the first probability value of the wearer speaking according to the power value within the human voice bandwidth and the voice activation detection value: result1 (m) = β 11 *p xf (m)+β 12 *p vad (m)

[0121] Among them, β 11 , β 12 Weighted value, β 11 +β 12 =1, p xf (m) is the power probability value within the human voice bandwidth, which indicates the probability that the wearer is speaking at the power value XF(m) within the human voice bandwidth, p vad (m) is the probability value of the voice activation detection value, p vad (m) can be equal to the voice activity detection value, where m is the number of frames.

[0122] In one embodiment, when XF(m) is less than or equal to a preset power reference value, the p xf (m) is the XF(m) divided by the preset power reference value, otherwise the p xf (m) is 1, and the preset power reference value is a power reference value set according to the sensitivity of the error microphone and the ADC gain.

[0123] S200: When the first probability value is greater than or equal to a first preset threshold, the power spectrum of the first audio signal and the power spectrum of the second audio signal are calculated, and the correlation coefficient between the first audio signal and the second audio signal is calculated, and a first sound source direction value is calculated based on the first audio signal and the second audio signal, and then a second probability value that the wearer is speaking is calculated based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value, where the second audio signal is a signal collected by the reference microphone.

[0124] In a specific embodiment, voice data of different people speaking while wearing the headset is collected, and the first preset threshold is determined based on this voice data and the sensitivity. This sensitivity can also be determined based on demand. The second and third preset thresholds in subsequent steps S300 and S400 are also determined in this way.

[0125] When the wearer speaks, the first audio signal and the second audio signal collected by the error microphone and the reference microphone have a significant correlation. Therefore, the correlation coefficient can be used to detect the probability of the wearer speaking. In one embodiment, the correlation coefficient of the first audio signal and the second audio signal can be calculated using the following formula:

[0126] Wherein, f1[n] is the sampling value of the first audio signal, f2[n+k] is the sampling value of the second audio signal, the sampling value is generally the signal amplitude value, n is the sampling point, and k is the number of sampling points of the second audio signal delayed relative to the first audio signal.

[0127] The sound source direction value can be calculated by the time difference between the signal reaching the error microphone and the reference microphone. Generally, the time delay estimation method can use calculation methods such as generalized cross-correlation and acoustic transfer function ratio. In one embodiment, as shown in Figure 2, the acoustic transfer function ratio can be used to calculate the time delay, as follows:

[0128] S201: Calculate the frequency response of each frame of the first audio signal and each frame of the second audio signal:

[0129] Where H1(m,k) is a frequency response that includes the phase information of the sound source, m is the number of frames, k is the frequency, Xf1(m,k) is the power spectrum of the first audio signal, and Xf2(m,k) is the power spectrum of the second audio signal.

[0130] S202: Calculate the phase information of each frame using the inverse tangent function according to the frequency response: phase(m,k)=∠H1(m,k)

[0131] S203: Smoothing the phase information: phase(m,k)=(1-α)phase(m-1,k)+α*phase(m,k)

[0132] Among them, α is the smoothing factor;

[0133] S204: Calculate the delay of frequency point k based on the smoothed phase information: t1 = phase(k) / (2*pi*k / L*Fs)

[0134] t1 is the direction value of the first sound source, L is the number of sampling points, and Fs is the sampling rate.

[0135] It should be noted that the sound source, sound source direction or sound source positioning is generally expressed by phase. In the present invention, the phase information is converted into delay t1 according to the need of calculating probability.

[0136] In one embodiment, the second probability value of the wearer speaking can be calculated based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient, and the first sound source direction value using the following formula: result (m) = β 21 *p pwr (m)+β 22 *p rec (m)+β 23 *p pha1 (m)

[0137] Among them, β 21 , β 22 , β 23 Weighted value, β 21 +β22 +β 23 =1, p pwr (m) is a weighted value of the power probability value of the first audio signal and the power probability value of the second audio signal, indicating the probability that the wearer is speaking under the power value XF1(m) of the first audio signal and the power value XF2(m) of the second audio signal;

[0138] The p rec (m) is obtained by normalizing the correlation coefficient between the first audio signal and the second audio signal;

[0139] The p pha1 (m) is based on the first sound source direction value t1, the preset first delay threshold T ref1 and time parameter δ T1 calculated;

[0140] m is the number of frames.

[0141] In one embodiment, the p pwr (m) is:

[0142] Calculating a power spectrum Xf1(m,k) of each frame of the first audio signal and a power spectrum Xf2(m,k) of each frame of the second audio signal;

[0143] Performing integration operations on Xf1(m,k) and Xf2(m,k) within a preset frequency range to obtain a first audio signal power XF1(m) and a second audio signal power XF2(m);

[0144] When the XF1(m) is less than or equal to the first preset power reference value, the p xf1 (m) is the XF1(m) divided by the first preset power reference value, otherwise the p xf1 (m) is 1, and the first preset power reference value is a power reference value set according to the error microphone sensitivity and the ADC gain;

[0145] When the XF2(m) is less than or equal to the second preset power reference value, the p xf2 (m) is the XF2(m) divided by the second preset power reference value, otherwise the p xf2 (m) is 1, and the second preset power reference value is a power reference value set according to the reference microphone sensitivity and ADC gain;

[0146] The p pwr For the p xf1 (m) and p xf2 The weighted value of (m).

[0147] S300: When the second probability value is greater than or equal to the second preset threshold, a second sound source direction value is calculated based on the second audio signal and the third audio signal, and a third probability value that the wearer is speaking is calculated based on the second sound source direction value, and the third audio signal is a signal collected by the call microphone.

[0148] The sound source direction value can also be calculated by the time difference between the signal reaching the reference microphone and the call microphone. Generally, the time delay estimation method can use calculation methods such as generalized cross-correlation and acoustic transfer function ratio. In one embodiment, as shown in Figure 3, the acoustic transfer function ratio can be used to calculate the time delay, as follows:

[0149] S301: Calculate the frequency response between each frame of the second audio signal and each frame of the third audio signal:

[0150] Where H2(m,k) is a frequency response that includes the phase information of the sound source, m is the number of frames, k is the frequency, Xf2(m,k) is the power spectrum of the second audio signal, and Xf3(m,k) is the power spectrum of the third audio signal.

[0151] S302: Calculate the phase information of each frame using the inverse tangent function according to the frequency response: phase(m,k)=∠H2(m,k)

[0152] S303: Smoothing the phase information: phase(m,k)=(1-α)phase(m-1,k)+α*phase(m,k)

[0153] α is the smoothing factor;

[0154] S304: Calculate the delay of frequency point k based on the smoothed phase information: t2 = phase(k) / (2*pi*k / L*Fs)

[0155] L is the number of sampling points, Fs is the sampling rate, and t2 is the second sound source direction value.

[0156] Similarly, the sound source, the direction of the sound source or the location of the sound source are generally expressed by phase. In the present invention, the phase information is converted into the delay t2 according to the need of calculating the probability.

[0157] In one embodiment, the third probability value of the wearer speaking can be calculated based on the second sound source direction value using the following formula: result3 (m) = p pha2 (m)

[0158] The p pha2 (m) is based on the second sound source direction value t2, the preset first delay threshold T ref2and time parameter δ T2 Calculated.

[0159] S400: When the third probability value is greater than or equal to a third preset threshold, the headset switches from the active noise reduction mode to the non-active noise reduction mode.

[0160] In one embodiment, as shown in FIG4 , the TWS headset mode switching method of the present invention further includes:

[0161] S500: When the headset is in the non-active noise reduction mode, within the preset time, the first probability value is calculated according to the S100. If the first probability value is less than the first preset threshold, the headset switches to the active noise reduction mode. Otherwise, the second probability value is calculated according to the S200. If the second probability value is less than the second preset threshold, the headset switches to the active noise reduction mode. Otherwise, the third probability value is calculated according to the S300. If the third probability value is less than the third preset threshold, the headset switches to the active noise reduction mode. Otherwise, the headset remains in the non-active noise reduction mode.

[0162] In a specific embodiment, the preset time can be set as needed, such as 5s, 10s, etc. If the wearer is detected speaking within the preset time, the transparent state is saved unchanged, and the time from recognition to the moment the wearer speaks is reset.

[0163] The TWS headphone mode switching method of the present invention calculates three probability values ​​based on power, correlation coefficient, and sound source direction, and combines them to determine whether the current mode is switched from active noise reduction mode to non-active noise reduction mode. The three probabilities are recursively related, and the second probability can only be calculated when the first probability is satisfied. Through software analysis, the intelligent, hands-free switching function of the TWS active noise reduction headphones is realized, improving the user experience. Compared with traditional solutions, this solution does not require the use of additional sensors such as bone conduction. Instead, it uses an error microphone, a reference microphone, a call microphone, and a speaker to jointly locate the sound source, ensuring the accuracy of identifying the wearer's speech and reducing costs.

[0164] The present invention also discloses a TWS earphone mode switching device for switching the earphone from an active noise reduction mode to a non-active noise reduction mode. The earphone includes an error microphone, a reference microphone, a call microphone and a speaker. The earphone is in the active noise reduction mode. As shown in Figure 5, the device includes a first probability value calculation module 1, a second probability value calculation module 2, a third probability value calculation module 3 and a mode switching module 4.

[0165] The first probability value calculation module 1 is used to calculate the cancellation power spectrum based on the first audio signal and the active noise reduction audio signal, calculate the voice activation detection value and the power value within the human voice bandwidth based on the cancellation power spectrum, and calculate the first probability value of the wearer speaking based on the voice activation detection value and the power value within the human voice bandwidth. The first audio signal is the signal collected by the error microphone, and the active noise reduction audio signal is the signal output to the speaker for playback.

[0166] In one embodiment, as shown in FIG6 , the first probability value calculation module 1 includes a cancellation power spectrum calculation unit 11, a power value calculation unit 12 within the human voice bandwidth, a voice activation detection value calculation unit 13, and a first probability value calculation unit 14. The cancellation power spectrum calculation unit 11 is configured to calculate the power spectrum of each frame of the first audio signal and the power spectrum of each frame of the active noise reduction audio signal, and to offset the power value of the active noise reduction audio signal at each frequency point by the power value of the first audio signal through calibration or adaptive means to obtain a cancellation power spectrum. The power value calculation unit 12 within the human voice bandwidth is configured to integrate the cancellation power spectrum within a preset frequency range to obtain a power value within the human voice bandwidth. The voice activation detection value calculation unit 13 is configured to calculate a voice activation detection value based on the cancellation power spectrum. The first probability value calculation unit 14 is configured to calculate a first probability value of the wearer speaking based on the power value within the human voice bandwidth and the voice activation detection value.

[0167] In one embodiment, the power value calculation unit 12 within the human voice bandwidth may perform an integration operation on the cancellation power spectrum within a preset frequency range to obtain the power value within the human voice bandwidth. In a specific embodiment, the preset frequency range may be 100 Hz to 1 kHz.

[0168] In one embodiment, the voice activity detection value calculation unit 13 may calculate the voice activity detection value based on energy features, endpoint detection, or deep learning according to the cancellation power spectrum.

[0169] The second probability value calculation module 20 is used to calculate the power spectrum of the first audio signal and the power spectrum of the second audio signal when the first probability value is greater than or equal to the first preset threshold, calculate the correlation coefficient of the first audio signal and the second audio signal, and calculate the first sound source direction value based on the first audio signal and the second audio signal, and then calculate the second probability value of the wearer speaking based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value, where the second audio signal is a signal collected by the reference microphone.

[0170] In one embodiment, as shown in Figure 7, the second probability value calculation module 20 includes a power spectrum calculation unit 21, a correlation coefficient calculation unit 22, a first sound source direction value calculation unit 23 and a second probability value calculation unit 24, the power spectrum calculation unit 21 is used to calculate the power spectrum of the first audio signal and the power spectrum of the second audio signal, the correlation coefficient calculation unit 22 is used to calculate the correlation coefficient of the first audio signal and the second audio signal, the first sound source direction value calculation unit 23 is used to calculate the first sound source direction value based on the power spectrum of the first audio signal and the power spectrum of the second audio signal, and the second probability value calculation unit 24 is used to calculate the second probability value of the wearer speaking based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value.

[0171] In one embodiment, as shown in Figure 8, the first sound source direction value calculation unit 23 includes a first frequency response calculation unit 231, a first phase calculation unit 232 and a first delay calculation unit 233, wherein the first frequency response calculation unit 231 is used to calculate the frequency response of each frame of the first audio signal and each frame of the second audio signal, the first phase calculation unit 232 is used to calculate the phase information of each frame through the inverse tangent function according to the frequency response, and smooth the phase information, and the first delay calculation unit 233 is used to calculate the delay of the frequency point k according to the smoothed phase information to obtain the first sound source direction value.

[0172] The third probability value calculation module 30 is used to calculate the second sound source direction value based on the second audio signal and the third audio signal when the second probability value is greater than or equal to the second preset threshold, and calculate the third probability value of the wearer speaking based on the second sound source direction value, where the third audio signal is a signal collected by the call microphone.

[0173] In one embodiment, as shown in Figure 9, the third probability value calculation module 30 includes a second sound source direction value calculation unit 31 and a third probability value calculation unit 32. The second sound source direction value calculation unit 31 is used to calculate the second sound source direction value based on the power spectrum of the second audio signal and the power spectrum of the third audio signal. The third probability value calculation unit 32 is used to calculate the third probability value of the wearer speaking based on the second sound source direction value.

[0174] In one embodiment, as shown in Figure 10, the second sound source direction value calculation unit 31 includes a second frequency response calculation unit 311, a second phase calculation unit 312 and a second delay calculation unit 313. The second frequency response calculation unit 311 is used to calculate the frequency response of each frame of the second audio signal and each frame of the third audio signal; the second phase calculation unit 312 is used to calculate the phase information of each frame through the inverse tangent function according to the frequency response, and smooth the phase information. The second delay calculation unit 313 is used to calculate the delay of the frequency point k based on the smoothed phase information.

[0175] The mode switching module 30 is configured to output a mode switching signal to switch the headset from the active noise reduction mode to the non-active noise reduction mode when the third probability value is greater than or equal to a third preset threshold.

[0176] In one embodiment, the mode switching module 30 is also used to calculate a first probability value through the first probability value calculation module 1 when the headset is in the non-active noise reduction mode. If the first probability value is less than the first preset threshold, the headset switches to the active noise reduction mode; otherwise, the second probability value is calculated by the second probability value calculation module 2. If the second probability value is less than the second preset threshold, the headset switches to the active noise reduction mode; otherwise, the third probability value is calculated by the third probability value calculation module 3. If the third probability value is less than the third preset threshold, the headset switches to the active noise reduction mode; otherwise, the headset remains in the non-active noise reduction mode.

[0177] The present invention also provides an earphone mode switching chip capable of executing the TWS earphone mode switching method of the present invention.

[0178] The present invention also provides a TWS headset, including the TWS headset mode switching device of the present invention, or the headset mode switching chip of the present invention.

[0179] The present invention also provides a storage medium storing a program, wherein the program is executed to implement the TWS earphone mode switching method of the present invention.

[0180] It will be understood by those skilled in the art that, under the premise of no conflict, the above-mentioned preferred embodiments can be freely combined and superimposed. Among them, the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. The numbering of each step in this article is only for the convenience of description and reference, and is not used to limit the order of execution. The specific execution order is determined by the technology itself, and those skilled in the art can determine various allowable and reasonable orders based on the technology itself.

[0181] It should be noted that the use of step numbers (letters or numbers) to refer to certain specific method steps in the present invention is solely for the purpose of descriptive convenience and brevity, and is in no way intended to limit the order of these method steps. Those skilled in the art will appreciate that the order of the relevant method steps is determined by the technology itself and should not be unduly limited by the presence of step numbers. Those skilled in the art can determine various permissible and reasonable step orders based on the technology itself.

[0182] Those skilled in the art will appreciate that, provided there is no conflict, the above preferred solutions can be freely combined and superimposed.

[0183] It should be understood that the above-mentioned embodiments are merely illustrative and non-restrictive. Without departing from the basic principles of the present invention, various obvious or equivalent modifications or substitutions that can be made by those skilled in the art to the above-mentioned details will be included in the scope of the claims of the present invention.

Claims

1. A TWS headset mode switching method, used to switch a headset from an active noise reduction mode to a non-active noise reduction mode, the headset comprising an error microphone, a reference microphone, a call microphone and a speaker, characterized in that: The method comprises: S100: Calculating a cancellation power spectrum according to a first audio signal and an active noise reduction audio signal, calculating a voice activation detection value and a power value within a human voice bandwidth according to the cancellation power spectrum, and calculating a first probability value that the wearer is speaking according to the voice activation detection value and the power value within the human voice bandwidth, wherein the first audio signal is a signal collected by the error microphone, and the active noise reduction audio signal is a signal output to the speaker for playback; S200: when the first probability value is greater than or equal to a first preset threshold, calculating the power spectrum of the first audio signal and the power spectrum of the second audio signal, calculating the correlation coefficient between the first audio signal and the second audio signal, and calculating a first sound source direction value according to the first audio signal and the second audio signal, and then calculating a second probability value that the wearer is speaking according to the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value, wherein the second audio signal is a signal collected by the reference microphone; S300: when the second probability value is greater than or equal to a second preset threshold, calculating a second sound source direction value according to the second audio signal and the third audio signal, and calculating a third probability value that the wearer is speaking according to the second sound source direction value, wherein the third audio signal is a signal collected by the call microphone; S400: When the third probability value is greater than or equal to a third preset threshold, the headset switches from the active noise reduction mode to the non-active noise reduction mode.

2. The TWS headset mode switching method according to claim 1, characterized in that: The calculation of the cancellation power spectrum according to the first audio signal and the active noise reduction audio signal in S100 includes: Calculate the power spectrum Xf1(m, k) of each frame of the first audio signal and the power spectrum Xf4(m, k) of each frame of the active noise reduction audio signal; By calibration or adaptive means, at each frequency point, the power value of the first audio signal is offset by the power value of the active noise reduction audio signal to obtain the offset power spectrum Xf(m,k): Xf(m,k)=Xf1(m,k)-a(k)Xf4(m,k); Among them, m is the number of frames, k is the frequency point, and a(k) is the adaptive or calibrated weight.

3. The TWS headset mode switching method according to claim 2, characterized in that: Calculating the power value within the human voice bandwidth according to the cancellation power spectrum in S100 includes: An integration operation is performed on the cancellation power spectrum within the preset frequency range to obtain a power value XF(m) within the human voice bandwidth.

4. The TWS headset mode switching method according to claim 3, characterized in that: The preset frequency range is 100 Hz to 1 khz.

5. The TWS headset mode switching method according to claim 2, characterized in that: Calculating the voice activation detection value according to the cancellation power spectrum includes: The voice activation detection value is calculated based on the cancellation power spectrum, energy characteristics, endpoint detection or deep learning method.

6. The TWS headset mode switching method according to claim 2, characterized in that: In S100, the first probability value of the wearer speaking is calculated based on the power value within the human voice bandwidth and the voice activation detection value: result1 (m) = β 11 *p xf (m)+β 12 *p vad (m) Among them, β 11 , β 12 Weighted value, β 11 +β 12 =1, p xf (m) is the power probability value within the human voice bandwidth, p vad (m) is the probability value of the voice activation detection value, p vad (m) is equal to the voice activation detection value, and m is the number of frames.

7. The TWS headset mode switching method according to claim 6, characterized in that: The p xf (m) is: When the XF(m) is less than or equal to the preset power reference value, the p xf (m) is the XF(m) divided by the first preset power reference value, otherwise the p xf (m) is 1, and the preset power reference value is a power reference value set according to the sensitivity of the error microphone and the ADC gain.

8. The TWS headset mode switching method according to claim 1, characterized in that: The correlation coefficient between the first audio signal and the second audio signal calculated in S200 is: Wherein, f1[n] is the sampling value of the first audio signal, f2[n+k] is the sampling value of the second audio signal, n is the sampling point, and k is the number of sampling points of the second audio signal delayed relative to the first audio signal.

9. The TWS headset mode switching method according to claim 1, characterized in that: Calculating the first sound source direction value according to the first audio signal and the second audio signal in S200 includes: S201: Calculate the frequency response between each frame of the first audio signal and each frame of the second audio signal: Wherein, H1(m,k) includes the sound source phase information, m is the frame number, k is the frequency point, Xf1(m,k) is the power spectrum of the first audio signal, and Xf2(m,k) is the power spectrum of the second audio signal; S202: Calculate the phase information of each frame using an inverse tangent function according to the frequency response: phase(m,k)=∠H1(m,k) S203: Smoothing the phase information: phase(m,k)=(1-α)phase(m-1,k)+α*phase(m,k) Among them, α is the smoothing factor; S204: Calculate the delay of frequency point k according to the smoothed phase information: t1=phase(k) / (2*pi*k / L*Fs) t1 is the first sound source direction value, L is the number of sampling points, and Fs is the sampling rate.

10. The TWS headset mode switching method according to claim 1, characterized in that: The step S200 of calculating the second probability value of the wearer speaking according to the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value includes: result (m) = β 21 *p pwr (m)+β 22 *p rec (m)+β 23 *p pha1 (m) Among them, β 21 , β 22 , β 23 Weighted value, β 21 +β 22 +β 23 =1, p pwr (m) is a weighted value of the power probability value of the first audio signal and the power probability value of the second audio signal; The p rec (m) is obtained by normalizing the correlation coefficient between the first audio signal and the second audio signal; The p pha1 (m) is based on the first sound source direction value t1, the preset first delay threshold T ref1 and the time parameter δ T1 Calculated; m is the number of frames.

11. The TWS headset mode switching method according to claim 10, characterized in that: The p pwr (m) is: Calculate a power spectrum Xf1(m, k) of each frame of the first audio signal and a power spectrum Xf2(m, k) of each frame of the second audio signal; Performing integration operations on Xf1(m, k) and Xf2(m, k) within a preset frequency range to obtain a first audio signal power XF1(m) and a second audio signal power XF2(m); When the XF1(m) is less than or equal to the first preset power reference value, the p xf1 (m) is the XF1(m) divided by the first preset power reference value, otherwise the p xf1 (m) is 1, and the first preset power reference value is a power reference value set according to the error microphone sensitivity and the ADC gain; When the XF2(m) is less than or equal to the second preset power reference value, the p xf2 (m) is the XF2(m) divided by the second preset power reference value, otherwise the p xf2 (m) is 1, and the second preset power reference value is a power reference value set according to the reference microphone sensitivity and the ADC gain; The p pwr For the p xf1 (m) and p xf2 The weighted value of (m).

12. The TWS headset mode switching method according to claim 1, characterized in that: Calculating the second sound source direction value according to the second audio signal and the third audio signal in S300 includes: S301: Calculate the frequency response between each frame of the second audio signal and each frame of the third audio signal: Wherein, H2(m,k) includes the sound source phase information, m is the frame number, k is the frequency point, Xf2(m,k) is the power spectrum of the second audio signal, and Xf3(m,k) is the power spectrum of the third audio signal; S302: Calculate the phase information of each frame according to the frequency response using an inverse tangent function: phase(m,k)=∠H2(m,k) S303: Smoothing the phase information. phase(m,k)=(1-α)phase(m-1,k)+α*phase(m,k) α is the smoothing factor; S304: Calculate the delay of frequency point k according to the smoothed phase information: t2=phase(k) / (2*pi*k / L*Fs) L is the number of sampling points, Fs is the sampling rate, and t2 is the second sound source direction value.

13. The TWS headset mode switching method according to claim 12, characterized in that: Calculating the third probability value of the wearer speaking according to the second sound source direction value in S300 includes: result3 (m) = p pha2 (m) The p pha2 (m) is based on the second sound source direction value t2, the preset first delay threshold T ref2 and the time parameter δ T2 Calculated.

14. The TWS headset mode switching method according to claims 1-13, characterized in that: The method further comprises: S500: When the headset is in the non-active noise reduction mode, within the preset time, the first probability value is calculated according to the S100, if the first probability value is less than the first preset threshold, the headset switches to the active noise reduction mode, otherwise, the second probability value is calculated according to the S200, if the second probability value is less than the second preset threshold, the headset switches to the active noise reduction mode, otherwise, the third probability value is calculated according to the S300, if the third probability value is less than the third preset threshold, the headset switches to the active noise reduction mode, otherwise, the headset remains in the non-active noise reduction mode.

15. A TWS headset mode switching device, used to switch the headset from an active noise reduction mode to a non-active noise reduction mode, the headset comprising an error microphone, a reference microphone, a call microphone and a speaker, characterized in that: The device comprises a first probability value calculation module, a second probability value calculation module, a third probability value calculation module and a mode switching module: The first probability value calculation module is used to calculate a cancellation power spectrum according to the first audio signal and the active noise reduction audio signal, calculate a voice activation detection value and a power value within the human voice bandwidth according to the cancellation power spectrum, and calculate a first probability value of the wearer speaking according to the voice activation detection value and the power value within the human voice bandwidth, wherein the first audio signal is a signal collected by the error microphone, and the active noise reduction audio signal is a signal output to the speaker for playback; The second probability value calculation module is used to calculate the power spectrum of the first audio signal and the power spectrum of the second audio signal when the first probability value is greater than or equal to a first preset threshold, calculate the correlation coefficient between the first audio signal and the second audio signal, and calculate a first sound source direction value according to the first audio signal and the second audio signal, and then calculate a second probability value that the wearer is speaking according to the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value, wherein the second audio signal is a signal collected by the reference microphone; The third probability value calculation module is used to calculate a second sound source direction value according to the second audio signal and the third audio signal when the second probability value is greater than or equal to a second preset threshold, and calculate a third probability value that the wearer is speaking according to the second sound source direction value, wherein the third audio signal is a signal collected by the call microphone; The mode switching module is used to output a mode switching signal when the third probability value is greater than or equal to a third preset threshold value, so as to switch the headset from the active noise reduction mode to the non-active noise reduction mode.

16. The TWS earphone mode switching device according to claim 15, characterized in that: The first probability value calculation module includes a power spectrum cancellation calculation unit, a power value calculation unit within the human voice bandwidth, a voice activation detection value calculation unit and a first probability value calculation unit. The cancellation power spectrum calculation unit is used to calculate the power spectrum of each frame of the first audio signal and the power spectrum of each frame of the active noise reduction audio signal, and to cancel the power value of the active noise reduction audio signal by the power value of the first audio signal at each frequency point in a calibrated or adaptive manner to obtain the cancellation power spectrum; The power value calculation unit within the human voice bandwidth is used to perform an integration operation on the cancellation power spectrum within a preset frequency range to obtain the power value within the human voice bandwidth; The voice activity detection value calculation unit is used to calculate the voice activity detection value according to the cancellation power spectrum; The first probability value calculation unit is used to calculate a first probability value of the wearer speaking according to the power value within the human voice bandwidth and the voice activation detection value.

17. The TWS earphone mode switching device according to claim 15, characterized in that: The second probability value calculation module includes a power spectrum calculation unit, a correlation coefficient calculation unit, a first sound source direction value calculation unit and a second probability value calculation unit. The power spectrum calculation unit is used to calculate the power spectrum of the first audio signal and the power spectrum of the second audio signal; The correlation coefficient calculation unit is used to calculate the correlation coefficient between the first audio signal and the second audio signal; The first sound source direction value calculation unit is used to calculate a first sound source direction value according to a power spectrum of the first audio signal and a power spectrum of the second audio signal; The second probability value calculation unit is used to calculate a second probability value that the wearer is speaking based on the power spectrum of the first audio signal, the power spectrum of the second audio signal, the correlation coefficient and the first sound source direction value.

18. The TWS earphone mode switching device according to claim 17, characterized in that: The first sound source direction value calculation unit includes a first frequency response calculation unit, a first phase calculation unit and a first delay calculation unit. The first frequency response calculation unit is used to calculate the frequency response of each frame of the first audio signal and each frame of the second audio signal; The first phase calculation unit is used to calculate the phase information of each frame according to the frequency response through an inverse tangent function, and smooth the phase information: The first delay calculation unit is used to calculate the frequency point k according to the smoothed phase information delay.

19. The TWS headset mode switching device according to claim 15, characterized in that: The third probability value calculation module includes a second sound source direction value calculation unit and a third probability value calculation unit. The second sound source direction value calculation unit is used to calculate a second sound source direction value according to the power spectrum of the second audio signal and the power spectrum of the third audio signal; The third probability value calculation unit is used to calculate a third probability value that the wearer is speaking according to the second sound source direction value.

20. The TWS earphone mode switching device according to claim 19, characterized in that: The second sound source direction value calculation unit includes a second frequency response calculation unit, a second phase calculation unit and a second delay calculation unit. The second frequency response calculation unit is used to calculate the frequency response of each frame of the second audio signal and each frame of the third audio signal; The second phase calculation unit is used to calculate the phase information of each frame according to the frequency response by using an inverse tangent function, and to perform smoothing on the phase information; The second delay calculation unit is used to calculate the delay of the frequency point k according to the smoothed phase information.

21. The TWS earphone mode switching device according to claim 15, characterized in that: The mode switching module is also used to calculate the first probability value through the first probability value calculation module within a preset time when the headset is in the non-active noise reduction mode. If the first probability value is less than the first preset threshold, the headset switches to the active noise reduction mode; otherwise, calculate the second probability value through the second probability value calculation module. If the second probability value is less than the second preset threshold, the headset switches to the active noise reduction mode; otherwise, calculate the third probability value through the third probability value calculation module. If the third probability value is less than the third preset threshold, the headset switches to the active noise reduction mode; otherwise, the headset remains in the non-active noise reduction mode.

22. A headphone mode switching chip, characterized in that: The method for switching the TWS headset mode as described in any one of claims 1 to 14 can be executed.

23. A TWS headset, characterized in that: It comprises a TWS earphone mode switching device as described in any one of claims 15 to 21, or an earphone mode switching chip as described in claim 22.

24. A storage medium, characterized in that The storage medium stores a program, wherein the program is executed to implement the TWS headset mode switching method as described in any one of claims 1-14.

Citation Information

Patent Citations

  • User interface for ANR headphones with active hear-through

    CN104871556A

  • Mode control method and device and terminal equipment

    CN113873379A

  • Earphone mode switching method and device, electronic equipment and storage medium

    CN116320872A

  • Method and apparatus for resisting noise based on adaptive nonlinear spectral subtraction

    CN1841500A

  • noise reduction and audiovisual voice activity detection

    DE60319796T2