Headphone control method, device, headphone, and computer-readable storage medium

By obtaining the spectrum amplitude difference between the environmental frequency domain signal and the ear channel frequency domain signal, and automatically switching the headphone mode to the transparent mode, the problem of inconvenience in noise reduction headphones in the prior art is solved, and the flexibility and user experience of headphone control are improved.

CN115396776BActive Publication Date: 2025-09-02BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211027381.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2025-09-02
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

When the wearer needs to communicate with others, the existing noise cancellation headset needs to manually adjust the noise cancellation mode or remove the headset, which leads to inconvenience in use.

Method used

By obtaining the spectrum amplitude difference between the environmental frequency domain signal and the ear channel frequency domain signal, the voice detection strategy is used to automatically switch the headphone mode to the transparent mode to avoid dependence of hardware sensors.

Benefits of technology

Improves the flexibility and user experience of headphone control, reduces power consumption, and can communicate without manual operation of users, reducing false detection situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115396776B_ABST
    Figure CN115396776B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, device, headset and computer-readable storage medium for controlling headphones, wherein the headset control method includes: obtaining an ambient frequency domain signal, which is a frequency domain expression of a sound signal in the environment surrounding the headphones; obtaining an ear canal frequency domain signal, which is a frequency domain expression of a sound signal in the ear canal of a user wearing the headphones; obtaining a spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal, the spectrum amplitude difference representing the difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal; if it is determined based on the spectrum amplitude difference and a preset voice detection strategy that the user wearing the headphones has voice activity, then controlling the mode of the headphones to switch to a transparent mode. This application determines whether the user wearing the headphones has voice activity by utilizing the spectrum amplitude difference obtained from the ambient frequency domain signal and the ear canal frequency domain signal, which not only improves the intelligence of the headset control, but also effectively reduces the power consumption of the headphones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of earphone technology, and in particular to an earphone control method and device, an earphone, and a computer-readable storage medium. Background Art

[0002] In recent years, with the rapid global adoption of new-generation consumer electronics devices like smartphones and tablets, headphone sales, particularly wireless headphones, have seen explosive growth. Noise-canceling headphones, which block out external noise and improve sound quality, are gaining popularity.

[0003] However, since noise-canceling headphones suppress both ambient noise and human speech, this can hinder the wearer's ability to communicate with others. In existing technology, if the wearer needs to communicate with others while wearing headphones, they must either remove the headphones or manually turn off the noise-canceling mode, which is very inconvenient. Therefore, finding better control over headphones is a pressing technical issue. Summary of the Invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a method and device for controlling an earphone, an earphone, and a computer-readable storage medium.

[0005] According to a first aspect of an embodiment of the present disclosure, a method for controlling an earphone is provided, the method comprising:

[0006] Acquire an ambient frequency domain signal, which is a frequency domain representation of the sound signal in the environment surrounding the headset;

[0007] Acquire an ear canal frequency domain signal, where the ear canal frequency domain signal is a frequency domain representation of a sound signal in the ear canal of a user wearing headphones;

[0008] Obtaining a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal, where the spectrum amplitude difference represents a difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal;

[0009] If it is determined based on the spectrum amplitude difference and the preset voice detection strategy that the user wearing the headset has voice activity, the mode of the headset is controlled to switch to the transparent mode.

[0010] Optionally, before controlling the mode of the headphones to switch to the transparency mode, the method further includes:

[0011] Get the power consumption mode of the headset, which includes low power consumption mode and high performance mode;

[0012] If the power consumption mode is the high-performance mode, determining whether the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset first voice detection strategy;

[0013] If the power consumption mode is the low power consumption mode, it is determined whether the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset second voice detection strategy.

[0014] Optionally, obtaining the power consumption mode of the headset includes:

[0015] Get the power consumption mode preset by the user, or

[0016] Acquire performance parameters of the headset, and determine a corresponding power consumption mode according to the performance parameters of the headset, wherein the performance parameters of the headset include at least one of a remaining battery charge and a CPU load.

[0017] Optionally, determining whether the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset first voice detection strategy includes:

[0018] The spectrum amplitude difference is input into a preset speech detection model to obtain a speech detection result, where the speech detection model is obtained by performing machine learning training on an initial model using a training set, wherein the training set includes a positive sample data set and a negative sample data set, the positive sample data set includes a plurality of positive sample data, each positive sample data includes first audio data and a first label corresponding to the first audio data, the first audio data includes a user's speaking voice when the user is wearing headphones, and the first label is used to indicate that the corresponding first audio data is the target audio, the negative sample data set includes a plurality of negative sample data, each negative sample data includes second audio data and a second label corresponding to the second audio data, the second audio data does not include the user's speaking voice when the user is wearing headphones, and the second label is used to indicate that the corresponding second audio data is not the target audio;

[0019] Based on the voice detection result, it is determined whether the user wearing the headset has voice activity.

[0020] Optionally, the spectrum amplitude difference includes a plurality of sub-amplitude differences, and determining whether the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset second voice detection strategy includes:

[0021] Obtaining a perceptual weighted value, which is the sum of the products of each sub-amplitude difference and the corresponding perceptual weighting coefficient, where the perceptual weighting coefficient is the weighted value corresponding to different frequency bands;

[0022] It is determined whether there is voice activity of the user wearing the headset based on the perception weight value.

[0023] Optionally, if it is determined based on the spectrum amplitude difference and a preset voice detection strategy that the user wearing the headset has voice activity, controlling the headset to switch to a transparent mode includes:

[0024] If it is determined based on the spectrum amplitude difference and a preset voice detection strategy that the user wearing the headset has voice activity, then obtaining the type of voice activity;

[0025] If the type of the voice activity is a conversation type, the mode of the headset is controlled to switch to a transparent mode;

[0026] If the voice activity is non-conversational, the headset remains in the same mode.

[0027] Optionally, obtaining a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal includes:

[0028] Acquire a speaker reference signal, where the speaker reference signal is an original sound signal of an audio signal to be played by a speaker of the headset, and the speaker reference signal includes multiple audio frames;

[0029] Determining the sum of energies of the respective audio frames included in the loudspeaker reference signal;

[0030] If the total energy of the audio frame is less than the preset energy, the spectrum amplitude difference is obtained according to the ambient frequency domain signal and the ear canal frequency domain signal.

[0031] Optionally, the method further includes:

[0032] If the total energy of the audio frame is greater than or equal to the preset energy, echo cancellation is performed on the ambient frequency domain signal and the ear canal frequency domain signal based on the frequency domain amplitude corresponding to the loudspeaker reference signal to obtain the ambient frequency domain signal and the ear canal frequency domain signal after echo cancellation;

[0033] The spectrum amplitude difference is obtained based on the ambient frequency domain signal and the ear canal frequency domain signal after echo cancellation.

[0034] Optionally, obtaining an ambient frequency domain signal includes:

[0035] Acquire the ambient time domain signal, which refers to the time domain expression of the sound signal in the environment surrounding the headset;

[0036] Performing time-frequency conversion on the ambient time domain signal to obtain the ambient frequency domain signal;

[0037] Acquire ear canal frequency domain signals, including:

[0038] Acquiring an ear canal time domain signal, where the ear canal time domain signal refers to a time domain representation of a sound signal in the ear canal of a user wearing headphones;

[0039] Perform time-frequency conversion on the ear canal time domain signal to obtain the ear canal frequency domain signal.

[0040] According to a second aspect of an embodiment of the present disclosure, a control device for an earphone is provided, wherein the earphone includes a feedforward microphone and a feedback microphone, and the device includes:

[0041] A first acquisition module is configured to acquire an ambient frequency domain signal, where the ambient frequency domain signal is a frequency domain expression of a sound signal in an environment surrounding the earphone;

[0042] a second acquisition module configured to acquire an ear canal frequency domain signal, where the ear canal frequency domain signal is a frequency domain expression of a sound signal in the ear canal of a user wearing the earphone;

[0043] a determination module configured to obtain a spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal, wherein the spectrum amplitude difference represents a difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal;

[0044] The switching module is configured to control the mode of the headset to switch to the transparent mode if it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset voice detection strategy.

[0045] According to a third aspect of an embodiment of the present disclosure, there is provided an earphone, comprising:

[0046] processor;

[0047] a memory for storing processor-executable instructions;

[0048] Feedforward microphone and feedback microphone;

[0049] Wherein, the processor is configured to:

[0050] Acquire an ambient frequency domain signal, which is a frequency domain representation of the sound signal in the environment surrounding the headset;

[0051] Acquire an ear canal frequency domain signal, where the ear canal frequency domain signal is a frequency domain representation of a sound signal in the ear canal of a user wearing headphones;

[0052] Obtaining a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal, where the spectrum amplitude difference represents a difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal;

[0053] If it is determined based on the spectrum amplitude difference and the preset voice detection strategy that the user wearing the headset has voice activity, the mode of the headset is controlled to switch to the transparent mode.

[0054] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the headset control method provided by the first aspect of the present disclosure are implemented.

[0055] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: when obtaining the ambient frequency domain signal and the ear canal frequency domain signal, the present application can obtain the spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal. On this basis, if it is determined that the user wearing the headphones has voice activity based on the spectrum amplitude difference and the preset voice detection strategy, the mode of the headphones is controlled to switch to the transparent mode. By utilizing the spectrum amplitude difference and the preset voice detection strategy to determine whether the user wearing the headphones has voice activity, the present application can not only improve the flexibility of the headphone control, but also reduce the power consumption of the headphone control, and can achieve control of the headphones without user intervention, thereby improving the user experience.

[0056] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0058] Figure 1 The figure is a flowchart of a method for controlling an earphone according to an exemplary embodiment.

[0059] Figure 2 The figure is a schematic structural diagram of an earphone in a method for controlling an earphone according to an exemplary embodiment.

[0060] Figure 3 This is an example diagram showing switching a headphone mode to a transparent mode in a headphone control method according to an exemplary embodiment.

[0061] Figure 4 The figure is a flow chart of a method for controlling an earphone according to another exemplary embodiment.

[0062] Figure 5 The figure is a block diagram of a device for controlling an earphone according to an exemplary embodiment.

[0063] Figure 6 The figure is a block diagram of a headset according to an exemplary embodiment. DETAILED DESCRIPTION

[0064] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0065] In recent years, the TWS (True Wireless Stereo) headphone market has continued to boom, with the industry expanding. Major headphone manufacturers have launched a variety of feature-rich products. Active noise cancellation and transparency are now essential features for TWS headphones. With the increasing demand for intelligent electronic products, many products are equipped with smart hands-free functions, which automatically switch to transparency mode in specific scenarios.

[0066] The intelligent hands-free function is also known as the adaptive transparency function. When the headset user is wearing the headset, if it detects that the user is talking, the headset will automatically switch from noise reduction mode to transparency mode, allowing the user to communicate with others without taking off the headset.

[0067] Most smart earphone removal solutions currently on the market rely on hardware sensors like bone vibration sensors. These solutions require additional hardware to achieve this, increasing the cost of the headset. Furthermore, these solutions are prone to false detection when vocal cord vibrations, such as coughing, occur.

[0068] In response to the above problems, the present application proposes a control method, device, headphones and computer-readable storage medium for headphones. By utilizing the spectrum amplitude difference obtained between the ambient frequency domain signal and the ear canal frequency domain signal to determine whether the user wearing the headphones has voice activity, it can not only effectively reduce false detection of vocal cord vibration scenarios such as coughing, but also effectively improve the user experience.

[0069] Figure 1 FIG. 1 is a flow chart showing a method for controlling an earphone according to an exemplary embodiment. Figure 1 As shown, the earphone control method is used in the earphone and includes the following steps.

[0070] In step S11, an environmental frequency domain signal is acquired.

[0071] In an embodiment of the present application, the earphones may be active noise reduction earphones, and the default operating mode of the earphones may be a noise reduction mode. The present application may obtain an ambient frequency domain signal when detecting that a user is wearing earphones, wherein the ambient frequency domain signal may be a frequency domain expression of a sound signal in the environment surrounding the earphones. Optionally, the embodiment of the present application may also obtain an ambient frequency domain signal when detecting that the noise reduction mode of the earphones is turned on. In addition, the embodiment of the present application may also obtain an ambient frequency domain signal when detecting that the noise reduction mode of the earphones is turned on.

[0072] As an optional method, the sound signal in the environment around the headset can be collected by the headset's feedforward microphone. Figure 2As shown, the headset in the embodiment of the present application may include a feed-forward microphone 101. The feed-forward microphone (FF MIC) 101 is set on the outside of the headset and can receive both the wearer's voice signal and the voice signals of other people around the wearer. Optionally, the present application can obtain the ambient frequency domain signal within the current sampling period.

[0073] As another optional method, when obtaining the ambient frequency domain signal, the present application can obtain the ambient time domain signal and perform time-frequency conversion on the ambient time domain signal to obtain the ambient frequency domain signal, wherein the ambient time domain signal refers to the time domain expression of the sound signal in the environment surrounding the headphone.

[0074] As an optional method, after collecting the sound signal in the environment surrounding the headphone, the embodiment of the present application can perform frame-by-frame overlapping and windowing processing on the sound signal in the environment surrounding the headphone to obtain a frame-level signal. The sampling frequency of the sound signal in the environment surrounding the headphone can be greater than a preset frequency, which can be 8000. For example, the sampling frequency of the sound signal in the environment surrounding the headphone can be 16000.

[0075] Optionally, when performing frame overlapping and windowing processing on the sound signal in the environment surrounding the headphone, the frame length of the frame overlapping can be any one of 16ms, 32ms, and 64ms. In the embodiment of the present application, the frame length of the frame overlapping can be 32ms. In addition, the frame shift in the embodiment of the present application can be 25%, 50%, or 75% of the frame length. The present application preferably uses a frame shift of 50% of the frame length. Optionally, the windowing in the embodiment of the present application can use a Hanning window, a Hamming window, etc. The present application prefers a Hanning window. In addition, the window length and the frame length when performing frame overlapping and windowing processing can be the same.

[0076] As an optional method, when performing frame-by-frame overlapping and windowing processing on the sound signal in the environment surrounding the headphone, the present application can first perform frame-by-frame processing on the sound signal in the environment surrounding the headphone to obtain multiple frames of acoustic signals. On this basis, the acoustic signal is converted into an ambient time domain signal to obtain multiple frames of ambient time domain signals. The calculation formula for the ambient time domain signal can be as follows.

[0077] ff f (m,n)=ff((m-1)*inc+n)*w(n),0≤n≤(L-1);

[0078] Among them, m represents the frame number index, n represents the data point index of the audio data, and ff f (m) represents the mth frame of the ambient time domain signal, ff f(m,n) represents the nth time domain sampling point of the mth frame ambient time domain signal. inc is the sampling point length of the frame shift, w(n) is the window function, and L is the sampling point length of the frame length. The calculation formula for the sampling point length of the frame length (L) can be: sampling rate (fs) * frame length (ft) / 1000. On this basis, the embodiment of the present application can perform a fast Fourier transform on the ambient time domain signal to obtain a frame-level feedforward microphone frequency domain amplitude signal (FF(m)), that is, to obtain the ambient frequency domain signal FF(m).

[0079] The calculation formula of the environmental frequency domain signal FF(m) can be shown as follows.

[0080]

[0081] Where m represents the index of the frame number, n represents the index of the frequency point, FF(m) represents the frequency domain amplitude signal of the feedforward microphone of the mth frame, that is, the ambient frequency domain signal, and FF(m,k) represents the kth frequency domain sampling point of the ambient frequency domain signal of the mth frame. Due to the conjugate symmetry of the discrete Fourier transform, this application can only take the forward frequency domain data. Point data.

[0082] In step S12, an ear canal frequency domain signal is acquired.

[0083] As an optional method, the present application can also obtain ear canal frequency domain signals when detecting that the user is wearing headphones, wherein the ear canal frequency domain signal can be a frequency domain expression of the sound signal in the ear canal of the user wearing headphones. Optionally, the embodiment of the present application can also obtain ear canal frequency domain signals when detecting that the noise reduction mode of the headphones is turned on. In addition, the embodiment of the present application can also obtain ear canal frequency domain signals when detecting that the noise reduction mode of the headphones is turned on.

[0084] As an optional method, the sound signal in the ear canal of the user wearing the earphone can be collected by the feedback microphone of the earphone. Figure 2 As shown, the earphones in the embodiment of the present application may include a feedback microphone 102. The feedback microphone (FB MIC) 102 is disposed on the inner side of the earphone and is used to receive the wearer's speech signal transmitted to the ear canal by the Eustachian tube. Optionally, the present application may obtain the ear canal frequency domain signal within the current sampling period.

[0085] As another optional method, when obtaining the ear canal frequency domain signal, the present application can obtain the ear canal time domain signal and perform time-frequency conversion on the ear canal time domain signal to obtain the ear canal frequency domain signal, wherein the ear canal time domain signal refers to the time domain expression of the sound signal in the ear canal of the user wearing headphones.

[0086] As an optional method, after collecting the sound signal in the ear canal of the user wearing the headphones, the embodiment of the present application can perform frame-by-frame overlapping and windowing processing on the sound signal in the ear canal of the user wearing the headphones to obtain a frame-level signal. The sampling frequency of the sound signal in the ear canal of the user wearing the headphones is similar to the sampling frequency of the sound signal in the environment surrounding the headphones, and will not be further described here.

[0087] In addition, the process of performing frame-by-frame overlapping and windowing processing on the sound signal in the ear canal of a user wearing headphones is similar to the process of performing frame-by-frame overlapping and windowing processing on the sound signal in the environment surrounding the headphones. Alternatively, the sound signal in the ear canal of a user wearing headphones can be framed to obtain multiple frames of acoustic signals. Based on this, the acoustic signal is converted into an ear canal time domain signal to obtain multiple frames of ear canal time domain signals. The calculation formula for the ear canal time domain signal can be as follows.

[0088] fb f (m, n)=fb((m-1)*inc+n)*w(n), 0≤n≤(L-1);

[0089] Among them, fb f (m) represents the mth frame of the ear canal time domain signal, fb f (m,n) represents the nth time domain sampling point of the mth frame of the ear canal time domain signal.

[0090] On this basis, embodiments of the present application can perform a fast Fourier transform on the ear canal time domain signal to obtain a frame-level feedback microphone frequency domain amplitude signal (FB(m)), that is, obtain the ear canal frequency domain signal FB(m). The calculation formula for the ear canal frequency domain signal FB(m) can be shown below.

[0091]

[0092] Wherein, FB(m) represents the frequency domain amplitude signal of the feedback microphone of the mth frame, that is, the ear canal frequency domain signal, and FF(m,k) represents the kth frequency domain sampling point of the ear canal frequency domain signal of the mth frame.

[0093] It should be noted that, in addition to the feedforward microphone and the feedback microphone, the headset may also include Figure 2 The speaker 103 shown can be used to output voice information. While obtaining the ambient frequency domain signal and the ear canal frequency domain signal, the embodiment of the present application can also obtain a speaker reference signal. The times corresponding to the ambient frequency domain signal, the ear canal frequency domain signal, and the speaker reference signal can be the same.

[0094] In step S13, a spectrum amplitude difference is obtained according to the ambient frequency domain signal and the ear canal frequency domain signal.

[0095] As an optional method, after obtaining the ambient frequency domain signal and the ear canal frequency domain signal, the embodiment of the present application can obtain the spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal, wherein the spectrum amplitude difference represents the difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal, and the calculation formula of the spectrum amplitude difference is shown below.

[0096]

[0097] Among them, FC(m) represents the spectrum amplitude difference of the mth frame, and FC(m,k) represents the kth frequency domain sampling point (frequency point) of the spectrum amplitude difference of the mth frame.

[0098] In step S14, if it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and the preset voice detection strategy, the headset mode is controlled to switch to the transparent mode.

[0099] As an optional method, after obtaining the spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal, the embodiment of the present application can determine whether the user wearing the headset has voice activity based on the spectrum amplitude difference and the preset voice detection strategy. If it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and the preset voice detection strategy, the embodiment of the present application can control the mode of the headset to switch to transparent mode. In the embodiment of the present application, transparent mode means that the headset collects ambient sound, filters the ambient sound and outputs it, and superimposes the sound leaked into the human ear, so that the human ear receives the complete ambient sound.

[0100] In other words, when it is determined that the user wearing the headset has voice activity, the present application can switch the headset mode from active noise reduction mode to transparency mode. The switching of the headset mode can be as follows: Figure 3 As shown. Figure 3 It is known that the embodiment of the present application can switch the mode from active noise reduction mode to transparency mode when voice activity is detected by the user wearing the headphones.

[0101] In addition, if it is determined based on the spectrum amplitude difference and the preset voice detection strategy that the user wearing the headset has no voice activity, the embodiment of the present application can keep the headset mode unchanged in the active noise reduction mode, that is, no mode switching operation is performed.

[0102] Optionally, if it is determined based on the spectrum amplitude difference and the preset voice detection strategy that the user wearing the headphones has no voice activity, the embodiment of the present application can also determine whether the user wearing the headphones is in a specified scene. If it is determined that the user wearing the headphones is in the specified scene, the embodiment of the present application can also switch the mode of the headphones from noise reduction mode to transparency mode.

[0103] As a specific implementation, when determining whether a user wearing headphones is in a specified scene, the present application can determine whether the headphones have collected a specified type of audio signal. If it is determined that the headphones have collected a specified type of audio signal, the present embodiment of the application can determine that the user is in the specified scene. The specified type of audio signal can be a whistle signal or a call signal.

[0104] In addition, if it is determined that the user wearing the headphones has no voice activity and is not in a specified scenario, the embodiment of the present application can keep the mode of the headphones unchanged, that is, keep the mode of the headphones in the noise reduction mode.

[0105] As an optional approach, when determining whether the user wearing the headset is engaging in voice activity based on the spectrum amplitude difference and a preset voice detection strategy, embodiments of the present application can utilize different voice detection strategies to determine whether the user wearing the headset is engaging in voice activity. The specific methods for utilizing different voice detection strategies to determine whether the user wearing the headset is engaging in voice activity will be described in detail in the following embodiments and will not be further elaborated upon here.

[0106] When the present application obtains the ambient frequency domain signal and the ear canal frequency domain signal, it can obtain the spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal. On this basis, if it is determined that the user wearing the headphones has voice activity based on the spectrum amplitude difference and the preset voice detection strategy, the headset mode is controlled to switch to the transparent mode. By using the spectrum amplitude difference and the preset voice detection strategy to determine whether the user wearing the headphones has voice activity, the present application can not only improve the flexibility of the headset control, but also reduce the power consumption of the headset control, and can achieve control of the headset without user intervention, thereby improving the user experience.

[0107] Figure 4 FIG. 1 is a flow chart showing a method for controlling an earphone according to an exemplary embodiment. Figure 4 As shown, the earphone control method is used in the earphone and includes the following steps.

[0108] In step S21, an environmental frequency domain signal is acquired.

[0109] In step S22, an ear canal frequency domain signal is acquired.

[0110] Among them, step S21 to step S22 have been described in detail in the above embodiment and will not be repeated here.

[0111] In step S23, a spectrum amplitude difference is obtained according to the ambient frequency domain signal and the ear canal frequency domain signal.

[0112] As an optional method, when the spectrum amplitude difference is obtained based on the ambient frequency domain signal and the ear canal frequency domain signal, the embodiment of the present application can also obtain a speaker reference signal, wherein the speaker reference signal can be the original sound signal of the audio signal to be played by the speaker of the earphone, and the speaker reference signal includes multiple audio frames.

[0113] On this basis, the embodiment of the present application can obtain the sum of the energies of each audio frame contained in the speaker reference signal, and then determine whether the sum of the energies of the audio frames is less than the preset energy. If it is determined that the sum of the energies of the audio frames is less than the preset energy, the embodiment of the present application can obtain the spectral amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal.

[0114] Specifically, when obtaining the summed energy of each audio frame contained in the speaker reference signal, embodiments of the present application may also perform frame-by-frame overlapping and windowing processing on the speaker reference signal to obtain a multi-frame acoustic signal. Based on this, the acoustic signal is converted into a time domain signal to obtain a multi-frame reference time domain signal. The calculation formula for this reference time domain signal can be as follows.

[0115] spk f (m, n)=spk((m-1)*inc+n)*w(n), 0≤n≤(L-1);

[0116] Among them, m represents the frame number index, n represents the data point index of the audio data, spk f (m) represents the reference time domain signal of the mth frame, spk f (m,n) represents the nth time domain sampling point of the mth frame reference signal. inc is the sampling point length of the frame shift, w(n) is the window function, and L is the sampling point length of the frame length. The sampling point length (L) of the frame length can be calculated as: sampling rate (fs) * frame length (ft) / 1000.

[0117] As an example, in the embodiment of the present application, the sampling frequency may be 16000, and the frame length ft may be 32, that is, the frame length L may be 16000*32 / 1000=512. In addition, the frame shift may be 50% of the frame length, that is, the sampling point length inc of the frame shift may be 256, and the window function w(n) may be a Hanning window.

[0118] On this basis, the embodiment of the present application can calculate the sum of the energies of each audio frame contained in the speaker reference signal. The calculation formula of the energy sum is as follows.

[0119]

[0120] Wherein, spk_e(m) is the sum of the energies of each audio frame contained in the speaker reference signal, m represents the frame number index, n represents the data point index of the audio data, and L is the sampling point length of the frame length.

[0121] As an optional approach, after obtaining the summed energy of each audio frame included in the speaker reference signal, embodiments of the present application can determine whether the summed energy of the audio frames is less than a preset energy value to determine whether the headset is playing content. If the summed energy of the audio frames is less than the preset energy value, it is determined that the headset is not playing content, and the spectrum amplitude difference can be directly calculated.

[0122] Optionally, if the total energy of the audio frames is determined to be greater than or equal to a preset energy, it is determined that the headphones are playing content. In this case, echo cancellation can be performed on the ambient frequency domain signal and the ear canal frequency domain signal based on the frequency domain amplitude corresponding to the speaker reference signal to obtain the echo-cancelled ambient frequency domain signal and the ear canal frequency domain signal. Based on this, a spectral amplitude difference is obtained based on the echo-cancelled ambient frequency domain signal and the ear canal frequency domain signal.

[0123] Before performing echo cancellation on the ambient frequency domain signal and the ear canal frequency domain signal based on the frequency domain amplitude corresponding to the speaker reference signal, embodiments of the present application can perform a fast Fourier transform on the speaker reference signal to obtain a frame-level spk frequency domain amplitude signal, that is, to obtain the frequency domain amplitude corresponding to the speaker reference signal. The specific calculation formula for the frame-level spk frequency domain amplitude signal is shown below.

[0124]

[0125] Wherein, SPK(m) represents the spectrum amplitude difference of the mth frame, and SPK(m, k) represents the kth frequency domain sampling point (frequency point) of the mth frame loudspeaker reference frequency domain amplitude signal.

[0126] On this basis, the embodiment of the present application can perform echo cancellation on the ambient frequency domain signal and the ear canal frequency domain signal based on the frequency domain amplitude corresponding to the speaker reference signal to obtain the ambient frequency domain signal and the ear canal frequency domain signal after echo cancellation. The formula for echo cancellation on the ambient frequency domain signal and the ear canal frequency domain signal can be shown as follows.

[0127] FF'(m,k)=AecFunction(FF(m,k),SPK(m,k));

[0128] FB'(m,k)=AecFunction(FB(m,k),SPK(m,k));

[0129] Wherein, FF′(m) represents the ambient frequency domain signal after the echo is removed, FB′(m) represents the ear canal frequency domain signal after the echo is removed, and AecFunction represents the echo cancellation process.

[0130] In this way, the embodiment of the present application can obtain the spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal after echo cancellation. The calculation formula of the spectrum amplitude difference is as follows.

[0131]

[0132] In summary, the method for obtaining the spectrum amplitude difference corresponding to when the earphone is playing content and when the earphone is not playing content is different, which can improve the accuracy of earphone control.

[0133] In step S24, if it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and the preset voice detection strategy, the power consumption mode of the headset is obtained.

[0134] In the embodiment of the present application, the power consumption mode of the headset may include a low power consumption mode and a high performance mode. When the power consumption mode of the headset is obtained, the embodiment of the present application may obtain the power consumption mode preset by the user.

[0135] Optionally, when the power consumption mode of the headset is obtained, embodiments of the present application may also obtain performance parameters of the headset and determine the corresponding power consumption mode based on the performance parameters of the headset. The headset performance parameters may include at least one of the remaining battery power, CPU (Central Processing Unit) load, etc.

[0136] As a specific implementation, when determining the corresponding power consumption mode based on the performance parameters of the headset, the embodiment of the present application can determine whether the remaining battery power of the headset is greater than a preset battery power. When it is determined that the remaining battery power of the headset is greater than the preset battery power, the power consumption mode of the headset is determined to be the high-performance mode. When it is determined that the remaining battery power of the headset is less than or equal to the preset battery power, the power consumption mode of the headset is determined to be the low-power mode.

[0137] As another specific implementation, when determining the corresponding power consumption mode based on the performance parameters of the headset, embodiments of the present application may determine whether the CPU load of the headset is less than a preset load. When it is determined that the CPU load of the headset is less than the preset load, the headset power consumption mode is determined to be the high-performance mode. When it is determined that the CPU load of the headset is greater than or equal to the preset load, the headset power consumption mode is determined to be the low-power mode.

[0138] As another specific embodiment, when obtaining the power consumption mode preset by the user, if high-performance detection indication information input by the user is received, the power consumption mode of the headset is determined to be high-performance mode. If low-power detection indication information input by the user is received, the power consumption mode of the headset is determined to be low-power mode. The high-performance detection indication information and the low-power detection indication information can be information input by the user by operating the headset, or can be indication information sent to the headset by the user through the terminal device by operating the terminal device.

[0139] As an optional method, after obtaining the power consumption mode of the headset, if the power consumption mode of the headset is determined to be high-performance mode, the embodiment of the present application can determine whether the user wearing the headset is performing voice activity based on the spectrum amplitude difference and a preset first voice detection strategy, and thus proceed to step S25. Alternatively, if the power consumption mode of the headset is determined to be low-power mode, the embodiment of the present application can determine whether the user wearing the headset is performing voice activity based on the spectrum amplitude difference and a preset second voice detection strategy, and thus proceed to step S26.

[0140] In step S25, if the power consumption mode is the high-performance mode, it is determined whether the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset first voice detection strategy.

[0141] As an optional approach, when the power consumption mode of the headset is determined to be high-performance mode, embodiments of the present application can determine whether the user wearing the headset is engaging in voice activity based on the spectrum amplitude difference and a preset first voice detection strategy. Specifically, embodiments of the present application can input the spectrum amplitude difference into a preset voice detection model to obtain a voice detection result.

[0142] The speech detection model may be obtained by performing machine learning training on an initial model using a training set, and the training set may include a positive sample data set and a negative sample data set. In addition, the positive sample data set may include multiple positive sample data, and each positive sample data may include first audio data and a first label corresponding to the first audio data, the first audio data including the user's speaking voice while wearing headphones, and the first label is used to indicate that the corresponding first audio data is target audio.

[0143] Optionally, the negative sample data set may include multiple negative sample data, each negative sample data may include second audio data and a second label corresponding to the second audio data, wherein the second audio data does not contain the speaking sound of the user wearing headphones, and the second label is used to indicate that the corresponding second audio data is not the target audio.

[0144] On this basis, the embodiment of the present application can determine whether the user wearing the headset has voice activity based on the voice detection result. In the embodiment of the present application, the voice detection model can include a Gaussian Mixture Model (GMM) or a Bayesian Gaussian Mixture Model (BGMM).

[0145] As an optional method, before inputting the spectrum amplitude difference into the speech detection model, the embodiment of the present application can first train the speech detection model, that is, use the training set to perform machine learning training on the initial model to obtain the speech detection model. Specifically, the embodiment of the present application can first create a positive and negative sample data set, that is, use a feedforward microphone, a feedback microphone, and a speaker to record the positive and negative sample data sets. The audio sampling frequency of the positive and negative sample data sets can be 16000.

[0146] As mentioned above, a positive sample dataset can be audio recorded while wearing headphones and speaking. This dataset can include recordings of at least 10 people, with a 50 / 50 split between male and female. Alternatively, the speech text can be the dialogue language of the scene. Furthermore, a negative sample dataset can be audio of non-headphone wearers speaking while wearing headphones, such as coughing, other people talking, and ambient sounds.

[0147] On this basis, the embodiment of the present application can extract features and construct training sets and test sets, use the recorded sample data set to calculate the frame-level spectrum amplitude difference features, and then perform data normalization on it. The processed data is used as the input of the initial model, and the initial model is trained using the training set to obtain a speech detection model. In addition, when constructing the positive and negative sample data sets, a corresponding label can be generated for each frame. The positive sample label can be 1, and the negative sample label can be 0. When training the speech detection model, the corresponding frame-level label can be used as another training input.

[0148] As an example, when the speech detection model is a Gaussian mixture model, the number of Gaussian distributions in the Gaussian mixture model can be 20, 30, or 40, and in the embodiment of the present application, the number of Gaussian distributions is preferably 30. The maximum number of iterations of the EM algorithm (Expectation-Maximum) in the Gaussian mixture model can be 150.

[0149] In the embodiment of the present application, the essence of using the speech detection model to obtain the speech detection results is to perform binary classification recognition, that is, the positive sample data and negative sample data for training the speech detection model are independent and identically distributed. These two types of data are used to train a GMM respectively, and eventually two corresponding GMMs will be trained.

[0150] As an optional method, when it is determined that the power consumption mode of the headset is a high-performance mode, the embodiment of the present application can input the spectrum amplitude difference into the speech detection model to obtain a speech detection result. Among them, the speech detection result may include a logarithmic maximum likelihood value. After receiving the spectrum amplitude difference, the speech detection model can obtain two logarithmic maximum likelihood values ​​through calculation. These two maximum likelihood values ​​are the likelihood values ​​corresponding to the positive sample and the negative sample respectively. When it is determined that the maximum likelihood value corresponding to the positive sample is greater than the maximum likelihood value corresponding to the negative sample, it indicates that the user wearing the headset has voice activity, and at this time the mode of the headset can be controlled to switch to the transparent mode.

[0151] In step S26, if the power consumption mode is the low power consumption mode, it is determined whether the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset second voice detection strategy.

[0152] Alternatively, if the power consumption mode of the headset is determined to be low power consumption mode, embodiments of the present application can determine whether the user wearing the headset is performing voice activity based on the spectrum amplitude difference and a preset second voice detection strategy. The spectrum amplitude difference may include multiple sub-amplitude differences.

[0153] Specifically, embodiments of the present application can obtain a perceptual weighting value, which can be the sum of the products of each sub-amplitude difference and the corresponding perceptual weighting coefficient. The perceptual weighting coefficient is the weighting value corresponding to different frequency bands. Based on this, whether the user wearing the headset is engaged in voice activity is determined based on the perceptual weighting value. The calculation formula for this perceptual weighting value can be as follows.

[0154]

[0155] Among them, fc_am(m) represents the perceptual weighted value, Indicates the perceptual weighting coefficient used to calculate the perceptual weighted value. The corresponding perceptual weighting coefficient may vary depending on the headset. The specific perceptual weighting coefficient is subject to actual conditions.

[0156] In this embodiment of the present application, the headphone user voice detection result of each frame can be represented by u_vad(m). If the value of u_vad(m) is 1, it indicates that the headphone user voice activity exists in this frame. If the value of u_vad(m) is 0, it indicates that the headphone user voice activity does not exist in this frame.

[0157] As an optional approach, when the perception weighted value fc_am(m) is greater than a preset threshold, the value of u_vad(m) can be set to 1; when the perception weighted value fc_am(m) is less than or equal to the preset threshold, the value of u_vad(m) can be set to 0. The preset threshold can be a threshold for determining headphone user voice activity per frame in low power mode.

[0158] In an embodiment of the present application, when determining whether a user wearing headphones has voice activity using low power mode or high performance mode, a specified number of consecutive frames of data may be counted to obtain a voice activity detection result. If it is determined that more than a specified percentage of results indicate the presence of headphone voice activity, it is determined that the user wearing the headphones has voice activity, and the headphones are controlled to switch to transparent mode, i.e., proceeding to step S27. As an example, the specified number may be 64 frames of data, which may be 1 second of data, and the specified percentage may be 80%. That is, when it is determined that more than 80% of the results indicate the presence of headphone voice activity, it is determined that the user wearing the headphones has voice activity.

[0159] The embodiment of the present application can realize intelligent non-removal of headphones through the spectral characteristics of the feedforward microphone and the feedback microphone. By setting low-power mode and high-performance mode for users to choose, the algorithm computing power requirement in low-power mode is low, which can effectively reduce the impact on the battery life of the headphones. In high-performance mode, the headphones can effectively reduce false detection of vocal cord vibration scenarios such as coughing through statistical learning, which can effectively improve the user experience.

[0160] In step S27, if it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and the preset voice detection strategy, the headset mode is controlled to switch to the transparent mode.

[0161] As described above, when it is determined based on the voice detection results that the user wearing the headphones has voice activity, the embodiment of the present application can control the mode of the headphones to switch to the transparent mode. During this process, the embodiment of the present application can obtain the type of voice activity and then determine whether the type of voice activity is a conversation type. If the type of voice activity is determined to be a conversation type, the embodiment of the present application can control the mode of the headphones to switch to the transparent mode.

[0162] As an optional method, obtaining the type of speech acquisition may include inputting the ear canal frequency domain signal into a semantic recognition model to obtain a semantic recognition result. The semantic recognition model may be obtained by machine learning training an initial semantic recognition model using a semantic training set, where the semantic training set includes a conversation sample dataset and a non-conversation sample dataset.

[0163] In an embodiment of the present application, a conversation sample data set may include a plurality of conversation sample data, each conversation sample data includes a third audio data and a third label corresponding to the third audio data, the third audio data includes the sound of the user wearing headphones having a conversation with other people, and the third label is used to indicate that the corresponding third audio data is the target conversation audio. In addition, a non-conversation sample data set includes a plurality of non-conversation sample data, each non-conversation sample data includes a fourth audio data and a fourth label corresponding to the fourth audio data, the fourth audio data includes the sound of the user wearing headphones not having a conversation with other people, and the fourth label is used to indicate that the corresponding fourth audio data is not the target conversation audio. On this basis, the type of speech activity corresponding to the semantic recognition result is obtained. Among them, the sound of the user wearing headphones not having a conversation with other people may include the sound of talking to oneself, or the sound of the user wearing headphones singing, etc.

[0164] Optionally, if it is determined based on the voice detection results that the user wearing the headset has voice activity, and if the voice activity is determined to be non-conversational, embodiments of the present application may maintain the headset mode unchanged. As an example, if it is determined that the user wearing the headset has voice activity, and if the voice activity is determined to be monologue, embodiments of the present application may maintain the headset mode unchanged.

[0165] As another example, when it is determined that the user wearing headphones has voice activity, if the voice activity is determined to be singing, the embodiment of the present application can keep the mode of the headphones unchanged, thus avoiding the mode switching affecting the user's normal use of the headphones.

[0166] It should be noted that after controlling the headset mode to switch to Transparency Mode, embodiments of the present application can continuously monitor whether the voice activity of the headset has ended. If the voice activity is determined to have ended, the present application can switch the headset from Transparency Mode back to Active Noise Cancellation Mode after the voice time reaches a specified duration. For example, 15 seconds after detecting the end of voice activity, embodiments of the present application can switch the headset mode back to Active Noise Cancellation Mode.

[0167] When the present application obtains the environmental frequency domain signal and the ear canal frequency domain signal, the present application can obtain the spectrum amplitude difference based on the environmental frequency domain signal and the ear canal frequency domain signal. On this basis, if it is determined that the user wearing the headphones has voice activity based on the spectrum amplitude difference and the preset voice detection strategy, the mode of the headphones is controlled to switch to the transparent mode. The present application determines whether the user wearing the headphones has voice activity by utilizing the spectrum amplitude difference and the preset voice detection strategy. This not only improves the flexibility of the headphone control, but also reduces the power consumption of the headphone control. The headphones can be controlled without user intervention, thereby improving the user experience. In addition, in the embodiment of the present application, the headphones can automatically switch from active noise reduction mode to transparent mode. During the conversation, the user does not need to take off the headphones to better communicate with others, thereby providing the user with a better intelligent experience.

[0168] Figure 5 FIG. 3 is a block diagram of a control device 300 for headphones according to an exemplary embodiment. Figure 5 The device includes a first acquisition module 301, a second acquisition module 302, a determination module 303 and a switching module 304.

[0169] The first acquisition module 301 is configured to acquire the environmental frequency domain signal, where the environmental frequency domain signal is a frequency domain expression of the sound signal in the environment surrounding the earphone;

[0170] The second acquisition module 302 is configured to acquire the ear canal frequency domain signal, where the ear canal frequency domain signal is a frequency domain expression of a sound signal in the ear canal of a user wearing the headset;

[0171] The determination module 303 is configured to obtain a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal;

[0172] The switching module 304 is configured to control the mode of the headset to switch to the transparent mode if it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset voice detection strategy.

[0173] In some implementations, the switching module 304 may include:

[0174] a mode acquisition submodule, configured to acquire a power consumption mode of the headset, where the power consumption mode includes a low power consumption mode and a high performance mode;

[0175] a first determining submodule configured to determine whether a user wearing the headset has voice activity based on the spectrum amplitude difference and a preset first voice detection strategy if the power consumption mode is a high-performance mode;

[0176] The second determining submodule is configured to determine whether the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset second voice detection strategy if the power consumption mode is the low power consumption mode.

[0177] In some embodiments, the mode acquisition submodule can also be configured to obtain a power consumption mode preset by the user, or obtain performance parameters of the headset, and determine the corresponding power consumption mode based on the performance parameters of the headset, wherein the performance parameters of the headset include at least one of the remaining battery power and the CPU load.

[0178] In some embodiments, the first determination submodule can also be configured to input the spectrum amplitude difference into a preset speech detection model to obtain a speech detection result, where the speech detection model is obtained by performing machine learning training on an initial model using a training set, wherein the training set includes a positive sample data set and a negative sample data set, the positive sample data set includes multiple positive sample data, each positive sample data includes first audio data and a first label corresponding to the first audio data, the first audio data includes the user's speaking voice when the user is wearing the headset, and the first label is used to indicate that the corresponding first audio data is the target audio, the negative sample data set includes multiple negative sample data, each negative sample data includes second audio data and a second label corresponding to the second audio data, the second audio data does not include the user's speaking voice when the user is wearing the headset, and the second label is used to indicate that the corresponding second audio data is not the target audio; based on the speech detection result, determine whether the user wearing the headset has voice activity.

[0179] In some embodiments, the spectrum amplitude difference includes multiple sub-amplitude differences, and the second determination submodule can also be configured to obtain a perceptual weighted value, which is the sum of the products of each sub-amplitude difference and the corresponding perceptual weighting coefficient, and the perceptual weighting coefficient is the weighted value corresponding to different frequency bands; based on the perceptual weighted value, it is determined whether the user wearing the headset has voice activity.

[0180] In some implementations, the switching module 304 may further include:

[0181] a type acquisition submodule configured to acquire a type of voice activity if it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset voice detection strategy;

[0182] a switching submodule, configured to control the mode of the headset to switch to a transparent mode if the type of the voice activity is a conversation type;

[0183] The holding module is configured to keep the mode of the headset unchanged if the voice activity is of a non-dialogue type.

[0184] In some implementations, the determination module 303 may include:

[0185] a reference signal acquisition submodule, configured to acquire a speaker reference signal, wherein the speaker reference signal is an original sound signal of an audio signal to be played by a speaker of the headset, and the speaker reference signal includes a plurality of audio frames;

[0186] an energy sum acquisition submodule, configured to determine the energy sum of each audio frame included in the loudspeaker reference signal;

[0187] The first spectrum amplitude difference acquisition submodule is configured to obtain a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal if the total energy of the audio frame is less than a preset energy.

[0188] In some implementations, the determining module 303 may further include:

[0189] an echo cancellation submodule, configured to, if the total energy of the audio frame is greater than or equal to a preset energy, perform echo cancellation on the ambient frequency domain signal and the ear canal frequency domain signal based on the frequency domain amplitude corresponding to the loudspeaker reference signal, to obtain the echo-cancelled ambient frequency domain signal and the ear canal frequency domain signal;

[0190] The second spectrum amplitude difference acquisition submodule is configured to obtain a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal after echo cancellation.

[0191] In some implementations, the first acquisition module 301 may include:

[0192] A first time domain signal acquisition submodule is configured to acquire an ambient time domain signal, where the ambient time domain signal refers to a time domain expression of a sound signal in an environment surrounding the headset;

[0193] A first time-frequency conversion module is configured to perform time-frequency conversion on the ambient time domain signal to obtain the ambient frequency domain signal;

[0194] The second acquisition module 302 may include:

[0195] a second time domain signal acquisition submodule configured to acquire an ear canal time domain signal, wherein the ear canal time domain signal refers to a time domain expression of a sound signal in the ear canal of a user wearing the headset;

[0196] The second time-frequency conversion module is configured to perform time-frequency conversion on the ear canal time domain signal to obtain the ear canal frequency domain signal.

[0197] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0198] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the earphone control method provided by the present disclosure are implemented.

[0199] Figure 6 FIG1 is a block diagram of a headset 800 for controlling a headset according to an exemplary embodiment. The headset 800 may include a speaker, a feedforward microphone, and a feedback microphone. In addition, the headset 800 may be a wireless headset or a wired headset, without limitation.

[0200] Reference Figure 6 The headset 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .

[0201] The processing component 802 generally controls the overall operation of the headset 800, such as operations associated with display, phone calls, data communications, camera operation, and recording. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the headset control method described above. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0202] The memory 804 is configured to store various types of data to support operations on the headset 800. Examples of such data include instructions for any application or method operating on the headset 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0203] The power supply assembly 806 provides power to the various components of the headset 800. The power supply assembly 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the headset 800.

[0204] The multimedia component 808 includes a screen that provides an output interface between the headset 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the headset 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0205] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the headset 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0206] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0207] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the headset 800. For example, the sensor assembly 814 can detect the open / closed state of the headset 800, the relative positioning of components, such as the display and keypad of the headset 800. The sensor assembly 814 can also detect changes in the position of the headset 800 or a component of the headset 800, the presence or absence of contact between the user and the headset 800, the orientation or acceleration / deceleration of the headset 800, and changes in the temperature of the headset 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0208] The communication component 816 is configured to facilitate wired or wireless communication between the headset 800 and other devices. The headset 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0209] In an exemplary embodiment, the headset 800 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-mentioned headset control method.

[0210] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions. The instructions can be executed by the processor 820 of the headset 800 to implement the headset control method described above. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0211] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned headset control method when executed by the programmable device.

Claims

1. A method for controlling headphones, characterized in that: include: Acquire an ambient frequency domain signal, where the ambient frequency domain signal is a frequency domain expression of a sound signal in an environment surrounding the headset; Acquire an ear canal frequency domain signal, where the ear canal frequency domain signal is a frequency domain expression of a sound signal in the ear canal of a user wearing the headset; Obtaining a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal, wherein the spectrum amplitude difference represents a difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal; If it is determined based on the spectrum amplitude difference and a preset voice detection strategy that the user wearing the headset has voice activity, controlling the mode of the headset to switch to a transparent mode; The method of determining whether the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset voice detection strategy includes: Before controlling the mode of the headset to switch to the transparency mode, obtaining a power consumption mode of the headset, where the power consumption mode includes a low power consumption mode and a high performance mode; If the power consumption mode is the low power consumption mode, determining whether the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset second voice detection strategy; The spectrum amplitude difference includes multiple sub-amplitude differences, and determining whether the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset second voice detection strategy includes: Obtaining a perceptual weighted value, where the perceptual weighted value is the sum of the products of each sub-amplitude difference and a corresponding perceptual weighting coefficient, where the perceptual weighting coefficient is a weighted value corresponding to different frequency bands; It is determined whether a user wearing the headset has voice activity based on the perception weight value.

2. The earphone control method according to claim 1, characterized in that: The determining, based on the spectrum amplitude difference and a preset voice detection strategy, that the user wearing the headset has voice activity includes: If the power consumption mode is a high-performance mode, determining whether the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset first voice detection strategy.

3. The earphone control method according to claim 2, characterized in that: The acquiring of the power consumption mode of the headset includes: Get the power consumption mode preset by the user, or Acquire performance parameters of the headset, and determine a corresponding power consumption mode according to the performance parameters of the headset, wherein the performance parameters of the headset include at least one of a remaining battery charge and a CPU load.

4. The earphone control method according to any one of claims 2 to 3, characterized in that: The determining, based on the spectrum amplitude difference and a preset first voice detection strategy, whether the user wearing the headset has voice activity includes: The spectral amplitude difference is input into a preset speech detection model to obtain a speech detection result, wherein the speech detection model is obtained by performing machine learning training on an initial model using a training set, wherein the training set includes a positive sample data set and a negative sample data set, the positive sample data set includes a plurality of positive sample data, each positive sample data includes first audio data and a first label corresponding to the first audio data, the first audio data includes a speaking voice of the user wearing the headset, and the first label is used to indicate that the corresponding first audio data is the target audio, the negative sample data set includes a plurality of negative sample data, each negative sample data includes second audio data and a second label corresponding to the second audio data, the second audio data does not include the speaking voice of the user wearing the headset, and the second label is used to indicate that the corresponding second audio data is not the target audio; Based on the voice detection result, it is determined whether a user wearing the headset has voice activity.

5. The earphone control method according to claim 1, characterized in that: If it is determined according to the spectrum amplitude difference and a preset voice detection strategy that the user wearing the headset has voice activity, controlling the mode of the headset to switch to the transparent mode includes: If it is determined that the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset voice detection strategy, obtaining the type of the voice activity; If the type of the voice activity is a conversation type, controlling the mode of the headset to switch to a transparent mode; If the voice activity is of a non-conversational type, the mode of the headset remains unchanged.

6. The earphone control method according to claim 1, characterized in that: The obtaining of a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal includes: Acquire a speaker reference signal, where the speaker reference signal is an original sound signal of an audio signal to be played by a speaker of the headset, and the speaker reference signal includes a plurality of audio frames; Determining the sum of energies of the audio frames included in the loudspeaker reference signal; If the total energy of the audio frame is less than the preset energy, the spectrum amplitude difference is obtained according to the ambient frequency domain signal and the ear canal frequency domain signal.

7. The earphone control method according to claim 6, characterized in that: The method further comprises: If the total energy of the audio frame is greater than or equal to a preset energy, performing echo cancellation on the ambient frequency domain signal and the ear canal frequency domain signal based on the frequency domain amplitude corresponding to the loudspeaker reference signal to obtain the ambient frequency domain signal and the ear canal frequency domain signal after echo cancellation; The spectrum amplitude difference is obtained according to the ambient frequency domain signal and the ear canal frequency domain signal after echo cancellation.

8. The earphone control method according to claim 1, characterized in that: The acquiring of the environmental frequency domain signal comprises: Acquire an ambient time domain signal, where the ambient time domain signal refers to a time domain expression of a sound signal in an environment surrounding the headset; Performing time-frequency conversion on the ambient time domain signal to obtain the ambient frequency domain signal; The obtaining of the ear canal frequency domain signal comprises: Acquiring an ear canal time domain signal, where the ear canal time domain signal refers to a time domain expression of a sound signal in the ear canal of a user wearing the headset; Performing time-frequency conversion on the ear canal time domain signal to obtain the ear canal frequency domain signal.

9. A headset control device, characterized in that: The device comprises: A first acquisition module is configured to acquire an environmental frequency domain signal, where the environmental frequency domain signal is a frequency domain expression of a sound signal in an environment surrounding the headset; a second acquisition module configured to acquire an ear canal frequency domain signal, where the ear canal frequency domain signal is a frequency domain expression of a sound signal in the ear canal of a user wearing the headset; a determination module configured to obtain a spectrum amplitude difference based on the ambient frequency domain signal and the ear canal frequency domain signal, wherein the spectrum amplitude difference represents a difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal; a switching module configured to control the mode of the headset to switch to a transparent mode if it is determined that the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset voice detection strategy; The switching module includes: a mode acquisition submodule, configured to acquire a power consumption mode of the headset, where the power consumption mode includes a low power consumption mode and a high performance mode; a second determining submodule configured to determine whether a user wearing the headset has voice activity based on the spectrum amplitude difference and a preset second voice detection strategy if the power consumption mode is the low power consumption mode; In which, the spectrum amplitude difference includes multiple sub-amplitude differences, and the second determination submodule is further configured to obtain a perceptual weighted value, which is the sum of the products of each sub-amplitude difference and the corresponding perceptual weighting coefficient, and the perceptual weighting coefficient is the weighted value corresponding to different frequency bands; based on the perceptual weighted value, determine whether the user wearing the headset has voice activity.

10. A headset, characterized in that: include: processor; a memory for storing processor-executable instructions; Feedforward microphone and feedback microphone; Wherein, the processor is configured to: Acquire an ambient frequency domain signal, where the ambient frequency domain signal is a frequency domain expression of a sound signal in an environment surrounding the headset; Acquire an ear canal frequency domain signal, where the ear canal frequency domain signal is a frequency domain expression of a sound signal in the ear canal of a user wearing the headset; Obtaining a spectrum amplitude difference according to the ambient frequency domain signal and the ear canal frequency domain signal, wherein the spectrum amplitude difference represents a difference between the amplitude of the ambient frequency domain signal and the amplitude of the ear canal frequency domain signal; If it is determined based on the spectrum amplitude difference and a preset voice detection strategy that the user wearing the headset has voice activity, controlling the mode of the headset to switch to a transparent mode; The method of determining whether the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset voice detection strategy includes: Before controlling the mode of the headset to switch to the transparency mode, obtaining a power consumption mode of the headset, where the power consumption mode includes a low power consumption mode and a high performance mode; If the power consumption mode is the low power consumption mode, determining whether the user wearing the headset has voice activity according to the spectrum amplitude difference and a preset second voice detection strategy; The spectrum amplitude difference includes multiple sub-amplitude differences, and determining whether the user wearing the headset has voice activity based on the spectrum amplitude difference and a preset second voice detection strategy includes: Obtaining a perceptual weighted value, where the perceptual weighted value is the sum of the products of each sub-amplitude difference and a corresponding perceptual weighting coefficient, where the perceptual weighting coefficient is a weighted value corresponding to different frequency bands; It is determined whether a user wearing the headset has voice activity based on the perception weight value.

11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Method and device for detecting voice of earphone wearer and storage medium

    CN111933140A

  • Earphone control method and device and earphone

    CN112770214A