Earphone sound leakage detection method, device, equipment and storage medium
By calculating the cross-correlation coefficient using feedback and feedforward microphones when audio is playing and not playing audio from the speaker, the accuracy problem of sound leakage detection in noise-canceling headphones is solved, achieving stable leakage detection and improving the listening experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAOMI TECH (WUHAN) CO LTD
- Filing Date
- 2023-03-13
- Publication Date
- 2026-07-21
AI Technical Summary
In existing technologies, noise-canceling headphones lack accuracy in detecting sound leakage, which affects the noise-canceling effect.
By acquiring audio signals using a feedback microphone and a feedforward microphone when audio is playing from the speaker and when no audio is playing, respectively, calculating the first and second cross-correlation coefficients, the sound leakage state of the headphones is determined, and different frequency bands are used to improve detection accuracy.
It improves the stability and accuracy of sound leakage detection in headphones under different scenarios, is suitable for low-power hardware resources, and improves the user's listening experience.
Smart Images

Figure CN116456237B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of audio data processing technology, and in particular to a method, apparatus, device, and storage medium for detecting sound leakage in headphones. Background Technology
[0002] Noise-canceling headphones are a type of headphone that reduces noise through active or passive noise cancellation. Active noise-canceling headphones, in particular, have noise-canceling circuitry that counteracts external noise. They mostly adopt a larger over-ear design, using earplugs and the headphone shell to block external noise for initial sound insulation, while also providing ample space for the active noise-canceling circuitry and power supply.
[0003] With the development of headphone technology and smart technology, active noise cancellation has become the preferred choice for manufacturers due to its ability to provide a quiet and comfortable listening experience in many scenarios. Active noise-canceling headphones first need to detect the current sound leakage, and then perform noise reduction according to the corresponding calculation strategy. Therefore, accurate detection of the sound leakage status of headphones is an important prerequisite for noise reduction. Summary of the Invention
[0004] This disclosure aims to at least partially address one of the technical problems in the related art.
[0005] The first aspect of this disclosure provides a method for detecting sound leakage in headphones, including:
[0006] When audio is played through the speaker of the headphones, the sound leakage state of the headphones is determined based on the first audio signal played through the speaker, the second audio signal collected synchronously with the feedback microphone of the headphones, and the first cross-correlation coefficient within the first frequency band.
[0007] When the speaker of the headphones is not playing audio, the sound leakage state of the headphones is determined by the third audio signal collected by the feedforward microphone of the headphones and the second audio signal collected by the feedback microphone of the headphones at the corresponding time, based on the second cross-correlation coefficient in the second frequency band.
[0008] The first frequency band range differs from the second frequency band range.
[0009] A second aspect of this disclosure provides a sound leakage detection device for headphones, comprising:
[0010] The first processing module is used to determine the sound leakage state of the headphones by using a first audio signal played by the speaker and a second audio signal collected synchronously with the feedback microphone of the headphones, based on a first cross-correlation coefficient within a first frequency band, when audio is played by the speaker of the headphones.
[0011] The second processing module is used to determine the sound leakage state of the headphones based on the third audio signal collected by the feedforward microphone of the headphones and the second audio signal collected synchronously by the feedback microphone of the headphones at the corresponding time, and the second cross-correlation coefficient in the second frequency band when the speaker of the headphones is not playing audio.
[0012] The first frequency band range differs from the second frequency band range.
[0013] A third aspect of this disclosure provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the sound leakage detection method for headphones as proposed in the first aspect of this disclosure.
[0014] The fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, implements the sound leakage detection method for headphones as proposed in the first aspect of this disclosure.
[0015] The headphone sound leakage detection method, apparatus, device, and storage medium disclosed herein have the following beneficial effects:
[0016] In this embodiment, when the headphone's speaker is playing audio, the sound leakage state of the headphone is determined by a first cross-correlation coefficient within a first frequency band, based on a first audio signal played by the speaker and a second audio signal synchronously collected by the headphone's feedback microphone at the corresponding moment. When the headphone's speaker is not playing audio, the sound leakage state of the headphone is determined by a second cross-correlation coefficient within a second frequency band, based on a third audio signal collected by the headphone's feedforward microphone and a second audio signal synchronously collected by the headphone's feedback microphone at the corresponding moment. This allows the headphone to determine a corresponding calculation strategy based on the current sound leakage state to improve the user's auditory experience. It is applicable to low-power hardware resources and ensures the stability of leakage detection in complex scenarios.
[0017] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description
[0018] The above and / or additional aspects and advantages of this disclosure will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, in which:
[0019] Figure 1 This is a schematic diagram of the operation of an earphone provided in an embodiment of the present disclosure;
[0020] Figure 2 This is a schematic flowchart illustrating a sound leakage detection method for headphones provided in an embodiment of this disclosure;
[0021] Figure 3 This is a schematic flowchart illustrating another method for detecting sound leakage in headphones provided in an embodiment of this disclosure;
[0022] Figure 4 A schematic diagram illustrating a threshold adjustment for leakage levels provided in an embodiment of this disclosure;
[0023] Figure 5 This is a structural block diagram of a sound leakage detection device for headphones provided in an embodiment of the present disclosure;
[0024] Figure 6 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0025] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0026] The following description, with reference to the accompanying drawings, outlines an embodiment of a sound leakage detection method, apparatus, computer device, and storage medium for headphones.
[0027] It should be noted that the execution subject of the headphone sound leakage detection method in this embodiment is the headphone sound leakage detection device, which can be implemented by software and / or hardware and can be configured in any electronic device. In the scenario proposed in this disclosure, the headphone can be used as the execution subject. The headphone sound leakage detection method proposed in this embodiment will be described below with "headphone" as the execution subject, and no limitation is made here.
[0028] More specifically, the sound leakage detection method for headphones provided in this disclosure can be applied to the use of TWS headphones (active noise-canceling headphones). To clearly illustrate the TWS headphone structure in this embodiment, as follows... Figure 1 This is a schematic diagram illustrating the operation of an earphone provided in an embodiment of this disclosure. Figure 1 The headphones contain a chip (i.e., a processor), a speaker, a feedforward microphone, and a feedback microphone. The feedforward microphone can detect ambient noise, while the feedback microphone typically receives sound from the vicinity of the speaker. The chip can generate corresponding anti-phase noise based on the current sound leakage to cancel out the noise, so that the human ear receives the noise-canceled sound.
[0029] Figure 2 This is a schematic flowchart of a sound leakage detection method for headphones provided in an embodiment of this disclosure.
[0030] like Figure 2 As shown, the sound leakage detection method for this headset may include the following steps:
[0031] Step 201: When the earphone's speaker is playing audio, determine the sound leakage state of the earphone based on the first audio signal played by the speaker, the second audio signal collected synchronously with the feedback microphone of the earphone, and the first cross-correlation coefficient within the first frequency band.
[0032] The acquisition time of the first audio signal and the second audio signal can be 2s, 1s, or 1.5s.
[0033] The preset sampling frequencies of the first and second audio signals can be predetermined sampling frequencies for audio signals, such as 48kHz, 44.1kHz, 24kHz, 16kHz, and 8kHz, and are not limited here. Optionally, in this disclosure, 16kHz can be selected as the preset sampling frequency, and is not limited here.
[0034] The first audio signal can be an audio signal obtained by collecting the sound emitted by the speaker, and the second audio signal is an audio signal obtained by the feedback microphone of the headphones. Furthermore, the first and second audio signals are collected synchronously, meaning the collection time for the first audio signal is the same as the collection time for the second audio signal.
[0035] For example, if the speaker of the current headphones plays 10 seconds of audio, with the initial playback time being T1 and the end playback time being T2, that is, the duration of T1 and T2 is 10 seconds, and if the feedback microphone collects the second audio signal from T1 to T3, and the duration of the second audio signal is 2 seconds, then the speaker can extract the audio from the time period T1 to T3, with a duration of 2 seconds, as the first audio signal. Therefore, at this time, the start and reception times of the first audio signal and the second audio signal are the same, and the duration of both is 2 seconds.
[0036] The current scene can be the current location of the headphones, such as a classroom, a car, an office, a park, an amusement park, etc., without any restrictions.
[0037] The first cross-correlation coefficient is the cross-correlation coefficient between the first audio signal and the second audio signal within a first frequency band. The cross-correlation coefficient, expressed using a cross-correlation function, represents the degree of correlation between two sequences, specifically describing the correlation between the values of the first and second audio signals.
[0038] The first frequency band is determined based on the audio frequency range of the audio played by the speaker. As one possible implementation, the first frequency band can be from 100Hz to 600Hz.
[0039] Optionally, the first audio signal and the second audio signal are synchronously acquired time-domain signals. Based on the first audio signal in the time domain, frame-by-frame overlapping and windowing are performed, as well as time-frequency transform (FFT) to obtain the frequency domain signals corresponding to each frame of the first audio signal. Similarly, the second audio signal in the time domain is framed, windowed, and subjected to time-frequency transform (FFT) to obtain the frequency domain signals corresponding to each frame of the second audio signal.
[0040] For example, let s(n) be the first audio signal in the time domain and f(n) be the second audio signal in the time domain. After obtaining s(n) and f(n), s(n) and f(n) can be divided into frames according to preset framing parameters. Here, n is the index of the sampling data in the time domain.
[0041] The preset framing parameters can include frame length and frame shift. The frame length can be 64ms, 32ms, 16ms, etc. Optionally, in this disclosure, 32ms can be selected as the frame length, and this is not limited. As an example, the frame shift can be half the frame length, i.e., 16ms, or it can be a quarter of the frame length, such as 8ms, and this is not limited.
[0042] Furthermore, when windowing the framed time-domain signal, it is necessary to window the time-domain signal for each frame. Specifically, windowing can be performed using the following formula:
[0043] s(k,n)=s((k-1)×inc+n)*w(n) (1)
[0044] f(k,n)=f((k-1)×inc+n)*w(n) (2)
[0045] In formula (1), s(k,n) is the sampled data with index n after windowing the first audio signal in the time domain of the kth frame, f(k,n) is the sampled data with index n after windowing the second audio signal in the time domain of the kth frame, inc is the frame shift, n (1≤n≤L) is the index of the sampled data in the frame, and L is the total number of samples in the time domain signal of each frame.
[0046] Where w(n) is a window function, such as a Hanning window or a Hamming window, which is not limited here. In this disclosure, a Hanning window can be selected as an example.
[0047] Furthermore, after frame-by-frame overlapping and windowing, a Fourier transform (FFT) operation can be performed on each frame of data to obtain the first frequency domain signal S(k,m) corresponding to the first audio signal of each frame, and the second frequency domain signal F(k,m) corresponding to the second audio signal of each frame. Here, k represents the k-th frame, and m is the index of the frequency point.
[0048] As one possible implementation, frequency points are filtered in the first frequency domain signal S(k,m) and the second frequency domain signal F(k,m), and only frequency points m' within the first frequency band are selected. The cross-correlation coefficient is then calculated in the frequency domain to obtain the first cross-correlation coefficient.
[0049] As another possible implementation, frequency points are filtered for the first frequency domain signal S(k,m) and the second frequency domain signal F(k,m), and only frequency points m' within the first frequency band are selected. Then, the Fourier inverse transform is performed to the time domain, and the cross-correlation coefficient is calculated to obtain the first cross-correlation coefficient.
[0050] It should be noted that the above example is only an illustrative example and is not intended to be limiting.
[0051] There is a mapping relationship between the first cross-correlation coefficient and the sound leakage state of the headphones. The sound leakage state of the headphones can be determined based on the first cross-correlation coefficient. This is because, when audio is played from a speaker, it is desirable for more audio to enter the ear canal and less to leak out, while simultaneously, less ambient noise should enter the ear canal. Figure 1 As shown, the microphone inside the ear canal is a feedback microphone. This means that the feedback microphone is expected to capture more audio from the speaker and less other noise. In signal processing, this translates to a higher degree of cross-correlation between the second audio signal from the feedback microphone and the first audio signal from the speaker; a larger first cross-correlation coefficient indicates less sound leakage from the headphones.
[0052] Conversely, when more audio from the speaker leaks out of the ear canal, less audio enters the ear canal. The feedback microphone inside the ear canal will then pick up less of the audio from the speaker and more of other noise. In signal processing, this translates to a low cross-correlation between the second audio signal from the feedback microphone and the first audio signal from the speaker (i.e., a low first cross-correlation coefficient), indicating significant sound leakage from the headphones.
[0053] In this embodiment of the disclosure, to ensure high differentiation between different sound leakage states of the headphones, the calculation of the first cross-correlation coefficient is specifically limited to a first frequency domain range of 100Hz to 600Hz, rather than the entire frequency domain. This is because the audio played by the speaker mainly consists of voice and music content, which is primarily concentrated in the 100Hz to 600Hz range. In scenarios where the speaker plays audio, the audio collected by the feedback microphone is mainly the audio played by the speaker. By calculating the first cross-correlation coefficient within the 100Hz to 600Hz range, the magnitude of the audio component played by the speaker in the audio collected by the feedback microphone can be reflected. Determining the sound leakage state based on the second cross-correlation coefficient within the 100Hz to 600Hz range has higher sensitivity and accuracy.
[0054] Step 202: When the speaker of the headphones is not playing audio, the sound leakage state of the headphones is determined based on the third audio signal collected by the feedforward microphone of the headphones and the second audio signal collected synchronously by the feedback microphone of the headphones at the corresponding time, and the second cross-correlation coefficient within the second frequency band.
[0055] The second cross-correlation coefficient is the cross-correlation coefficient between the third audio signal and the second audio signal within the second frequency band. The cross-correlation coefficient, expressed using a cross-correlation function, represents the degree of correlation between the two sequences, specifically describing the correlation between the values of the third audio signal and the second audio signal.
[0056] The first and second frequency bands differ because the primary sources of the audio collected by the feedback microphone differ between scenarios where audio is playing from the headphone speaker and when no audio is playing. The corresponding frequency bands for these audio sources also differ. For example, when audio is playing from the headphone speaker, the primary source is the speaker itself; while when no audio is playing, the primary source is ambient noise. This embodiment considers the impact of this difference on the discriminative power of determining the headphone's sound leakage state, setting different frequency bands for these two scenarios, which improves the accuracy of determining the headphone's sound leakage state. Furthermore, since calculations are only required for a portion of the frequency bands, the computational load is reduced to some extent.
[0057] The second frequency band is determined based on the noise audio frequency domain range of the ambient noise. As one possible implementation, the second frequency band could be from 1000Hz to 2000Hz.
[0058] The process of acquiring the third and second audio signals, as well as the process of calculating the second cross-correlation coefficient, can be found in the relevant descriptions in the preceding steps, and will not be repeated in this embodiment.
[0059] There is a mapping relationship between the second cross-correlation coefficient and the sound leakage state of the headphones. The sound leakage state of the headphones can be determined based on the second cross-correlation coefficient because less ambient noise is expected to enter the ear canal when the speaker is not playing audio. Figure 1 As shown, the microphone inside the ear canal is a feedback microphone, meaning that it is desirable for the feedback microphone to pick up less ambient noise. In signal processing, this translates to a low degree of cross-correlation between the second audio signal from the feedback microphone and the third audio signal from the front microphone outside the ear canal; a smaller second cross-correlation coefficient indicates less sound leakage from the headphones.
[0060] Conversely, when there is a lot of ambient noise entering the ear canal, the feedback microphone inside the ear canal will pick up more of the ambient noise. In signal processing, this means there is a high degree of cross-correlation between the second audio signal from the feedback microphone and the third audio signal from the front microphone; that is, a large second cross-correlation coefficient indicates significant sound leakage from the headphones.
[0061] In this embodiment of the disclosure, to ensure high differentiation between different sound leakage states of the headphones, the calculation of the second cross-correlation coefficient is specifically limited to a second frequency domain range of 1000Hz to 2000Hz, rather than the entire frequency domain. This is because ambient noise is mainly concentrated in the 1000Hz to 2000Hz range. In scenarios where the speaker is not playing audio, the audio collected by the feedback microphone is mainly ambient noise. By calculating the second cross-correlation coefficient within the 1000Hz to 2000Hz range, the magnitude of the ambient noise component in the audio collected by the feedback microphone can be reflected. Determining the sound leakage state based on the second cross-correlation coefficient within the 1000Hz to 2000Hz range has higher sensitivity and accuracy.
[0062] In this embodiment, when the headphone's speaker is playing audio, the sound leakage state of the headphone is determined by a first cross-correlation coefficient within a first frequency band, based on a first audio signal played by the speaker and a second audio signal synchronously collected by the headphone's feedback microphone. When the headphone's speaker is not playing audio, the sound leakage state of the headphone is determined by a second cross-correlation coefficient within a second frequency band, based on a third audio signal collected by the headphone's feedforward microphone and a second audio signal synchronously collected by the headphone's feedback microphone at the corresponding moment. This facilitates the headphone in determining a corresponding calculation strategy to improve the user's auditory experience based on the current sound leakage state, is applicable to low-power hardware resources, and ensures the stability of leakage detection in complex scenarios.
[0063] Figure 3 This is a schematic flowchart illustrating another method for detecting sound leakage in headphones provided in an embodiment of this disclosure.
[0064] like Figure 3 As shown, the sound leakage detection method for this headset may include the following steps:
[0065] Step 301: Determine whether the speaker is playing audio.
[0066] Step 302: When the speaker is playing audio, determine the first cross-correlation coefficient within the first frequency band based on the amplitude spectrum of the first audio signal played by the speaker and the amplitude spectrum of the second audio signal collected at the synchronization time of the feedback microphone of the headphones.
[0067] The amplitude spectrum is obtained by performing a Fourier transform on each frame of the time-domain signal, referring to the method provided in the aforementioned embodiments. The amplitude spectrum can be obtained by calculating the geometric mean of the real part (indicating amplitude) and imaginary part (indicating phase) of each frequency point in the spectrum, or the real part of each frequency point can be directly taken as the amplitude spectrum. That is, the horizontal axis of the amplitude spectrum is similar to that of the spectrum, representing the frequency of each frequency point, while the vertical axis represents the amplitude.
[0068] Optionally, referring to the foregoing embodiments, the first audio signal in the time domain is subjected to frame-by-frame overlapping windowing, time-frequency domain transform (FFT), and frequency points within the first frequency band are selected. The amplitude spectrum corresponding to any frame of the frequency domain signal is denoted as A1(k,m'). Similarly, the second audio signal in the time domain is subjected to frame-by-frame overlapping windowing, time-frequency domain transform (FFT), and frequency points within the first frequency band are selected. The amplitude spectrum corresponding to any frame of the frequency domain signal is denoted as A2(k,m'). A first cross-correlation coefficient is obtained by cross-correlation calculation of A1(k,m') and A2(k,m').
[0069] Step 303: Based on the first cross-correlation number within the first frequency band, query the first mapping table to determine the leakage level corresponding to the first cross-correlation number.
[0070] The first cross-correlation coefficient and the leakage level are inversely related, and the leakage level is used to indicate the sound leakage status of the headphones.
[0071] The first mapping table is used to indicate the threshold values of the first cross-correlation coefficients corresponding to each leakage level. Optionally, each leakage level corresponds to a first cross-correlation coefficient interval, and each interval includes two threshold values: an upper threshold value and a lower threshold value.
[0072] As one possible implementation, the leakage level of the current headphones can be calculated directly based on the first cross-correlation coefficient obtained at the moment.
[0073] For example, if the first cross-correlation coefficient is divided into three intervals from smallest to largest, namely D1, D2, and D3, the corresponding leakage levels are level 3, level 2, and level 1 respectively, according to the leakage from most to least. If the first cross-correlation coefficient is in the D2 interval, it means that the leakage level of the current earphone is level 2.
[0074] As another possible approach, to avoid abrupt changes in the leakage level due to measurement errors, the boundary thresholds of each interval of the first cross-correlation coefficient in the first mapping table can be dynamically adjusted.
[0075] In some scenarios, when the current leakage level is N, and the first cross-correlation coefficient is determined based on the first mapping table to gradually approach the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold between the leakage level N and the leakage level N+1 is reduced by a preset adjustment amount, where N is a positive number.
[0076] In other scenarios, when the current leakage level is N+1, and the first cross-correlation coefficient is determined based on the first mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold between the leakage level N and the leakage level N+1 is increased by the preset adjustment amount.
[0077] Among them, the preset adjustment amount (margin) n The threshold adjustment amount is set to a preset threshold value. Based on the adjusted threshold value, the corresponding leakage level is determined.
[0078] Step 304: When the speaker is not playing audio, determine the second cross-correlation coefficient in the second frequency band based on the amplitude spectrum of the third audio signal collected by the feedforward microphone of the earphone and the amplitude spectrum of the second audio signal collected at the corresponding moment of the feedback microphone of the earphone.
[0079] The amplitude spectrum and the second cross-correlation coefficient can be referred to in the foregoing steps and the relevant descriptions in the embodiments, which will not be repeated in this embodiment.
[0080] Step 305: Based on the second cross-correlation coefficient within the second frequency band, query the second mapping table to determine the leakage level corresponding to the second cross-correlation coefficient.
[0081] The second cross-correlation coefficient has a positive relationship with the leakage level, and the leakage level is used to indicate the sound leakage status of the headphones.
[0082] The second mapping table contains pre-stored threshold values for the second cross-correlation coefficients corresponding to each leakage level. Optionally, each leakage level corresponds to a second cross-correlation coefficient interval, and each interval includes two threshold values: an upper threshold value and a lower threshold value.
[0083] One possible approach is to directly calculate the leakage level of the current headphones based on the currently calculated second cross-correlation coefficient.
[0084] For example, if the second cross-correlation coefficient is divided into three intervals from smallest to largest, namely D1, D2, and D3, the corresponding leakage levels are level 1, level 2, and level 3 respectively, from least to most leakage. If the second cross-correlation coefficient is in the D2 interval, it means that the current leakage level of the headphones is level 2.
[0085] As another possible approach, to avoid abrupt changes in the leakage level due to measurement errors, the boundary thresholds of each interval of the second cross-correlation coefficient in the second mapping table can be dynamically adjusted.
[0086] In some scenarios, when the current leakage level is N, and the second cross-correlation coefficient is determined based on the second mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold between the leakage level N and the leakage level N+1 is increased by a preset adjustment amount, where N is a positive number.
[0087] In other scenarios, when the current leakage level is N+1, and the second cross-correlation coefficient is determined based on the second mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold between the leakage level N and the leakage level N+1 is reduced by the preset adjustment amount.
[0088] Among them, the preset adjustment amount (margin) n The threshold adjustment amount is set to a preset threshold value. Based on the adjusted threshold value, the corresponding leakage level is determined.
[0089] like Figure 4 As shown, Figure 4 It includes three leakage levels: n, n-1, and n+1, where the boundary between n and n+1 is defined as TH. n In actual detection scenarios, if the leakage level is n, and the leakage characteristic value gradually approaches TH... n At this point, the interval boundary (i.e., the threshold between leakage level N and leakage level N+1) can be adjusted to TH. n +margin n Conversely, if the leakage level is n+1, and the second cross-correlation coefficient gradually approaches TH...n At this time, the limit threshold TH can be set. n Adjusted to TH n -margin n Among them, margin n This refers to the preset adjustment amount mentioned above. This significantly increases the stability and robustness of the algorithm, thereby improving the user's listening experience and effectively ensuring the stability of leak detection in complex scenarios.
[0090] Furthermore, after determining the leakage level, a noise reduction strategy corresponding to that leakage level can be implemented.
[0091] The noise reduction strategy corresponds to the leakage level, with different leakage levels requiring different noise reduction strategies. The headphones can adaptively generate the inverse phase noise to be generated based on the noise reduction strategy corresponding to the leakage level, thereby compensating for leakage and ensuring a consistent listening experience for users in various scenarios.
[0092] In this embodiment, when the headphone's speaker is playing audio, the sound leakage state of the headphone is determined by a first cross-correlation coefficient within a first frequency band, based on a first audio signal played by the speaker and a second audio signal synchronously collected by the headphone's feedback microphone. When the headphone's speaker is not playing audio, the sound leakage state of the headphone is determined by a second cross-correlation coefficient within a second frequency band, based on a third audio signal collected by the headphone's feedforward microphone and a second audio signal synchronously collected by the headphone's feedback microphone at the corresponding moment. This facilitates the headphone in determining a corresponding calculation strategy to improve the user's auditory experience based on the current sound leakage state, is applicable to low-power hardware resources, and ensures the stability of leakage detection in complex scenarios.
[0093] To achieve the above embodiments, this disclosure also proposes a sound leakage detection device for headphones.
[0094] Figure 5 This is a structural block diagram of the headphone sound leakage detection device provided in the fourth embodiment of this disclosure.
[0095] like Figure 5 As shown, the sound leakage detection device 500 for the headphones may include:
[0096] The first processing module 501 is used to determine the sound leakage state of the headphones by using a first audio signal played by the speaker and a second audio signal collected synchronously with the feedback microphone of the headphones, based on a first cross-correlation coefficient within a first frequency band, when audio is played by the speaker of the headphones.
[0097] The second processing module 502 is used to determine the sound leakage state of the headphones by, based on the third audio signal collected by the feedforward microphone of the headphones and the second audio signal collected at the corresponding moment of the feedback microphone of the headphones, and the second cross-correlation coefficient within the second frequency band, when the speaker of the headphones is not playing audio.
[0098] The first frequency band range differs from the second frequency band range.
[0099] In some possible implementations of the embodiments of this disclosure, the upper frequency limit of the first frequency band is less than the lower frequency limit of the second frequency band.
[0100] In some possible implementations of the embodiments of this disclosure, the first frequency band range is determined based on the audio frequency domain range of the audio played by the speaker; and / or, the second frequency band range is determined based on the noise audio range.
[0101] In some possible implementations of the embodiments of this disclosure, the first processing module 501 is configured to: determine a first cross-correlation coefficient within a first frequency band based on the amplitude spectrum of the first audio signal played by the speaker and the amplitude spectrum of the second audio signal collected synchronously with the feedback microphone of the earphone; and query a first mapping table based on the first cross-correlation coefficient within the first frequency band to determine the leakage level corresponding to the first cross-correlation coefficient; wherein the first cross-correlation coefficient and the leakage level are inversely related, and the leakage level is used to indicate the sound leakage state of the earphone.
[0102] Furthermore, the first processing module 501 is also configured to: reduce the boundary threshold by a preset adjustment amount when the current leakage level is N and the first cross-correlation coefficient is determined to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1 based on the first mapping table, where N is a positive number; or, increase the boundary threshold by the preset adjustment amount when the current leakage level is N+1 and the first cross-correlation coefficient is determined to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1 based on the first mapping table.
[0103] In some possible implementations of the embodiments of this disclosure, the second processing module 502 is used to: determine a second cross-correlation coefficient within a second frequency band based on the amplitude spectrum of the third audio signal collected by the feedforward microphone of the earphone and the amplitude spectrum of the second audio signal collected synchronously by the feedback microphone of the earphone at the corresponding time; and query a second mapping table based on the second cross-correlation coefficient within the second frequency band to determine the leakage level corresponding to the second cross-correlation coefficient; wherein the second cross-correlation coefficient and the leakage level have a positive relationship, and the leakage level is used to indicate the sound leakage state of the earphone.
[0104] Furthermore, the second processing module 502 is also configured to: when the current leakage level is N, and the second cross-correlation coefficient is determined based on the second mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, increase the boundary threshold by a preset adjustment amount, where N is a positive number; or, when the current leakage level is N+1, and the second cross-correlation coefficient is determined based on the second mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, decrease the boundary threshold by the preset adjustment amount.
[0105] To implement the above embodiments, this disclosure also proposes a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the headphone sound leakage detection method proposed in the foregoing embodiments of this disclosure.
[0106] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing a computer program, which, when executed by a processor, implements the headphone sound leakage detection method as proposed in the foregoing embodiments of this disclosure.
[0107] To implement the above embodiments, this disclosure also proposes a computer program product that, when the instruction processor in the computer program product is executed, performs the headphone sound leakage detection method as proposed in the foregoing embodiments of this disclosure.
[0108] Figure 6 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Figure 6 The computer device 12 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0109] like Figure 6As shown, the computer device 12 is represented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0110] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0111] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0112] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 30 and / or cache memory 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 6 Not shown; usually referred to as a "hard drive"). Although Figure 6Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disc drive for reading and writing to a removable non-volatile optical disc (e.g., a compact disc read-only memory (CD-ROM), a digital video disc read-only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of this disclosure.
[0113] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this disclosure.
[0114] Computer device 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), and with one or more devices that enable a user to interact with computer device 12, and / or with any device that enables computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via input / output (I / O) interface 22. Furthermore, computer device 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of computer device 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0115] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.
[0116] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0117] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0118] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.
[0119] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0120] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0121] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0122] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0123] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A method for detecting sound leakage in headphones, characterized in that, include: When audio is played through the speaker of the headphones, the sound leakage state of the headphones is determined based on the first audio signal played through the speaker, the second audio signal collected synchronously with the feedback microphone of the headphones, and the first cross-correlation coefficient within the first frequency band. When the speaker of the headphones is not playing audio, the sound leakage state of the headphones is determined by the third audio signal collected by the feedforward microphone of the headphones and the second audio signal collected by the feedback microphone of the headphones at the corresponding time, based on the second cross-correlation coefficient in the second frequency band. Wherein, the first frequency band range differs from the second frequency band range; the first cross-correlation coefficient is determined within the first frequency band range based on the amplitude spectrum of the first audio signal and the amplitude spectrum of the second audio signal; the second cross-correlation coefficient is determined within the second frequency band range based on the amplitude spectrum of the third audio signal and the amplitude spectrum of the second audio signal; the first cross-correlation coefficient and the second cross-correlation coefficient represent the degree of correlation between the two audio signals using a cross-correlation function; The step of determining the sound leakage state of the headphones based on the first audio signal played by the speaker, the second audio signal collected synchronously with the feedback microphone of the headphones, and the first cross-correlation coefficient within a first frequency band includes: Based on the first cross-correlation number within the first frequency band, a first mapping table is queried to determine the leakage level corresponding to the first cross-correlation number; wherein, the first cross-correlation number and the leakage level are inversely related, and the leakage level is used to indicate the sound leakage status of the headphones; The first mapping table is used to indicate the threshold value of the first cross-correlation coefficient corresponding to each leakage level; the method further includes: When the current leakage level is N, and the first cross-correlation number is determined based on the first mapping table to gradually approach the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold is reduced by a preset adjustment amount, where N is a positive number; or, If the current leakage level is N+1, and the first cross-correlation coefficient is determined based on the first mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold is increased by the preset adjustment amount.
2. The method according to claim 1, characterized in that, The upper frequency limit of the first frequency band is less than the lower frequency limit of the second frequency band.
3. The method according to claim 1, characterized in that, The first frequency band range is determined based on the audio frequency range of the audio played by the speaker; and / or, The second frequency band range is determined based on the noise audio range.
4. The method according to any one of claims 1-3, characterized in that, The step of determining the sound leakage state of the headphones based on the third audio signal collected by the feedforward microphone of the headphones and the second audio signal collected synchronously by the feedback microphone of the headphones at the corresponding time, using a second cross-correlation coefficient within a second frequency band, includes: Based on the second cross-correlation number within the second frequency band, a second mapping table is queried to determine the leakage level corresponding to the second cross-correlation number; wherein, the second cross-correlation number and the leakage level have a positive relationship, and the leakage level is used to indicate the sound leakage status of the headphones.
5. The method according to claim 4, characterized in that, The second mapping table is used to indicate the threshold values of the second cross-correlation coefficient corresponding to each leakage level; the method further includes: When the current leakage level is N, and the second cross-correlation coefficient is determined based on the second mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold is increased by a preset adjustment amount, where N is a positive number; or, If the current leakage level is N+1, and the second cross-correlation coefficient is determined based on the second mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold is reduced by the preset adjustment amount.
6. A sound leakage detection device for headphones, characterized in that, include: The first processing module is used to determine the sound leakage state of the headphones by using a first audio signal played by the speaker and a second audio signal collected synchronously with the feedback microphone of the headphones, based on a first cross-correlation coefficient within a first frequency band, when audio is played by the speaker of the headphones. The second processing module is used to determine the sound leakage state of the headphones based on the third audio signal collected by the feedforward microphone of the headphones and the second audio signal collected synchronously by the feedback microphone of the headphones at the corresponding time, and the second cross-correlation coefficient in the second frequency band when the speaker of the headphones is not playing audio. Wherein, the first frequency band range differs from the second frequency band range; the first cross-correlation coefficient is determined within the first frequency band range based on the amplitude spectrum of the first audio signal and the amplitude spectrum of the second audio signal; the second cross-correlation coefficient is determined within the second frequency band range based on the amplitude spectrum of the third audio signal and the amplitude spectrum of the second audio signal; the first cross-correlation coefficient and the second cross-correlation coefficient represent the degree of correlation between the two audio signals using a cross-correlation function; The first processing module is further configured to query a first mapping table based on a first cross-correlation number within the first frequency band to determine the leakage level corresponding to the first cross-correlation number; wherein the first cross-correlation number and the leakage level are inversely related, and the leakage level is used to indicate the sound leakage status of the headphones; The first mapping table is used to indicate the threshold values of the first cross-correlation coefficient corresponding to each leakage level; it also includes: When the current leakage level is N, and the first cross-correlation number is determined based on the first mapping table to gradually approach the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold is reduced by a preset adjustment amount, where N is a positive number; or, If the current leakage level is N+1, and the first cross-correlation coefficient is determined based on the first mapping table to be gradually approaching the boundary threshold between the leakage level N and the leakage level N+1, the boundary threshold is increased by the preset adjustment amount.
7. An electronic device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the sound leakage detection method for headphones as described in any one of claims 1-5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the headphone sound leakage detection method as described in any one of claims 1-5.