TWS earphones and their playback methods and devices

By using a combination of crossover and multiple speakers in TWS earphones, the impact of ANC function on sound quality and call clarity is resolved, achieving a balance between high sound quality and broadband HD calls.

CN115250397BActive Publication Date: 2025-11-14HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110467311.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-28
Publication Date
2025-11-14
Estimated Expiration
2041-04-28

AI Technical Summary

Technical Problem

Existing TWS earphones, when implementing ANC (Active Noise Cancellation), will cancel some music or call sound, affecting sound quality and call clarity, and cannot balance high-quality music and broadband HD calls.

Method used

The design employs at least two speakers, using a crossover to divide the audio signal into multiple frequency bands of sub-audio signals, which are then played by each speaker. Combined with SP filters and feedback microphone processing, this ensures that the speakers maintain high sound quality across all frequency bands and support ultra-wideband voice calls.

Benefits of technology

It achieves high sound quality across all frequency bands of the audio source while supporting ultra-wideband voice calls, avoiding the noise cancellation of music and calls caused by ANC, thus improving sound quality and call clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115250397B_ABST
    Figure CN115250397B_ABST
Patent Text Reader

Abstract

This application provides a TWS earphone and a method and apparatus for playing TWS earphones. The TWS earphone includes: an audio signal processing path, a crossover, and at least two speakers; wherein the output terminal of the audio signal processing path is connected to the input terminal of the crossover; the output terminal of the crossover is connected to the at least two speakers; the audio signal processing path is configured to output a speaker driving signal; the crossover is configured to divide the speaker driving signal into at least two frequency bands of sub-audio signals, the at least two frequency bands corresponding to the main operating frequency bands of the at least two speakers; adjacent frequency bands in the at least two frequency bands partially overlap, or adjacent frequency bands in the at least two frequency bands do not overlap; the at least two speakers are configured to play the corresponding sub-audio signals. This application can achieve high sound quality across all frequency bands of the audio source and support ultra-wideband voice calls.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to audio processing technology, and more particularly to a true wireless stereo (TWS) headset and a method and apparatus for playing TWS headsets. Background Technology

[0002] In users' headphone usage scenarios, the common requirement is to achieve stable active noise cancellation (ANC) or hear-through (HT) functions while providing high-quality music or broadband HD calls.

[0003] While current TWS earbuds offer ANC (Active Noise Cancellation) functionality, they also cancel out some music or call audio, affecting sound quality and call clarity. Summary of the Invention

[0004] This application provides a TWS earphone and a method and apparatus for playing TWS earphones, which can not only deliver high sound quality across all frequency bands of the audio source, but also support ultra-wideband voice calls.

[0005] In a first aspect, this application provides a TWS earphone, comprising: an audio signal processing path, a crossover, and at least two speakers; wherein the output terminal of the audio signal processing path is connected to the input terminal of the crossover; the output terminal of the crossover is connected to the at least two speakers; the audio signal processing path is configured to output a speaker driving signal after performing noise reduction or pass-through processing on an audio source; the audio source is original music or call voice; or, the audio source includes a voice signal processed by human voice enhancement and the original music or call voice; the crossover is configured to divide the speaker driving signal into at least two frequency bands of sub-audio signals, the at least two frequency bands corresponding to the main operating frequency bands of the at least two speakers; adjacent frequency bands in the at least two frequency bands partially overlap, or adjacent frequency bands in the at least two frequency bands do not overlap; the at least two speakers are configured to play the corresponding sub-audio signals.

[0006] The frequency band of the processed speaker drive signal corresponds to the frequency band of the audio source and may include the entire low, mid and high frequency bands. However, since the main operating frequency band of a single speaker may only cover a part of the low, mid and high frequency bands, the single speaker cannot reproduce high sound quality across the entire frequency band.

[0007] This application controls the crossover to divide the speaker drive signal according to a preset method by setting the parameters of the crossover. The crossover can be configured to divide the speaker drive signal based on the main operating frequency bands of at least two speakers, obtaining sub-audio signals of at least two frequency bands corresponding to the main operating frequency bands of the at least two speakers, and then each speaker plays the corresponding sub-audio signal, thereby allowing the speaker to maintain optimal frequency response when playing the sub-audio signals transmitted to it. The crossover is controlled to divide the speaker drive signal according to a preset method by setting the parameters of the crossover.

[0008] The TWS earphones of this application are equipped with at least two speakers, and the main operating frequency bands of the at least two speakers are not exactly the same. A crossover can divide the speaker driving signal into sub-audio signals of at least two frequency bands. Adjacent frequency bands of the at least two frequency bands may partially overlap or not overlap. In this way, each sub-audio signal is transmitted to a frequency band-matched speaker. The aforementioned frequency band matching may mean that the main operating frequency band of the speaker covers the frequency band of the sub-audio signal transmitted to it. In this way, the speaker maintains the optimal frequency response when playing the transmitted sub-audio signal, which can not only reflect high sound quality in all frequency bands of the audio source, but also support ultra-wideband voice calls.

[0009] In one possible implementation, the audio signal processing path includes a secondary path SP filter configured to prevent the noise reduction or pass-through processing from eliminating sound from the audio source when the noise reduction or pass-through processing is concurrent with the audio source.

[0010] In one possible implementation, the audio signal processing path further includes: a feedback FB microphone and a feedback filter; wherein the FB microphone is configured to pick up an ear canal signal, the ear canal signal including residual noise signals inside the ear canal and the music or voice call; the SP filter is configured to input the audio source, process the audio source, and output a signal superimposed on the ear canal signal before transmitting it to the feedback filter; the feedback filter is configured to generate a signal for the noise reduction or pass-through processing, the noise reduction or pass-through processing signal being one of the superimposed signals used to generate the speaker drive signal.

[0011] In one possible implementation, the feedback FB microphone, the feedback filter, and the SP filter are set in the codec.

[0012] In one possible implementation, the SP filter is configured to input the audio source, process the audio source, and output a signal that is a superimposed signal of one of the speaker drive signals.

[0013] In one possible implementation, the SP filter is located in a digital signal processing (DSP) chip.

[0014] This application determines the SP filter (including a fixed SP filter or an adaptive SP filter) through the above method, which can achieve noise reduction or pass-through functions during music playback or transmission, and can also prevent the pass-through or noise reduction technology from eliminating the music or call sound.

[0015] In one possible implementation, it further includes: a first digital-to-analog converter (DAC); the input of the first DAC is connected to the output of the audio signal processing path, and the output of the first DAC is connected to the input of the crossover; the first DAC is configured to convert the speaker drive signal from digital form to analog form; correspondingly, the crossover is an analog crossover.

[0016] After receiving the speaker drive signal, if the signal is in digital form before passing through the DAC, but the signal played by the speaker needs to be in analog form, the digital speaker drive signal can be converted into an analog speaker drive signal through the first DAC, and then the analog speaker drive signal can be divided into at least two frequency bands of sub-audio signals through the analog crossover.

[0017] This embodiment uses a structure of first conversion and then frequency division.

[0018] In one possible implementation, it further includes: at least two second DACs; the inputs of the at least two second DACs are all connected to the output of the crossover, and the outputs of the at least two second DACs are respectively connected to one of the at least two speakers; the second DACs are configured to convert one of the sub-audio signals of the at least two frequency bands from digital form to analog form; correspondingly, the crossover is a digital crossover.

[0019] This embodiment uses a structure of frequency division followed by conversion.

[0020] In one possible implementation, the primary operating frequency bands of the at least two speakers are not exactly the same.

[0021] In one possible implementation, the at least two speakers include a moving coil speaker and a balanced armature speaker.

[0022] The TWS earphones in this embodiment are equipped with two speakers: a dynamic speaker and a balanced armature speaker. The main operating frequency band of the dynamic speaker is below 8.5kHz, while the main operating frequency band of the balanced armature speaker is above 8.5kHz. This allows for the use of a crossover to divide the analog speaker drive signal into sub-audio signals below 8.5kHz and sub-audio signals above 8.5kHz. The dynamic speaker can maintain optimal frequency response when playing sub-audio signals below 8.5kHz, and the balanced armature speaker can maintain optimal frequency response when playing sub-audio signals above 8.5kHz. As a result, the TWS earphones can deliver high sound quality across all frequency bands of the audio source and support ultra-wideband voice calls.

[0023] In one possible implementation, the at least two loudspeakers include a moving coil loudspeaker, a moving iron loudspeaker, a microelectromechanical system (MEMS) loudspeaker, and a planar diaphragm.

[0024] The TWS earphones in this embodiment are equipped with four speakers: a dynamic speaker, a balanced armature speaker, a MEMS speaker, and a planar diaphragm. The main operating frequency band of the dynamic speaker is below 8.5kHz, the main operating frequency band of the balanced armature speaker is above 8.5kHz, and the main operating frequency band of the MEMS speaker depends on the application. The main operating frequency band of in-ear earphones is the full frequency band, while the main operating frequency band of over-ear earphones is weaker below 7kHz, with the main operating frequency band being the high frequency above 7kHz. The main operating frequency band of the planar diaphragm is 10kHz to 20kHz. This allows for the use of a crossover to divide the analog speaker drive signal into four sub-bands. The dynamic speaker can maintain optimal frequency response when playing sub-audio signals below 8.5kHz, the balanced armature speaker can maintain optimal frequency response when playing sub-audio signals above 8.5kHz, the MEMS speaker can maintain optimal frequency response when playing sub-audio signals above 7kHz, and the planar diaphragm can maintain optimal frequency response when playing sub-audio signals above 10kHz. This allows the TWS earphones to exhibit high sound quality across all frequency bands of the audio source and support ultra-wideband voice calls.

[0025] Secondly, this application provides a playback method for TWS earphones, which is applied to TWS earphones as described in any one of the first aspects above; the method includes: acquiring an audio source, wherein the audio source is original music or call voice, or the audio source includes a voice signal processed by human voice enhancement and the original music or call voice; performing noise reduction or pass-through processing on the audio source to obtain a speaker driving signal; dividing the speaker driving signal into sub-audio signals of at least two frequency bands, wherein adjacent frequency bands of the at least two frequency bands partially overlap, or the adjacent frequency bands of the at least two frequency bands do not overlap; and playing one of the sub-audio signals of the at least two frequency bands through at least two speakers respectively.

[0026] In one possible implementation, the step of performing noise reduction or pass-through processing on the audio source to obtain the speaker driving signal includes: obtaining a fixed secondary path SP filter through a codec; processing the audio source according to the fixed SP filter to obtain a filtered signal; and performing noise reduction or pass-through processing on the filtered signal to obtain the speaker driving signal.

[0027] In one possible implementation, obtaining the fixed secondary path SP filter via the codec includes: obtaining an estimated SP filter based on a pre-set speaker drive signal and an ear canal signal picked up by a feedback FB microphone, wherein the ear canal signal includes residual noise signals inside the ear canal; and determining the estimated SP filter as the fixed SP filter when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range.

[0028] In one possible implementation, after obtaining the estimated SP filter based on the preset speaker drive signal and the ear canal signal picked up by the feedback FB microphone, the method further includes: when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, obtaining the parameters of the cascaded second-order filter based on the target frequency response of the estimated SP filter and the preset frequency division requirement; obtaining the SP cascaded second-order filter based on the parameters of the cascaded second-order filter, and using the SP cascaded second-order filter as the fixed SP filter.

[0029] In one possible implementation, the step of performing noise reduction or pass-through processing on the audio source to obtain the speaker driving signal includes: obtaining an adaptive SP filter through a digital signal processing (DSP) chip; processing the audio source according to the adaptive SP filter to obtain a filtered signal; and performing noise reduction or pass-through processing on the filtered signal to obtain the speaker driving signal.

[0030] In one possible implementation, obtaining the adaptive SP filter via a digital signal processing (DSP) chip includes: obtaining a real-time noise signal; obtaining an estimated SP filter based on the audio source and the real-time noise signal; and determining the estimated SP filter as the adaptive SP filter when the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range.

[0031] In one possible implementation, acquiring the real-time noise signal includes: acquiring an external signal picked up by a feedforward (FF) microphone and an ear canal signal picked up by a feedback (FB) microphone, wherein the external signal includes an external noise signal and the music or call voice, and the ear canal signal includes residual noise signal inside the ear canal and the music or call voice; acquiring a voice signal picked up by a main microphone; and subtracting the external signal and the ear canal signal from the voice signal to obtain the real-time noise signal.

[0032] In one possible implementation, the primary operating frequency bands of the at least two speakers are not exactly the same.

[0033] In one possible implementation, the at least two loudspeakers include a moving coil loudspeaker and a balanced armature loudspeaker.

[0034] In one possible implementation, the at least two loudspeakers include a moving coil loudspeaker, a moving iron loudspeaker, a microelectromechanical system (MEMS) loudspeaker, and a planar diaphragm.

[0035] Thirdly, this application provides a playback device for TWS earphones, which is applied to the TWS earphones described in the first aspect above; the device includes: an acquisition module for acquiring an audio source, wherein the audio source is original music or call voice, or the audio source includes a voice signal processed by human voice enhancement and the original music or call voice; a processing module for performing noise reduction or pass-through processing on the audio source to obtain a speaker driving signal; a frequency division module for dividing the speaker driving signal into sub-audio signals of at least two frequency bands, wherein adjacent frequency bands of the at least two frequency bands partially overlap, or the adjacent frequency bands of the at least two frequency bands do not overlap; and a playback module for playing one of the sub-audio signals of the at least two frequency bands through at least two speakers respectively.

[0036] In one possible implementation, the processing module is specifically configured to obtain a fixed secondary path SP filter through a codec; process the audio source according to the fixed SP filter to obtain a filtered signal; and perform noise reduction or pass-through processing on the filtered signal to obtain the speaker drive signal.

[0037] In one possible implementation, the processing module is specifically configured to obtain an estimated SP filter based on a pre-set speaker drive signal and an ear canal signal picked up by a feedback FB microphone, wherein the ear canal signal includes residual noise signals inside the ear canal and the music or voice call; when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the estimated SP filter is determined as the fixed SP filter.

[0038] In one possible implementation, the processing module is further configured to, when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, obtain the parameters of the cascaded second-order filter according to the target frequency response of the estimated SP filter and the preset frequency division requirement; obtain the SP cascaded second-order filter according to the parameters of the cascaded second-order filter, and use the SP cascaded second-order filter as the fixed SP filter.

[0039] In one possible implementation, the processing module is specifically configured to obtain an adaptive SP filter through a digital signal processing (DSP) chip; process the audio source according to the adaptive SP filter to obtain a filtered signal; and perform noise reduction or pass-through processing on the filtered signal to obtain the speaker driving signal.

[0040] In one possible implementation, the processing module is specifically configured to acquire a real-time noise signal; acquire an estimated SP filter based on the audio source and the real-time noise signal; and determine the estimated SP filter as the adaptive SP filter when the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range.

[0041] In one possible implementation, the processing module is specifically configured to acquire an external signal picked up by the feedforward (FF) microphone and an ear canal signal picked up by the feedback (FB) microphone, wherein the external signal includes external noise signal and the music or call speech, and the ear canal signal includes residual noise signal inside the ear canal and the music or call speech; acquire a speech signal picked up by the main microphone; subtract the external signal and the ear canal signal from the speech signal to obtain a signal difference; and acquire the estimated SP filter based on the audio source and the signal difference.

[0042] In one possible implementation, the primary operating frequency bands of the at least two speakers are not exactly the same.

[0043] In one possible implementation, the at least two loudspeakers include a moving coil loudspeaker and a balanced armature loudspeaker.

[0044] In one possible implementation, the at least two loudspeakers include a moving coil loudspeaker, a moving iron loudspeaker, a microelectromechanical system (MEMS) loudspeaker, and a planar diaphragm.

[0045] Fourthly, this application provides a computer-readable storage medium including a computer program that, when executed on a computer, causes the computer to perform any of the methods described in the second aspect above.

[0046] Fifthly, this application provides a computer program that, when executed by a computer, performs any of the methods described in the second aspect above. Attached Figure Description

[0047] Figure 1 This is an exemplary structural diagram of a TWS earphone related to related technologies;

[0048] Figure 2a This is an exemplary structural diagram of a TWS earphone related to related technologies;

[0049] Figure 2b This is an exemplary structural diagram of a TWS earphone related to related technologies;

[0050] Figure 3 This is an exemplary structural diagram of a TWS earphone according to this application;

[0051] Figure 4 This is an exemplary structural diagram of a TWS earphone according to this application;

[0052] Figure 5a A flowchart illustrating an exemplary process for obtaining the fixed SP filter in this application.

[0053] Figure 5b A flowchart illustrating an exemplary process for obtaining the fixed SP filter in this application.

[0054] Figure 6 This is a schematic diagram of an exemplary structure of a TWS earphone according to this application.

[0055] Figure 7a This is an exemplary schematic diagram of signal frequency division in this application;

[0056] Figure 7b This is an exemplary schematic diagram of signal frequency division in this application;

[0057] Figure 7c This is an exemplary schematic diagram of signal frequency division in this application;

[0058] Figure 7d This is an exemplary schematic diagram of signal frequency division in this application;

[0059] Figure 8a This is an exemplary structural diagram of a TWS earphone according to this application;

[0060] Figure 8b This is an exemplary structural diagram of a TWS earphone according to this application;

[0061] Figure 8c This is an exemplary structural diagram of a TWS earphone according to this application;

[0062] Figure 8d This is an exemplary structural diagram of a TWS earphone according to this application;

[0063] Figure 8e This is an exemplary structural diagram of a TWS earphone according to this application;

[0064] Figure 9a This is an exemplary structural diagram of a TWS earphone according to this application;

[0065] Figure 9b This is an exemplary structural diagram of a TWS earphone according to this application;

[0066] Figure 9c This is an exemplary structural diagram of a TWS earphone according to this application;

[0067] Figure 9d This is an exemplary structural diagram of a TWS earphone according to this application;

[0068] Figure 9e This is an exemplary structural diagram of a TWS earphone according to this application;

[0069] Figure 10 An exemplary flowchart of the playback method for TWS earphones in this application;

[0070] Figure 11 This is an exemplary structural diagram of the playback device for the TWS earphones of this application. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0072] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0073] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0074] Figure 1 This is an exemplary structural diagram of a TWS (True Wireless Stereo) earphone, a related technology. Figure 1 As shown, the TWS earphone includes three types of microphones: a main microphone, a feedforward (FF) microphone, and a feedback (FB) microphone. The main microphone is used to pick up human voices during calls, the FF microphone is used to pick up external noise signals, and the FB microphone is used to pick up residual noise signals inside the ear canal. The TWS earphone also includes a dynamic speaker, which is used to play processed music or call audio.

[0075] based on Figure 1 The structure shown, Figure 2a This is an exemplary structural diagram of a TWS (True Wireless Stereo) earphone, a related technology. Figure 2a As shown, the main microphone and the FF microphone are connected to the input of the voice enhancement filter, respectively. The output of the voice enhancement filter and the audio source (including music and voice calls) are superimposed by the superimposed signal of the superimposed signal to the superimposed signal of the superimposed signal to the input of the superimposed signal to the superimposed signal to the input of the superimposed signal to the superimposed signal to the input of the superimposed signal to the secondary path (SP) filter. The FF microphone is also connected to the input of the feedforward filter, and the output of the feedforward filter is connected to the other input of the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the superimposed signal to the the superimposed signal to the the the thesofon. The FF microphone is connected to one input of the superimposed signal to the superimposed signal to the superimposed signal to the the thesis. The output of the superimposed signal to the superimposed signal to the superimposed signal to the thesis. The output of ...

[0076] The Active Noise Cancellation & Passthrough & Enhanced Hearing (ANC & HT & AH, AHA) joint controller is connected to the voice enhancement filter, feedforward filter, feedback filter, and SP filter, respectively. The function of the AHA joint controller is to handle abnormal situations by the TWS earbuds themselves, ensuring their normal and stable operation. The AHA joint controller analyzes multiple signals to determine the current state of ANC, HT, or enhanced hearing (AH), and then determines whether abnormal situations such as howling or clipping have occurred, implementing corresponding processing measures. System control is achieved by controlling the parameter values ​​of the aforementioned filters.

[0077] In the above structure, the human voice enhancement filter, audio source and AHA joint controller are set in the digital signal processing (DSP) chip, while the feedforward filter, feedback filter, SP filter and DAC are set in the coder-decoder (CODEC).

[0078] based on Figure 1 The structure shown, Figure 2b This is an exemplary structural diagram of a TWS (True Wireless Stereo) earphone, a related technology. Figure 2b As shown, with Figure 2a The structural difference shown is that the human voice enhancement filter has been moved from the DSP chip to the CODEC.

[0079] based on Figures 1 to 2b With the structure shown, TWS earbuds can achieve the following functions:

[0080] Active noise cancellation: The front-facing (FF) microphone picks up external noise signals and generates a feedforward noise-canceling signal through a feedforward filter; the front-facing (FB) microphone picks up residual noise signals inside the ear canal and generates a feedback noise-canceling signal through a feedback filter. The feedforward noise-canceling signal, the feedback noise-canceling signal, and the downlink signal are superimposed to form the final speaker drive signal, which is then converted from digital to analog by a DAC to generate an analog speaker drive signal. When this analog speaker drive signal is played in reverse in a dynamic speaker, it cancels out the audio signals in the space, thus obtaining a noise-canceling analog audio signal. In this way, low-frequency noise within a certain frequency band can be eliminated, thereby achieving the purpose of noise reduction.

[0081] Pass-through: The FF microphone picks up external noise signals and generates a feedforward compensation signal through a feedforward filter; the FB microphone picks up acoustic signals inside the ear canal and generates a feedback suppression signal through a feedback filter. The feedforward compensation signal, feedback suppression signal, and downlink signal are superimposed to form the final speaker drive signal, which is then converted from digital to analog by a DAC to generate an analog speaker drive signal. This analog speaker drive signal is played in a dynamic speaker, which can reduce or suppress low-frequency noise and compensate for high-frequency signals, thus obtaining an analog audio signal that achieves acoustic compensation. In this way, the "blockage" effect of active sound production (the wearer hearing their own voice) or the "stethoscope" effect of sound production from body vibrations (such as walking, chewing, scratching, etc. while wearing headphones) is reduced or suppressed, and high-frequency sounds within a certain frequency band of passive sound production (such as human voices or music in the environment) are compensated, thereby achieving the purpose of pass-through.

[0082] It should be noted that in the above steps of superimposing to obtain the final speaker drive signal, the superimposed signal may include all three: the feedforward compensation signal, the feedback suppression signal, and the downlink signal, or it may include two or one of these three signals. For example, when noise reduction is performed while the user is listening to music, the superimposed signal may include all three; when the user only wants to use headphones for noise reduction to create a quiet environment, the superimposed signal may include the feedforward compensation signal and the feedback suppression signal, but not the downlink signal.

[0083] Voice enhancement and concurrent music / call playback: When implementing voice enhancement on the DSP side, the signals from the FF microphone and main microphone are sent to the voice enhancement filter for processing, resulting in a signal where ambient noise is suppressed and the voice is preserved. This signal is then sent down to the CODEC, where it is superimposed with the signals output from the feedforward filter and feedback filter to obtain the speaker drive signal. After digital-to-analog conversion by the DAC, an analog speaker drive signal is output for the dynamic speaker to play. Since the sound played by the dynamic speaker is picked up by the FB microphone, to prevent the played sound from being denoised by the feedback filter, the estimated signal played by the dynamic speaker can be subtracted from the signal picked up by the FB microphone. This part of the signal will not be denoised by the feedback filter, thus achieving concurrent voice enhancement and music / call playback. ANC / HT functions can also be implemented on this basis.

[0084] Figure 2a In this approach, implementing vocal enhancement in the DSP ensures computational overhead, resulting in good noise reduction and stable playback performance. However, it also has a longer latency, which can lead to a reverberant sound. Figure 2b In this study, voice enhancement is implemented in CODEC. The algorithm is relatively fixed and simple, with short latency and low reverberation, but the noise reduction effect is limited.

[0085] Figures 1 to 2b The TWS earphone structure shown may severely degrade the sound quality of high-frequency music, especially music signals above 8.5kHz, affecting the music playback effect; it also does not support high-frequency voice, especially voice signals above 8.5kHz, resulting in limited bandwidth for voice calls.

[0086] This application provides a speaker structure for TWS earphones, which can improve the above-mentioned technical problems.

[0087] Figure 3 This is an exemplary structural diagram of a TWS earphone according to this application. Figure 3 As shown, the TWS earphone 30 includes three types of microphones: a main microphone 31, an FF microphone 32, and an FB microphone 33. The main microphone 31 is used to pick up the human voice during a call, the FF microphone 32 is used to pick up external noise signals, and the FB microphone 33 is used to pick up residual noise signals inside the ear canal.

[0088] The TWS earphones 30 also include a dynamic speaker 34a and a balanced armature speaker 34b. The main operating frequency band of the dynamic speaker 34a is less than 8.5kHz, and the main operating frequency band of the balanced armature speaker 34b is greater than 8.5kHz. It should be noted that this application does not specifically limit the number of speakers, as long as there are at least two. These at least two speakers can have different main operating frequency bands, or some speakers can have the same main operating frequency band, while others have a different main operating frequency band. For example, with three speakers (1-3), the main operating frequency bands of speaker 1 and speaker 2 are the same, while the main operating frequency band of speaker 3 is different from that of speakers 1 and 2. Another example is with three speakers (1-3), where the main operating frequency bands of speakers 1, 2, and 3 are all different. Optionally, at least two loudspeakers may include dynamic loudspeakers, balanced armature loudspeakers, micro-electro-mechanical systems (MEMS) loudspeakers, and planar diaphragms. The main operating frequency band of dynamic loudspeakers is below 8.5kHz, the main operating frequency band of balanced armature loudspeakers is above 8.5kHz, and the main operating frequency band of MEMS loudspeakers depends on the application. The main operating frequency band of in-ear headphones is the full frequency band, the main operating frequency band of over-ear headphones is weaker below 7kHz, and the main operating frequency band is the high frequency above 7kHz. The main operating frequency band of planar diaphragms is 10kHz to 20kHz.

[0089] based on Figure 3 The TWS earphones shown Figure 4 This is an exemplary structural diagram of a TWS earphone according to this application. Figure 4 As shown, the TWS earphone 40 includes: an audio signal processing path 41, a crossover 42, and at least two speakers 43; wherein, the output terminal of the audio signal processing path 41 is connected to the input terminal of the crossover 42; and the output terminal of the crossover 42 is connected to at least two speakers 43.

[0090] The audio signal processing path 41 is configured to output a speaker drive signal after performing noise reduction or pass-through processing on the audio source; the audio source is the original music or voice call; or, the audio source includes a voice signal that has been enhanced with human voice and the original music or voice call.

[0091] Crossover 42 is configured to divide the speaker drive signal into at least two sub-audio signals, the at least two frequency bands corresponding to the main operating frequency bands of at least two speakers 43; adjacent frequency bands of the at least two frequency bands partially overlap, or adjacent frequency bands of the at least two frequency bands do not overlap. At least two speakers 43 are configured to play the corresponding sub-audio signals.

[0092] The frequency band of the processed speaker drive signal corresponds to the frequency band of the audio source and may include the entire low, mid and high frequency bands. However, since the main operating frequency band of a single speaker may only cover a part of the low, mid and high frequency bands, the single speaker cannot reproduce high sound quality across the entire frequency band.

[0093] The crossover 42 of this application can be configured to divide the speaker drive signal based on the main operating frequency bands of at least two speakers, obtaining sub-audio signals of at least two frequency bands corresponding to the main operating frequency bands of the at least two speakers respectively. Each speaker then plays the corresponding sub-audio signal, thereby maintaining optimal frequency response when playing the sub-audio signals transmitted to it. The parameters of the crossover 42 can be set to control the crossover 42 to divide the speaker drive signal according to a preset method.

[0094] In one possible implementation, the aforementioned audio signal processing path 41 can employ... Figure 2a or Figure 2b The structure shown indicates that the SP filter is set in the coder-decoder (CODEC), and a fixed SP filter is obtained through the CODEC.

[0095] CODEC can obtain an estimated SP filter based on a pre-set speaker drive signal and an ear canal signal picked up by an FB microphone. The ear canal signal includes residual noise signals inside the ear canal. When the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the estimated SP filter is determined as a fixed SP filter.

[0096] Figure 5a Here is an exemplary flowchart of obtaining the fixed SP filter in this application, as follows: Figure 5a As shown, a speaker drive signal x[n] is preset, which drives at least two speakers to emit sound after passing through a frequency divider. The sound is transmitted to the FB microphone and picked up by the FB microphone. These sounds are converted into digital signals y[n].

[0097] Typically, the transfer function of an SP filter can be assumed to be a high-order FIR filter, and iterative modeling can be performed using the Least Mean Square Error (LMS) algorithm. The actual transfer function S(z) of the SP filter is unknown, but all its information is contained in the speaker drive signal x[n] and the digital signal y[n] picked up by the FB microphone. Therefore, it can be modeled using a high-order FIR filter. As an estimation SP filter to simulate S(z). Input x[n] get For y[n] and The difference is used to obtain the error signal e[n]. When e[n] is within a set range, the algorithm is considered to have converged to a satisfactory state, and the high-order FIR filter at this point can be considered to be in a satisfactory state. Since it is approximately equal to the true S(z), it can be determined as the transfer function of the fixed SP filter. The above process can be iteratively modeled using the least mean square (LMS) algorithm.

[0098] Furthermore, in addition to simulating a fixed SP filter using a finite impulse response (FIR) filter, a fixed SP filter can also be simulated using an infinite impulse response (IIR) filter on the CODEC side. Figure 5b Here is an exemplary flowchart of obtaining the fixed SP filter in this application, as follows: Figure 5b As shown, when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within the set range, the parameters of the cascaded second-order filter are obtained according to the target frequency response of the estimated SP filter and the pre-set frequency division requirements; the SP cascaded second-order filter is obtained according to the parameters of the cascaded second-order filter, and the SP cascaded second-order filter is used as the fixed SP filter.

[0099] (1) CODEC uses the already obtained FIR filter Calculate the target frequency response.

[0100] s = [s0 s1 … s N-1], ∈N represents the filter order, and k represents the number of frequency points. The corresponding target response is calculated using formula (1):

[0101]

[0102] (2) The target frequency response is divided into frequencies, and different weighting coefficients ω[k] are set in different frequency bands.

[0103] (3) Obtain the IIR filter parameters by converting the FIR filter to the IIR filter. This process requires that the target responses of the two filters be as consistent as possible.

[0104] From a mathematical perspective, this means finding the objective function of formula (2) in order to obtain the coefficients of multiple IIR filters:

[0105]

[0106] Where B[k] and A[k] are the complex frequency responses of the IIR filter coefficients b and a, respectively, and their calculation methods are the same as those in formula (1).

[0107] (4) Implementation of IIR filters

[0108] By using optimization algorithms, a set of IIR filter parameters with minimum mean square can be found. The IIR filter obtained from this set of IIR filter parameters can be used as the final fixed SP filter and put into the CODEC for hardening implementation.

[0109] In one possible implementation, the aforementioned audio signal processing path 41 can employ... Figure 6 The structure shown indicates that the SP filter is located in the digital signal processing (DSP) chip, and the adaptive SP filter is obtained through the DSP chip.

[0110] The DSP chip can acquire real-time noise signals and obtain an estimated SP filter based on the audio source (the audio source is the original music or voice call; or, the audio source includes voice-enhanced voice signals and the original music or voice call) and the real-time noise signal. When the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range, the estimated SP filter is determined as an adaptive SP filter.

[0111] The above-mentioned acquisition of real-time noise signal can be achieved by acquiring external signals picked up by the FF microphone (external signals include external noise signals and music or call voice) and ear canal signals picked up by the FB microphone (ear canal signals include residual noise signals inside the ear canal and music or call voice), and then acquiring the voice signal picked up by the main microphone. The external signals and ear canal signals are then subtracted from the voice signal to obtain the real-time noise signal.

[0112] The difference between this embodiment and the embodiment described above for obtaining a fixed SP filter is that: x[n] represents the audio source, while y[n] represents the real-time noise signal, i.e., x[n] = dnlink[n], y[n] = fb[n] - ff[n]*A(z) - fb[n-1]*C(z), where fb[n] represents the speech signal picked up by the main microphone, ff[n]*A(z) represents the external signal picked up by the FF microphone, and fb[n-1]*C(z) represents the ear canal signal picked up by the FB microphone. This embodiment can also use... Figure 5a The procedure shown yields a high-order FIR filter corresponding to the minimum error signal e[n]. Since the audio source is real-time, the noise reduction or pass-through function of TWS earphones will run concurrently with music or voice calls. At this time, the signal picked up by the FB microphone cannot be directly estimated by the SP filter. It is necessary to remove the influence of other signals before modeling and analysis, that is, to obtain the real-time noise signal. Therefore, the SP filter obtained at this time can adapt to the real-time situation of the audio source and noise signal rather than be fixed.

[0113] This application determines the SP filter (including a fixed SP filter or an adaptive SP filter) through the above method, which can achieve noise reduction or pass-through functions during music playback or transmission, and can also prevent the pass-through or noise reduction technology from eliminating the music or call sound.

[0114] For example, suppose TWS earbuds include two speakers: a dynamic speaker and a balanced armature speaker. The dashed line represents the frequency response curve of the dynamic speaker, the dotted line represents the frequency response curve of the balanced armature speaker, and the solid line represents the crossover line.

[0115] Figure 7a This is an exemplary schematic diagram of signal frequency division for this application, such as... Figure 7a As shown, the crossover 42 is configured to attenuate the non-primary operating frequency band of the moving-coil speaker while retaining the power of the moving-coil speaker in the primary operating frequency band, attenuate the non-primary operating frequency band of the balanced-iron speaker while retaining the power of the balanced-iron speaker in the primary operating frequency band, and perform frequency division at the intersection frequency point within the attenuation frequency bands of the moving-coil speaker and the balanced-iron speaker to obtain sub-audio signals of two frequency bands.

[0116] Figure 7b This is an exemplary schematic diagram of signal frequency division for this application, such as... Figure 7b As shown, the crossover 42 is configured to attenuate the non-primary operating frequency band of the balanced iron speaker, retain the power of the balanced iron speaker in the primary operating frequency band, and perform frequency division at a certain frequency point within the attenuation frequency band of the balanced iron speaker to obtain sub-audio signals of two frequency bands.

[0117] Figure 7c This is an exemplary schematic diagram of signal frequency division for this application, such as... Figure 7c As shown, the crossover 42 is configured to attenuate the non-primary operating frequency band of the moving-coil speaker in two stages, preserving the power of the moving-coil speaker in the primary operating frequency band; attenuate the non-primary operating frequency band of the balanced-iron speaker in two stages, preserving the power of the balanced-iron speaker in the primary operating frequency band; and perform a frequency division at the transition point between the first and second attenuation frequency bands of the balanced-iron speaker, and a frequency division at the transition point between the first and second attenuation frequency bands of the moving-coil speaker, to obtain sub-audio signals of three frequency bands.

[0118] For example, suppose TWS earbuds include three speakers: a dynamic speaker, a balanced armature speaker, and a MEMS speaker. The dashed line represents the frequency response curve of the dynamic speaker, the single-dot line represents the frequency response curve of the balanced armature speaker, the double-dot line represents the frequency response curve of the MEMS speaker, and the solid line represents the crossover line.

[0119] Figure 7d This is an exemplary schematic diagram of signal frequency division for this application, such as... Figure 7d As shown, in Figure 7c Based on the example shown, the crossover 42 is configured to attenuate the non-primary operating frequency band of the MEMS speaker, retain the power of the MEMS speaker in the primary operating frequency band, and perform a crossover at the intersection frequency of the attenuation frequency bands of the moving iron speaker and the MEMS speaker, resulting in a total of four sub-audio signals.

[0120] It should be noted that, Figures 7a to 7d The examples shown are several examples of how the crossover 42 divides the speaker drive signal. This application does not limit the specific frequency division method of the crossover 42.

[0121] The TWS earphones of this application are equipped with at least two speakers, and the main operating frequency bands of the at least two speakers are not exactly the same. A crossover can divide the speaker driving signal into sub-audio signals of at least two frequency bands. Adjacent frequency bands of the at least two frequency bands may partially overlap or not overlap. In this way, each sub-audio signal is transmitted to a frequency band-matched speaker. The aforementioned frequency band matching may mean that the main operating frequency band of the speaker covers the frequency band of the sub-audio signal transmitted to it. In this way, the speaker maintains the optimal frequency response when playing the transmitted sub-audio signal, which can not only reflect high sound quality in all frequency bands of the audio source, but also support ultra-wideband voice calls.

[0122] In one possible implementation, Figure 8a This is an exemplary structural diagram of a TWS earphone according to this application. Figure 8a As shown, in Figure 4Based on the structure shown, the TWS earphone 40 also includes a first DAC 44. The input terminal of the first DAC 44 is connected to the output terminal of the audio signal processing path 41, and the output terminal of the first DAC 44 is connected to the input terminal of the crossover 42.

[0123] The first DAC 44 is configured to convert the speaker drive signal from digital to analog form. Correspondingly, the crossover 42 is an analog crossover.

[0124] Figure 2a or Figure 2b In the structure shown, after the speaker drive signal is obtained, it is in digital form before it passes through the DAC. However, the signal played by the speaker needs to be in analog form. Therefore, the digital speaker drive signal can be converted into an analog speaker drive signal through the first DAC, and then the analog speaker drive signal can be divided into at least two frequency bands of sub-audio signals through the analog crossover.

[0125] This embodiment uses a structure of first conversion and then frequency division.

[0126] For example, Figure 8b This is an exemplary structural diagram of a TWS earphone according to this application. Figure 8b As shown, the structure of this embodiment is Figure 8a A more detailed implementation of the structure shown.

[0127] The main microphone 601 and FF microphone 602 in the TWS earphone 60 are connected to the input of the voice enhancement filter 603. The output of the voice enhancement filter 603 and the audio source 604 (including music and call voice) are superimposed by the superimposed amplifier 1 to obtain the downlink signal. The downlink signal is transmitted to one input of the superimposed amplifier 2. The downlink signal is also transmitted to the input of the SP filter 605. The FF microphone 602 is also connected to the input of the feedforward filter 606, and the output of the feedforward filter 606 is connected to the other input of the superimposed amplifier 2. The FB microphone 607 is connected to one input of the superimposed amplifier 3. The output of the SP filter 605 is connected to the other input of the superimposed amplifier 3. The output of the superimposed amplifier 3 is connected to the input of the feedback filter 608, and the output of the feedback filter 608 is connected to the third input of the superimposed amplifier 2. The output of the superimposed amplifier 2 is connected to the input of the digital-to-analog converter (DAC) 609. The output of 609 is connected to the input of analog crossover 610, and the output of analog crossover 610 is connected to dynamic speaker 611a and balanced armature speaker 611b. AHA combined controller 612 is connected to vocal enhancement filter 603, feedforward filter 606, feedback filter 608 and SP filter 605 respectively.

[0128] The TWS earphone 60 in this embodiment is equipped with two speakers, namely a dynamic speaker 611a and a balanced armature speaker 611b. The main operating frequency band of the dynamic speaker 611a is below 8.5kHz, and the main operating frequency band of the balanced armature speaker 611b is above 8.5kHz. This allows the crossover 42 to divide the analog speaker drive signal into sub-audio signals below 8.5kHz and sub-audio signals above 8.5kHz. The dynamic speaker 611a can maintain the best frequency response when playing sub-audio signals below 8.5kHz, and the balanced armature speaker 611b can maintain the best frequency response when playing sub-audio signals above 8.5kHz. Thus, the TWS earphone 60 can not only exhibit high sound quality in all frequency bands of the audio source, but also support ultra-wideband voice calls.

[0129] For example, Figure 8c This is an exemplary structural diagram of a TWS earphone according to this application. Figure 8c As shown, the structure of this embodiment is Figure 8a Another more detailed implementation of the structure shown.

[0130] and Figure 8b The difference in the structure shown is that the output of the analog crossover 610 is connected to a moving coil speaker 611a, a moving iron speaker 611b, a MEMS speaker 611c, and a planar diaphragm 611d.

[0131] The TWS earphone 60 in this embodiment includes four speakers: a dynamic speaker 611a, a balanced armature speaker 611b, a MEMS speaker 611c, and a planar diaphragm 611d. The main operating frequency band of the dynamic speaker 611a is below 8.5kHz, the main operating frequency band of the balanced armature speaker 611b is above 8.5kHz, and the main operating frequency band of the MEMS speaker 611c depends on the application. The main operating frequency band of in-ear earphones is the entire frequency band, while the main operating frequency band of over-ear earphones is weaker below 7kHz, with the main operating frequency band being the high frequency above 7kHz. The main operating frequency band of the planar diaphragm 611d is 10kHz. With a frequency range of z~20kHz, the crossover 42 can be configured to divide the analog speaker drive signal into four sub-bands. The dynamic speaker 611a can maintain the best frequency response when playing sub-audio signals below 8.5kHz, the balanced armature speaker 611b can maintain the best frequency response when playing sub-audio signals above 8.5kHz, the MEMS speaker 611c can maintain the best frequency response when playing sub-audio signals above 7kHz, and the planar diaphragm 611d can maintain the best frequency response when playing sub-audio signals above 10kHz. This allows the TWS earphones 60 to not only exhibit high sound quality across all frequency bands of the audio source, but also support ultra-wideband voice calls.

[0132] In the above structure, the human voice enhancement filter 603, the audio source 604 and the AHA joint controller 612 are located in the digital signal processing (DSP) chip, while the feedforward filter 606, the feedback filter 608, the SP filter 605 and the DAC 609 are located in the coder-decoder (CODEC).

[0133] For example, Figure 8d This is an exemplary structural diagram of a TWS earphone according to this application. Figure 8d As shown, the structure of this embodiment is Figure 8a A more detailed implementation of the structure shown.

[0134] and Figure 8b The difference in the structure shown is that the human voice enhancement filter 603 has been moved from the DSP chip to the CODEC.

[0135] For example, Figure 8e This is an exemplary structural diagram of a TWS earphone according to this application. Figure 8e As shown, the structure of this embodiment is Figure 8a A more detailed implementation of the structure shown.

[0136] and Figure 8c The difference in the structure shown is that the human voice enhancement filter 603 has been moved from the DSP chip to the CODEC.

[0137] In one possible implementation, Figure 9a This is an exemplary structural diagram of a TWS earphone according to this application. Figure 9a As shown, in Figure 4 Based on the structure shown, the TWS earphone 40 also includes at least two second DACs 45. The input terminals of the at least two second DACs 45 are all connected to the output terminal of the crossover 42, and the output terminals of the at least two second DACs 45 are respectively connected to one of the at least two speakers 43.

[0138] The second DAC 45 is configured to convert one of the sub-audio signals from at least two frequency bands from digital to analog form. Correspondingly, the frequency divider 42 is a digital frequency divider.

[0139] Figure 2a or Figure 2bIn the structure shown, after the speaker drive signal is obtained, it is in digital form before it passes through the DAC. However, the signal played by the speaker needs to be in analog form. Therefore, the digital speaker drive signal can be divided into at least two frequency band sub-audio signals by a digital frequency divider. Then, at least two second DACs convert the digital sub-audio signals transmitted to them into analog sub-audio signals respectively.

[0140] This embodiment uses a structure of frequency division followed by conversion.

[0141] For example, Figure 9b This is an exemplary structural diagram of a TWS earphone according to this application. Figure 9b As shown, the structure of this embodiment is Figure 9a A more detailed implementation of the structure shown.

[0142] In the TWS earphone 70, the main microphone 701 and the FF microphone 702 are connected to the input of the voice enhancement filter 703. The output of the voice enhancement filter 703 and the audio source 704 (including music and call voice) are superimposed by the superimposed signal of the superimposed signal generator 1 to obtain the downlink signal. The downlink signal is transmitted to one input of the superimposed signal generator 2. The downlink signal is also transmitted to the input of the SP filter 705. The FF microphone 702 is also connected to the input of the feedforward filter 706, and the output of the feedforward filter 706 is connected to the other input of the superimposed signal generator 2. The FB microphone 707 is connected to one input of the superimposed signal generator 3. The output of the SP filter 705 is connected to the other input of the superimposed signal generator 3. The output of the superimposed signal generator 3 is connected to the input of the feedback filter 708, and the output of the feedback filter 708 is connected to the third input of the superimposed signal generator 2. The output of the superimposed signal generator 2 is connected to the input of the digital crossover 709, and the output of the digital crossover 709 is connected to two DACs. The input terminals of 710a and 710b are connected; the output terminal of DAC710a is connected to the dynamic speaker 711a, and the output terminal of DAC 710b is connected to the balanced armature speaker 711b. The AHA combined controller 712 is connected to the vocal enhancement filter 703, the feedforward filter 706, the feedback filter 708, and the SP filter 705, respectively.

[0143] For example, Figure 9c This is an exemplary structural diagram of a TWS earphone according to this application. Figure 9c As shown, the structure of this embodiment is Figure 9a A more detailed implementation of the structure shown.

[0144] The TWS earphone 70 in this embodiment includes four speakers: a dynamic speaker 711a, a balanced armature speaker 711b, a MEMS speaker 711c, and a planar diaphragm 711d. The main operating frequency band of the dynamic speaker 711a is below 8.5kHz, the main operating frequency band of the balanced armature speaker 711b is above 8.5kHz, and the main operating frequency band of the MEMS speaker 711c depends on the application. The main operating frequency band of in-ear earphones is the entire frequency band, while the main operating frequency band of over-ear earphones is weaker below 7kHz, with the main operating frequency band being the high frequency above 7kHz. The main operating frequency band of the planar diaphragm 711d is 10kHz. With a frequency range of z~20kHz, the crossover 42 can be configured to divide the analog speaker drive signal into four sub-bands. The dynamic speaker 711a can maintain the best frequency response when playing sub-audio signals below 8.5kHz, the balanced armature speaker 711b can maintain the best frequency response when playing sub-audio signals above 8.5kHz, the MEMS speaker 711c can maintain the best frequency response when playing sub-audio signals above 7kHz, and the planar diaphragm 711d can maintain the best frequency response when playing sub-audio signals above 10kHz. This allows the TWS earphones 70 to not only exhibit high sound quality across all frequency bands of the audio source, but also support ultra-wideband voice calls.

[0145] In the above structure, the human voice enhancement filter 703, the audio source 704 and the AHA joint controller 712 are located in the DSP chip, while the feedforward filter 706, the feedback filter 708, the SP filter 705 and the DAC 709 are located in the CODEC.

[0146] For example, Figure 9d This is an exemplary structural diagram of a TWS earphone according to this application. Figure 9d As shown, the structure of this embodiment is Figure 9a A more detailed implementation of the structure shown.

[0147] and Figure 9b The difference in the structure shown is that the human voice enhancement filter 703 has been moved from the DSP chip to the CODEC.

[0148] For example, Figure 9e This is an exemplary structural diagram of a TWS earphone according to this application. Figure 9e As shown, the structure of this embodiment is Figure 9a A more detailed implementation of the structure shown.

[0149] and Figure 9c The difference in the structure shown is that the human voice enhancement filter 703 has been moved from the DSP chip to the CODEC.

[0150] Figure 10An exemplary flowchart of the playback method of the TWS earphones of this application is shown below. Figure 10 As shown, the method of this embodiment can be applied to the TWS earphones in the above embodiments. The method may include:

[0151] Step 1001: Obtain the audio source.

[0152] Optionally, the audio source is the original music or call audio; that is, the audio source can be music or video sound that the user is listening to with headphones, or call audio that the user is making with headphones. This audio source can come from the player of an electronic device. Optionally, the audio source includes a voice-enhanced audio signal and the original music or call audio; that is, in addition to the music or call audio in the above two cases, the audio source can also superimpose a voice-enhanced external audio signal. This voice-enhanced external audio signal can be... Figure 2a or Figure 2b The human voice enhancement filter in the structure shown is obtained, and will not be described in detail here.

[0153] Step 1002: Perform noise reduction or pass-through processing on the audio source to obtain the speaker driving signal.

[0154] In one possible implementation, a fixed secondary path SP filter can be obtained through a CODEC. The audio source is then processed using this fixed SP filter to obtain a filtered signal. This filtered signal is then subjected to noise reduction or pass-through processing to obtain a speaker drive signal. Specifically, the CODEC can obtain an estimated SP filter based on a pre-set speaker drive signal and the ear canal signal picked up by the feedback FB microphone. The ear canal signal includes residual noise signals within the ear canal. When the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the estimated SP filter is determined as the fixed SP filter. Optionally, when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the parameters of a cascaded second-order filter are obtained based on the target frequency response of the estimated SP filter and pre-set frequency division requirements. An SP-cascaded second-order filter is then obtained based on these parameters and used as the fixed SP filter.

[0155] In one possible implementation, an adaptive SP filter can be obtained through a DSP chip. The audio source is then processed using the adaptive SP filter to obtain a filtered signal. This filtered signal is then subjected to noise reduction or pass-through processing to obtain a speaker drive signal. Specifically, the DSP chip can acquire a real-time noise signal and obtain an estimated SP filter based on the audio source and the real-time noise signal. When the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range, the estimated SP filter is determined as the adaptive SP filter. Optionally, the DSP chip can first acquire the external signal picked up by the FF microphone and the ear canal signal picked up by the FB microphone. The external signal includes external noise and music or conversation speech, while the ear canal signal includes residual noise and music or conversation speech within the ear canal. Then, it can acquire the speech signal picked up by the main microphone. Finally, the external signal and the ear canal signal are subtracted from the speech signal to obtain the real-time noise signal.

[0156] The noise reduction and pass-through processing in this application can be referred to the above embodiments, and will not be repeated here.

[0157] Step 1003: Divide the speaker drive signal into at least two frequency bands of sub-audio signals.

[0158] The frequency band of the processed speaker drive signal corresponds to the frequency band of the audio source and may include the entire low, mid and high frequency bands. However, since the main operating frequency band of a single speaker may only cover a part of the low, mid and high frequency bands, the single speaker cannot reproduce high sound quality across the entire frequency band.

[0159] The crossover of this application can be configured to divide the speaker drive signal based on the main operating frequency bands of at least two speakers, obtaining sub-audio signals of at least two frequency bands corresponding to the main operating frequency bands of the at least two speakers respectively. Each speaker then plays the corresponding sub-audio signal, thereby maintaining optimal frequency response when playing the sub-audio signal transmitted to it. The adjacent frequency bands of the aforementioned at least two frequency bands may partially overlap, or the adjacent frequency bands of the at least two frequency bands may not overlap. For example, the crossover divides the speaker drive signal into two frequency bands, high and low, with the high-frequency and low-frequency bands completely separated and without overlap; or, the high-frequency and low-frequency bands may partially overlap. As another example, the crossover divides the speaker drive signal into three frequency bands: high, mid, and low. The high-frequency and mid-frequency bands are completely separated and without overlap, while the mid-frequency and low-frequency bands partially overlap; or, the high-frequency, mid-frequency, and low-frequency bands are completely separated and without overlap; or, the high-frequency and mid-frequency bands partially overlap, while the mid-frequency and low-frequency bands are completely separated and without overlap.

[0160] This application allows setting the parameters of the crossover 42 to control the crossover 42 to divide the speaker drive signal according to a preset method. The crossover implementation method can be found in [reference needed]. Figures 7a to 7d This will not be elaborated upon here.

[0161] Step 1004: Play one of the sub-audio signals of at least two frequency bands through at least two speakers.

[0162] The TWS earphones of this application are equipped with at least two speakers, and the main operating frequency bands of the at least two speakers are not exactly the same. A crossover can divide the speaker driving signal into sub-audio signals of at least two frequency bands. Adjacent frequency bands of the at least two frequency bands may partially overlap or not overlap. In this way, each sub-audio signal is transmitted to a frequency band-matched speaker. The aforementioned frequency band matching may mean that the main operating frequency band of the speaker covers the frequency band of the sub-audio signal transmitted to it. In this way, the speaker maintains the optimal frequency response when playing the transmitted sub-audio signal, which can not only reflect high sound quality in all frequency bands of the audio source, but also support ultra-wideband voice calls.

[0163] Figure 11 This is an exemplary structural diagram of the playback device for the TWS earphones of this application, as shown below. Figure 11 As shown, the device 1100 of this embodiment can be applied to the TWS earphones in the above embodiments. The device 1100 includes: an acquisition module 1101, a processing module 1102, a frequency division module 1103, and a playback module 1104. Wherein,

[0164] The acquisition module 1101 is used to acquire an audio source, wherein the audio source is original music or call voice, or the audio source includes a voice signal that has been processed by human voice enhancement and the original music or call voice; the processing module 1102 is used to perform noise reduction or pass-through processing on the audio source to obtain a speaker driving signal; the frequency division module 1103 is used to divide the speaker driving signal into sub-audio signals of at least two frequency bands, wherein adjacent frequency bands of the at least two frequency bands partially overlap, or the adjacent frequency bands of the at least two frequency bands do not overlap; and the playback module 1104 is used to play one of the sub-audio signals of the at least two frequency bands through at least two speakers respectively.

[0165] In one possible implementation, the processing module 1102 is specifically used to obtain a fixed secondary path SP filter through a codec; process the audio source according to the fixed SP filter to obtain a filtered signal; and perform noise reduction or pass-through processing on the filtered signal to obtain the speaker driving signal.

[0166] In one possible implementation, the processing module 1102 is specifically configured to obtain an estimated SP filter based on a preset speaker drive signal and an ear canal signal picked up by a feedback FB microphone, wherein the ear canal signal includes residual noise signals inside the ear canal and the music or voice call; when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the estimated SP filter is determined as the fixed SP filter.

[0167] In one possible implementation, the processing module 1102 is further configured to, when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, obtain the parameters of the cascaded second-order filter according to the target frequency response of the estimated SP filter and the preset frequency division requirement; obtain the SP cascaded second-order filter according to the parameters of the cascaded second-order filter, and use the SP cascaded second-order filter as the fixed SP filter.

[0168] In one possible implementation, the processing module 1102 is specifically used to obtain an adaptive SP filter through a digital signal processing (DSP) chip; process the audio source according to the adaptive SP filter to obtain a filtered signal; and perform noise reduction or pass-through processing on the filtered signal to obtain the speaker driving signal.

[0169] In one possible implementation, the processing module 1102 is specifically used to acquire a real-time noise signal; acquire an estimated SP filter based on the audio source and the real-time noise signal; and determine the estimated SP filter as the adaptive SP filter when the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range.

[0170] In one possible implementation, the processing module 1102 is specifically configured to acquire an external signal picked up by the feedforward (FF) microphone and an ear canal signal picked up by the feedback (FB) microphone, wherein the external signal includes external noise signal and the music or call speech, and the ear canal signal includes residual noise signal inside the ear canal and the music or call speech; acquire a speech signal picked up by the main microphone; subtract the external signal and the ear canal signal from the speech signal to obtain a signal difference; and acquire the estimated SP filter based on the audio source and the signal difference.

[0171] In one possible implementation, the primary operating frequency bands of the at least two speakers are not exactly the same.

[0172] In one possible implementation, the at least two loudspeakers include a moving coil loudspeaker and a balanced armature loudspeaker.

[0173] In one possible implementation, the at least two loudspeakers include a moving coil loudspeaker, a moving iron loudspeaker, a microelectromechanical system (MEMS) loudspeaker, and a planar diaphragm.

[0174] The apparatus of this embodiment can be used to perform Figure 10 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.

[0175] In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware encoding processor, or implemented by a combination of hardware and software modules in the encoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0176] The memory mentioned in the above embodiments can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0177] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0178] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0179] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0180] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0181] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0182] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0183] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A true wireless stereo (TWS) earphone, characterized in that, include: An audio signal processing path, a crossover, and at least two speakers; wherein, The output of the audio signal processing path is connected to the input of the crossover; the output of the crossover is connected to the at least two speakers. The audio signal processing path is configured to output a speaker driving signal after performing noise reduction or pass-through processing on the audio source; the audio source is the original music or voice call; or, the audio source includes a voice signal that has undergone voice enhancement processing and the original music or voice call. The frequency divider is configured to divide the speaker drive signal into at least two frequency bands of sub-audio signals, the at least two frequency bands corresponding to the main operating frequency bands of the at least two speakers; adjacent frequency bands in the at least two frequency bands partially overlap, or adjacent frequency bands in the at least two frequency bands do not overlap; The at least two speakers are configured to play corresponding sub-audio signals; The audio signal processing path includes: The secondary path SP filter is configured to prevent the noise reduction or pass-through processing from eliminating the sound of the audio source when the noise reduction or pass-through processing is concurrent with the audio source. The SP filter includes a fixed SP filter or an adaptive SP filter; The fixed SP filter is obtained through a codec (codec). The codec obtains an estimated SP filter based on a pre-set speaker drive signal and an ear canal signal picked up by a feedback FB microphone. The ear canal signal includes residual noise signals inside the ear canal. When the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the estimated SP filter is determined as the fixed SP filter. The adaptive SP filter is obtained through a digital signal processing (DSP) chip. The DSP chip acquires a real-time noise signal and obtains an estimated SP filter based on the audio source and the real-time noise signal. When the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range, the estimated SP filter is determined as the adaptive SP filter.

2. The TWS earphone according to claim 1, characterized in that, The audio signal processing path further includes: a feedback FB microphone and a feedback filter; wherein... The FB microphone is configured to pick up ear canal signals, which include residual noise signals inside the ear canal and the music or call voice. The SP filter is configured to input the audio source, process the audio source, and output a signal that is superimposed on the ear canal signal and then transmitted to the feedback filter. The feedback filter is configured to generate a signal for the noise reduction or pass-through processing, the noise reduction or pass-through processing signal being one of the superimposed signals used to generate the speaker drive signal.

3. The TWS earphone according to claim 2, characterized in that, The feedback FB microphone, the feedback filter, and the SP filter are configured in the codec.

4. The TWS earphone according to claim 1, characterized in that, The SP filter is configured to input the audio source, process the audio source, and output a signal that is one of the superimposed signals of the speaker drive signal.

5. The TWS earphone according to claim 4, characterized in that, The SP filter is located in the digital signal processing (DSP) chip.

6. The TWS earphone according to any one of claims 1-5, characterized in that, Also includes: First digital-to-analog converter (DAC); The input terminal of the first DAC is connected to the output terminal of the audio signal processing path, and the output terminal of the first DAC is connected to the input terminal of the frequency divider; The first DAC is configured to convert the speaker drive signal from digital form to analog form; Accordingly, the frequency divider is an analog frequency divider.

7. The TWS earphone according to any one of claims 1-5, characterized in that, Also includes: At least two second DACs; the input terminals of the at least two second DACs are all connected to the output terminal of the crossover, and the output terminals of the at least two second DACs are respectively connected to one of the at least two speakers; The second DAC is configured to convert one of the sub-audio signals of the at least two frequency bands from digital form to analog form; Accordingly, the frequency divider is a digital frequency divider.

8. The TWS earphone according to any one of claims 1-5, characterized in that, The primary operating frequency bands of the at least two speakers are not exactly the same.

9. The TWS earphone according to claim 8, characterized in that, The at least two loudspeakers include a dynamic loudspeaker and a balanced armature loudspeaker.

10. The TWS earphone according to claim 8, characterized in that, The at least two loudspeakers include a moving coil loudspeaker, a moving iron loudspeaker, a microelectromechanical system (MEMS) loudspeaker, and a planar diaphragm.

11. A playback method for true wireless stereo (TWS) earphones, characterized in that, The method is applied to the TWS earphones according to any one of claims 1-10; the method includes: Acquire an audio source, wherein the audio source is original music or call voice, or the audio source includes a voice signal that has been enhanced with human voice and the original music or call voice; The audio source is subjected to noise reduction or pass-through processing to obtain the speaker driving signal; The loudspeaker drive signal is divided into at least two frequency bands of sub-audio signals, wherein adjacent frequency bands of the at least two frequency bands partially overlap, or, adjacent frequency bands of the at least two frequency bands do not overlap; One of the sub-audio signals of the at least two frequency bands is played through at least two speakers respectively; The step of performing noise reduction or pass-through processing on the audio source to obtain the speaker drive signal includes: Obtain the fixed secondary path SP filter through the codec (CODEC). An estimated SP filter is obtained based on a pre-set speaker drive signal and an ear canal signal picked up by a feedback FB microphone, wherein the ear canal signal includes residual noise signals inside the ear canal. When the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the estimated SP filter is determined as the fixed SP filter. The audio source is processed using the fixed SP filter to obtain a filtered signal; The filtered signal is subjected to noise reduction or pass-through processing to obtain the speaker drive signal; or, An adaptive SP filter is obtained through a digital signal processing (DSP) chip. Acquire real-time noise signals; The estimated SP filter is obtained based on the audio source and the real-time noise signal; When the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range, the estimated SP filter is determined as the adaptive SP filter. The audio source is processed using the adaptive SP filter to obtain a filtered signal; The speaker drive signal is obtained by performing noise reduction or pass-through processing on the filtered signal.

12. The method according to claim 11, characterized in that, After obtaining the estimated SP filter based on the preset speaker drive signal and the ear canal signal picked up by the feedback FB microphone, the method further includes: When the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the parameters of the cascaded second-order filter are obtained according to the target frequency response of the estimated SP filter and the preset frequency division requirements. The SP cascaded second-order filter is obtained based on the parameters of the cascaded second-order filter, and the SP cascaded second-order filter is used as the fixed SP filter.

13. The method according to claim 11, characterized in that, The acquisition of real-time noise signals includes: The system acquires external signals picked up by the feedforward (FF) microphone and ear canal signals picked up by the feedback (FB) microphone. The external signals include external noise signals and the music or call voice. The ear canal signals include residual noise signals inside the ear canal and the music or call voice. Acquire the voice signal picked up by the main microphone; The real-time noise signal is obtained by subtracting the external signal and the ear canal signal from the speech signal.

14. The method according to any one of claims 11-13, characterized in that, The primary operating frequency bands of the at least two speakers are not exactly the same.

15. The method according to claim 14, characterized in that, The at least two loudspeakers include a dynamic loudspeaker and a balanced armature loudspeaker.

16. The method according to claim 14, characterized in that, The at least two loudspeakers include a moving coil loudspeaker, a moving iron loudspeaker, a microelectromechanical system (MEMS) loudspeaker, and a planar diaphragm.

17. A playback device for true wireless stereo (TWS) earphones, characterized in that, The device is applied to the TWS earphones according to any one of claims 1-10; the device comprises: An acquisition module is used to acquire an audio source, wherein the audio source is original music or call voice, or the audio source includes a voice signal that has been processed with human voice enhancement and the original music or call voice; The processing module is used to perform noise reduction or pass-through processing on the audio source to obtain a speaker driving signal; A frequency division module is used to divide the speaker drive signal into at least two frequency bands of sub-audio signals, wherein adjacent frequency bands of the at least two frequency bands partially overlap, or, adjacent frequency bands of the at least two frequency bands do not overlap; A playback module for playing one of the sub-audio signals of the at least two frequency bands through at least two speakers; Specifically, the processing module is used to obtain a fixed SP filter through a codec; process the audio source according to the fixed SP filter to obtain a filtered signal; and perform noise reduction or pass-through processing on the filtered signal to obtain the speaker driving signal. The processing module is specifically configured to obtain an estimated SP filter based on a pre-set speaker drive signal and an ear canal signal picked up by a feedback FB microphone. The ear canal signal includes residual noise signals inside the ear canal and the music or call audio. When the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, the estimated SP filter is determined as the fixed SP filter; or... The processing module is specifically used to obtain an adaptive SP filter through a digital signal processing (DSP) chip; process the audio source according to the adaptive SP filter to obtain a filtered signal; and perform noise reduction or pass-through processing on the filtered signal to obtain the speaker driving signal. The processing module is specifically used to acquire a real-time noise signal; acquire an estimated SP filter based on the audio source and the real-time noise signal; and determine the estimated SP filter as the adaptive SP filter when the difference between the signal obtained by the estimated SP filter and the real-time noise signal is within a set range.

18. The apparatus according to claim 17, characterized in that, The processing module is further configured to, when the difference between the signal obtained by the estimated SP filter and the ear canal signal is within a set range, obtain the parameters of the cascaded second-order filter according to the target frequency response of the estimated SP filter and the preset frequency division requirement; obtain the SP cascaded second-order filter according to the parameters of the cascaded second-order filter, and use the SP cascaded second-order filter as the fixed SP filter.

19. The apparatus according to claim 17, characterized in that, The processing module is specifically used to acquire external signals picked up by the feedforward (FF) microphone and ear canal signals picked up by the feedback (FB) microphone. The external signals include external noise signals and the music or call voice, and the ear canal signals include residual noise signals inside the ear canal and the music or call voice. The module also acquires the voice signal picked up by the main microphone, subtracts the external signals and the ear canal signals from the voice signal to obtain the signal difference, and obtains the estimated SP filter based on the audio source and the signal difference.

20. The apparatus according to any one of claims 17-19, characterized in that, The primary operating frequency bands of the at least two speakers are not exactly the same.

21. The apparatus according to claim 20, characterized in that, The at least two loudspeakers include a dynamic loudspeaker and a balanced armature loudspeaker.

22. The apparatus according to claim 20, characterized in that, The at least two loudspeakers include a moving coil loudspeaker, a moving iron loudspeaker, a microelectromechanical system (MEMS) loudspeaker, and a planar diaphragm.

23. A computer-readable storage medium, characterized in that, Includes a computer program, which, when executed on a computer, causes the computer to perform the method of any one of claims 11-16.

24. A computer program, characterized in that, When the computer program is executed by a computer, it is used to perform the method according to any one of claims 11-16.

Citation Information

Patent Citations

  • Noise reduction device and method

    CN111836147A

  • Multi-vibration-unit TWS earphone with embedded frequency division circuit

    CN208850008U