Audio processing method and related apparatus

By integrating a microphone into the headphones to collect ambient sound and generate target or inverted signals, the problem of headphone sound leakage is solved, privacy protection and improved signal-to-noise ratio in the ear are achieved, and the complexity and high cost of cavity design are avoided.

WO2026113326A1PCT designated stage Publication Date: 2026-06-04HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-06-10
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Both air conduction and bone conduction headphones suffer from sound leakage at high volumes, posing a risk of user privacy breaches. Existing technologies rely on complex and costly cavity designs, which cannot meet the demands of rapid iteration.

Method used

By integrating a microphone into the headphones to collect ambient sound, generating an error signal, and generating a target signal based on the filter coefficients to cancel out sound leakage, or by generating an inverted signal through feedforward filter coefficients to physically cancel out sound leakage, the cavity design is avoided.

Benefits of technology

It effectively reduces sound leakage, protects user privacy, improves the signal-to-noise ratio in the ear, adapts to different environments, and reduces design complexity and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025100212_04062026_PF_FP_ABST
    Figure CN2025100212_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the field of audio processing. Disclosed are an audio processing method and a related apparatus. The method comprises: playing a downlink signal by means of a sound-emitting unit; collecting sound in an environment by means of at least one error microphone, so as to obtain at least one error signal; on the basis of the at least one error signal and at least one filter coefficient, determining a target signal; and playing the target signal by means of a first air-conduction speaker, so as to reduce sound leakage during the playback of the downlink signal by the sound-emitting unit. Since sound in an environment comprises sound leakage generated when a sound-emitting unit plays a downlink signal, the playback of a target signal generated on the basis of an error signal can cancel out the sound leakage during the playback of the downlink signal by the sound-emitting unit, thereby achieving the effect of reducing sound leakage.
Need to check novelty before this filing date? Find Prior Art

Description

Audio processing methods and related devices

[0001] This application claims priority to Chinese Patent Application No. 202411761208.5, filed on November 29, 2024, entitled "Audio Processing Method and Related Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of audio processing, and in particular to an audio processing method and related apparatus. Background Technology

[0003] Open-back headphones are wearable audio devices available in both ear-hook and clip-on styles. They are also categorized by their sound-producing unit: air conduction headphones and bone conduction headphones. Air conduction headphones transmit audio signals through a speaker and via the air, while bone conduction headphones transmit signals through bone conduction elements. These signals are then conducted to the cochlea via contact between the headphone shell and the ear's cartilage / skin, thus producing the sensation of hearing, without requiring air transmission.

[0004] However, at high volumes, the sound emitted by the speakers of air conduction headphones not only travels into the ear but also out through the air. Similarly, some of the sound emitted by the bone conduction vibrator of bone conduction headphones is also radiated out through the air. In other words, both air conduction and bone conduction headphones have a certain degree of sound leakage, which poses a risk of privacy breaches for users. Therefore, there is an urgent need for an audio processing method to reduce sound leakage and protect user privacy. Summary of the Invention

[0005] This application provides an audio processing method and related apparatus that can reduce sound leakage from headphones. The technical solution is as follows:

[0006] In a first aspect, an audio processing method is provided for use in headphones, the headphones including a first air-conducting speaker, a sound-generating unit, and at least one error microphone, the method comprising: playing a downlink signal through the sound-generating unit; acquiring ambient sound through the at least one error microphone to obtain at least one error signal; determining a target signal based on the at least one error signal and at least one filter coefficient; and playing the target signal through the first air-conducting speaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit.

[0007] By integrating at least one microphone into the headphones to acquire ambient sound in real time, an error signal is obtained. A target signal is then generated based on the error signal and at least one filter coefficient, and played through a first air conduction unit. Since ambient sound includes sound leakage generated when the sound-emitting unit plays the downlink signal, playing the target signal generated based on the error signal can cancel out the sound leakage during the downlink signal playback process, thus reducing sound leakage. Furthermore, compared to related techniques that reduce sound leakage through special cavities, the various methods provided in this application do not rely on cavity design, thus offering greater flexibility.

[0008] In one possible implementation, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0009] Based on the description of the above application scenarios, since bone conduction headphones generate hearing through bone conduction and do not rely on air conduction, users can hear audio clearly even in noisy environments. In contrast, air conduction headphones are still difficult to hear downlink audio at high volumes in noisy environments. Therefore, if the sound-generating unit of this application is a bone conduction vibrator, the vibration of the bone conduction vibrator is less likely to be masked by ambient sound, thus enabling users to hear audio clearly even in noisy environments, thereby improving the user experience of the headphones. Furthermore, combined with the subsequent processing steps of this application, it is possible to effectively reduce sound leakage outside the ear through bone conduction, enhance the signal-to-noise ratio inside the ear, and thus effectively improve the quality of audio playback.

[0010] In one possible implementation, the number of at least one error microphone is 1.

[0011] It should be noted that the filter coefficients in this application are actually a set of at least one coefficient. Each filter coefficient corresponds one-to-one with at least one filtering algorithm (also known as a filter). That is, each filter coefficient includes the coefficients corresponding to the filter. The number of coefficients corresponding to the filter can be one or more.

[0012] In one possible implementation, at least one filter coefficient includes at least one feedback filter coefficient, which corresponds one-to-one with at least one error microphone. In this case, based on the at least one feedback filter coefficient, the error signals acquired by the corresponding error microphones in at least one error signal can be processed to obtain at least one feedback inverted signal, which is opposite in phase to the corresponding error signal. Based on the at least one feedback inverted signal, the target signal is determined.

[0013] In one possible implementation, for any one of the at least one feedback filter coefficients, the error signal acquired by the error microphone corresponding to that feedback filter coefficient is processed to obtain the feedback inverted signal corresponding to that feedback filter coefficient. Processing each of the at least one feedback filter coefficients in the same manner yields at least one feedback inverted signal.

[0014] It should be noted that the above-mentioned at least one feedback filter coefficient corresponds one-to-one with at least one feedback filter (i.e., feedback filtering algorithm). Processing the error signal collected by the corresponding error microphone in the at least one error signal based on at least one feedback filter coefficient means that filtering is performed on the error signal collected by the corresponding error microphone in the at least one error signal based on the coefficients corresponding to the feedback filter.

[0015] There are several ways to determine the target signal based on at least one feedback inverted signal. Three of these methods will be introduced below.

[0016] The first implementation method involves superimposing at least one feedback inverted signal to obtain the target signal.

[0017] Since the ambient sound includes the leakage sound generated when the sound unit plays the downlink signal, the error signal in the environment collected by the error microphone can be filtered to generate a feedback inverse signal that is opposite in phase to the original error signal. By playing the feedback inverse signal, the leakage sound generated by the downlink signal can be canceled out when they meet, thereby reducing the leakage sound.

[0018] The second implementation method is to superimpose the downlink signal with at least one feedback inverted signal to obtain the target signal.

[0019] Since the target signal contains a downlink signal, playing this target signal through the first air-conduction speaker enables both the first air-conduction speaker and the sound-generating unit to play audio simultaneously, thus effectively improving the in-ear signal-to-noise ratio. Furthermore, since ambient sound includes leakage sound generated when the sound-generating unit plays the downlink signal, filtering the error signal collected by the error microphone generates a feedback inverted signal with the opposite phase to the original error signal. Playing this feedback inverted signal allows it to cancel out the leakage sound generated by the downlink signal when they meet, thereby reducing leakage sound. In summary, the target signal contains not only a downlink signal but also at least one feedback inverted signal for eliminating leakage sound. Playing this target signal through the first air-conduction speaker can enhance the downlink signal, improve the in-ear signal-to-noise ratio, and simultaneously eliminate leakage sound.

[0020] In the third implementation, at least one filter coefficient also includes a feedforward filter coefficient. In this case, the downlink signal is processed based on the feedforward filter coefficient to obtain a feedforward inverted signal, which is out of phase with the leakage sound in the environment. The feedforward inverted signal is superimposed with at least one feedback inverted signal to obtain the target signal.

[0021] It should be noted that the above feedforward filter coefficients correspond one-to-one with the feedforward filter (i.e., the feedforward filtering algorithm). Processing the downlink signal based on the feedforward filter coefficients means filtering the downlink signal based on the coefficients corresponding to the feedforward filter.

[0022] Feedback filtering algorithms generate a feedback inverse signal based on the error signal (i.e., sound leakage in the environment) collected by the error microphone, thus accurately canceling out sound leakage. Feedforward filtering algorithms, on the other hand, predict potential sound leakage based on the downlink signal (i.e., the original audio played by the speaker unit), generating a corresponding feedforward inverse signal. By combining feedback and feedforward active signal control methods, they can more flexibly adapt to different sound environments and headphone usage, thereby improving the sound leakage prevention effect.

[0023] In one possible implementation, at least one filter coefficient may be determined before determining the target signal based on at least one error signal and at least one filter coefficient.

[0024] In one possible implementation, each error microphone in at least one error microphone corresponds to a first path and a second path. The first path is the air propagation path from the sound-emitting unit to the corresponding error microphone, and the second path is the air propagation path from the first air-conducting loudspeaker to the corresponding error microphone. In this case, the frequency response of at least one filter coefficient can be determined based on the transfer function of the first path and the transfer function of the second path corresponding to each error microphone. Based on the frequency response of at least one filter coefficient, at least one filter coefficient is obtained according to a relevant time-frequency conversion algorithm.

[0025] In one possible implementation, before determining the target signal based on at least one error signal and at least one filter coefficient, a target leakage signal-to-noise ratio can also be determined based on at least one error signal, the target leakage signal-to-noise ratio being used to describe the magnitude relationship between leakage sound and noise in the environment; if the leakage signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold, then the step of determining the target signal based on at least one error signal and at least one filter coefficient is performed.

[0026] For any one of the at least one error signals, the leakage signal-to-noise ratio (SNR) can be determined based on that error signal using a relevant algorithm. Processing each of the at least one error signal in the same way yields at least one leakage SNR. In this case, the maximum, minimum, median, mode, or average of the at least one leakage SNR can be determined as the target leakage SNR.

[0027] If the target sound leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, it indicates that the sound leakage of the headphones is large. Therefore, subsequent sound leakage prevention steps need to be performed. Thus, subsequent steps can be performed to determine the target signal based on at least one error signal and at least one filter coefficient.

[0028] If the target signal-to-noise ratio (SNR) of the leaked sound is less than the SNR threshold, then at least one filter coefficient is adjusted so that the average absolute value of the filter coefficients corresponding to the adjusted at least one filter algorithm is less than the average absolute value of the filter coefficients before the adjustment, and the step of determining the target signal based on at least one error signal and at least one filter coefficient is executed; or, the step of determining the target signal based on at least one error signal and at least one filter coefficient is not executed.

[0029] If the target sound leakage signal-to-noise ratio (SNR) is less than the SNR threshold, it indicates that the headphone's sound leakage is relatively small. The filter coefficients can be adjusted to be smaller to reduce the sound leakage prevention effect. Therefore, at least one filter coefficient can be adjusted so that the average absolute value of the adjusted filter coefficient is less than the average absolute value of the unadjusted filter coefficient. Then, the step of determining the target signal based on at least one error signal and at least one filter coefficient can be performed. Alternatively, the subsequent sound leakage prevention function can be omitted entirely, i.e., the step of determining the target signal based on at least one error signal and at least one filter coefficient can be skipped.

[0030] In one possible implementation, the filter coefficients that need to be adjusted are determined from at least one filter coefficient to obtain at least one target filter coefficient. For each filter coefficient in the at least one target filter coefficient, if the filter coefficient is positive, the target value is subtracted from the filter coefficient, and if the filter coefficient is negative, the target value is added to the filter coefficient. In this way, the absolute value of the filter coefficient can be reduced.

[0031] In one possible implementation, the headphones also include at least one monitoring microphone; in this case, during the playback of the target signal through the first air conduction speaker, the sound played by the first air conduction speaker is collected by the at least one monitoring microphone to obtain at least one monitoring signal; if it is determined based on the at least one monitoring signal that there is an abnormality in the first air conduction speaker, the playback of the target signal through the first air conduction speaker is stopped, or at least one filter coefficient is cleared to zero.

[0032] Among the abnormal situations are uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space (i.e., the sound generated by the first loudspeaker has phenomena such as attenuation and divergence), and / or, the sound generated by the first air-conducting loudspeaker has a whistling phenomenon.

[0033] For the implementation of determining the abnormality of the first air-conducting loudspeaker based on at least one monitoring signal, please refer to relevant howling detection and audio divergence detection technologies; this application will not elaborate on this.

[0034] In one possible implementation, the at least one monitoring microphone can be a microphone independent of the at least one error microphone mentioned above. In another possible implementation, when there is only one error microphone, it can integrate the functions of the monitoring microphone; when there are multiple error microphones, any one or more of these error microphones can integrate the functions of the monitoring microphone. In other words, the at least one monitoring microphone and the at least one error microphone can operate independently, each performing its corresponding function. Of course, one or more of the at least one error microphones can also be reused to implement the functions of the monitoring microphone.

[0035] Secondly, an audio processing method is provided for use in headphones, the headphones including a first air-conducting loudspeaker and a sound-generating unit, the method including: playing a downlink signal through the sound-generating unit; processing the downlink signal based on a feedforward filtering coefficient to obtain a feedforward inverted signal; playing the feedforward inverted signal through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit, wherein the feedforward inverted signal is out of phase with the sound leakage.

[0036] By processing the downlink signal through feedforward filtering coefficients, a feedforward inverted signal can be generated. When this signal is played through the first air-conducting speaker, it physically cancels out the original sound leakage generated by the sound-generating unit, thus significantly reducing sound leakage. In other words, the various audio processing methods provided in this application all use the first air-conducting speaker to play the generated signal to eliminate sound leakage, effectively preventing others from hearing the content played through the headphones and protecting user privacy. Furthermore, compared to related technologies that reduce sound leakage through special cavities, the various methods provided in this application do not rely on cavity design, thus offering greater flexibility.

[0037] In one possible implementation, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0038] Based on the description of the above application scenarios, since bone conduction headphones generate hearing through bone conduction and do not rely on air conduction, users can hear audio clearly even in noisy environments. In contrast, air conduction headphones are still difficult to hear downlink audio at high volumes in noisy environments. Therefore, if the sound-generating unit of this application is a bone conduction vibrator, the vibration of the bone conduction vibrator is less likely to be masked by ambient sound, thus enabling users to hear audio clearly even in noisy environments, thereby improving the user experience of the headphones. Furthermore, combined with the subsequent processing steps of this application, it is possible to effectively reduce sound leakage outside the ear through bone conduction, enhance the signal-to-noise ratio inside the ear, and thus effectively improve the quality of audio playback.

[0039] It should be noted that the feedforward filter coefficients in this application are actually a set of at least one coefficient. The feedforward filter coefficients correspond to the feedforward filter algorithm (also known as the feedforward filter). In other words, the feedforward filter coefficients include the coefficients corresponding to the feedforward filter, and the number of these coefficients can be one or more.

[0040] In one possible implementation, the feedforward filter coefficients can be determined before processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal.

[0041] In one possible implementation, the headphones further include at least one leakage sampling point, each leakage sampling point corresponding to a first path and a second path. The first path is the air propagation path from the sound-generating unit to the corresponding leakage sampling point, and the second path is the air propagation path from the first air-conducting loudspeaker to the corresponding leakage sampling point. In this case, the frequency response of the feedforward filter coefficients can be determined based on the transfer function of the first path and the transfer function of the second path corresponding to each leakage sampling point. Based on the frequency response of the feedforward filter coefficients, the feedforward filter coefficients are obtained according to the relevant time-frequency conversion algorithm.

[0042] In one possible implementation, each of the at least one leaky sound sampling points is equipped with an error microphone. In another possible implementation, the number of at least one leaky sound sampling point is one.

[0043] In one possible implementation, before processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal, at least one error signal can be obtained by acquiring ambient sound through at least one error microphone. Based on the at least one error signal, the target leakage signal-to-noise ratio is determined, and the target sound energy is used to describe the relationship between the leakage sound and ambient noise in the environment. If the target leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, the step of processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal is performed.

[0044] For any one of the at least one error signals, the leakage signal-to-noise ratio (SNR) can be determined based on that error signal using a relevant algorithm. Processing each of the at least one error signal in the same way yields at least one leakage SNR. In this case, the maximum, minimum, median, mode, or average of the at least one leakage SNR can be determined as the target leakage SNR.

[0045] If the target sound leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, it indicates that the sound leakage of the headphones is large. Therefore, it is necessary to perform subsequent sound leakage prevention steps. Thus, the subsequent step of processing the downlink signal based on the feedforward filter coefficient to obtain the feedforward inverted signal can be performed.

[0046] If the target signal-to-noise ratio (SNR) of the leaky audio is less than the SNR threshold, the feedforward filter coefficients are adjusted so that the absolute value of the adjusted feedforward filter coefficients is less than the absolute value of the original feedforward filter coefficients. Then, the step of processing the downlink signal based on the feedforward filter coefficients to obtain a feedforward inverted signal is performed. Alternatively, the step of processing the downlink signal based on the feedforward filter coefficients to obtain a feedforward inverted signal is not performed.

[0047] If the target leakage signal-to-noise ratio (SNR) is less than the SNR threshold, it indicates that the headphone leakage is relatively small. The filter coefficient can be adjusted to a smaller value to reduce the leakage prevention effect. Therefore, the feedforward filter coefficient can be adjusted so that the absolute value of the adjusted feedforward filter coefficient is less than the absolute value of the original feedforward filter coefficient. Then, the step of processing the downlink signal based on the feedforward filter coefficient to obtain a feedforward inverted signal can be performed. Alternatively, the subsequent leakage prevention function can be omitted entirely, i.e., the step of processing the downlink signal based on the feedforward filter coefficient to obtain a feedforward inverted signal can be skipped.

[0048] In one possible implementation, if the feedforward filter coefficient is positive, the target value is subtracted from the filter coefficient; if the filter coefficient is negative, the target value is added to the filter coefficient. In this way, the absolute value of the filter coefficient can be reduced.

[0049] In one possible implementation, the headphones also include at least one monitoring microphone; in this case, during the process of playing a feedforward inverted signal through the first air conduction speaker, the sound played by the first air conduction speaker is collected by the at least one monitoring microphone to obtain at least one monitoring signal; if it is determined based on the at least one monitoring signal that there is an abnormality in the first air conduction speaker, the playback of the feedforward inverted signal through the first air conduction speaker is stopped, or the feedforward filter coefficient is cleared to zero.

[0050] Among the abnormal situations are uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space (i.e., the sound generated by the first loudspeaker has phenomena such as attenuation and divergence), and / or, the sound generated by the first air-conducting loudspeaker has a whistling phenomenon.

[0051] For the implementation of determining the abnormality of the first air-conducting loudspeaker based on at least one monitoring signal, please refer to relevant howling detection and audio divergence detection technologies; this application will not elaborate on this.

[0052] In one possible implementation, the at least one monitoring microphone can be a microphone independent of the at least one error microphone mentioned above. In another possible implementation, when there is only one error microphone, it can integrate the functions of the monitoring microphone; when there are multiple error microphones, any one or more of these error microphones can integrate the functions of the monitoring microphone. In other words, the at least one monitoring microphone and the at least one error microphone can operate independently, each performing its corresponding function. Of course, one or more of the at least one error microphones can also be reused to implement the functions of the monitoring microphone.

[0053] Thirdly, an audio processing apparatus is provided, which has the function of implementing the audio processing method described in the first aspect. The audio processing apparatus includes at least one module for implementing the audio processing method provided in the first aspect.

[0054] Fourthly, an audio processing apparatus is provided, which has the function of implementing the audio processing method described in the second aspect above. The audio processing apparatus includes at least one module for implementing the audio processing method provided in the second aspect above.

[0055] Fifthly, an earphone is provided, the earphone including a first air-conducting speaker, a sound-generating unit, at least one error microphone, and a processor, the processor being configured to implement the audio processing method provided in the first aspect above.

[0056] In one possible implementation, the headphones further include a memory for storing a computer program that performs the audio processing method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the audio processing method described in the first aspect.

[0057] In one possible implementation, the headset may also include a communication bus for establishing a connection between the processor and the memory.

[0058] In a sixth aspect, an earphone is provided, the earphone including a first air-conducting speaker, a sound-generating unit, and a processor, the processor being used to implement the audio processing method provided in the second aspect above.

[0059] In one possible implementation, the headphones further include a memory for storing a computer program that performs the audio processing method provided in the second aspect above. The processor is configured to execute the computer program stored in the memory to implement the audio processing method described in the second aspect above.

[0060] In one possible implementation, the headset may also include a communication bus for establishing a connection between the processor and the memory.

[0061] In a seventh aspect, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is run on a computer or processor, the computer or processor performs the steps of the audio processing method described in the first aspect, or performs the steps of the audio processing method described in the second aspect.

[0062] Eighthly, a computer program product is provided, the computer program product comprising computer instructions that, when executed on a computer or processor, cause the computer to perform the steps of the audio processing method described in the first aspect, or to perform the steps of the audio processing method described in the second aspect. Alternatively, a computer program is provided that, when executed on a computer or processor, causes the computer or processor to perform the steps of the audio processing method described in the first aspect, or to perform the steps of the audio processing method described in the second aspect.

[0063] The technical effects achieved by the third to eighth aspects mentioned above are similar to those achieved by the corresponding technical means in the first and second aspects, and will not be repeated here. Attached Figure Description

[0064] Figure 1 is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0065] Figure 2 is a schematic diagram of the hardware structure of an earphone provided in an embodiment of this application;

[0066] Figure 3 is a schematic diagram of the hardware structure of another type of earphone provided in an embodiment of this application;

[0067] Figure 4 is a schematic diagram of the hardware structure of another type of earphone provided in an embodiment of this application;

[0068] Figure 5 is a schematic diagram of the hardware structure of another type of earphone provided in an embodiment of this application;

[0069] Figure 6 is a schematic diagram of another type of earphone provided in an embodiment of this application;

[0070] Figure 7 is a flowchart of an audio processing method provided in an embodiment of this application;

[0071] Figure 8 is a schematic diagram of an earphone provided in an embodiment of this application;

[0072] Figure 9 is a schematic diagram of an earphone provided in an embodiment of this application;

[0073] Figure 10 is a schematic diagram of an earphone provided in an embodiment of this application;

[0074] Figure 11 is a schematic diagram of an earphone provided in an embodiment of this application;

[0075] Figure 12 is a flowchart of another audio processing method provided in an embodiment of this application;

[0076] Figure 13 is a schematic diagram of an earphone provided in an embodiment of this application;

[0077] Figure 14 is a schematic diagram of an earphone provided in an embodiment of this application;

[0078] Figure 15 is a schematic diagram of an audio processing device provided in an embodiment of this application;

[0079] Figure 16 is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application. Detailed Implementation

[0080] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0081] To facilitate understanding, before providing a detailed explanation of the audio processing method provided in the embodiments of this application, the terms, application scenarios, and implementation environments involved in the embodiments of this application will be introduced first.

[0082] First, the terms used in the embodiments of this application will be introduced.

[0083] Sound leakage: While the sound unit emits sound towards the ear canal, it also emits sound in other directions. In this process, the directional sounds that are easily heard by others (such as sounds that are directed towards the outside of the ear and not perpendicular to the ground) are called sound leakage.

[0084] Bone conduction: refers to the process by which a vibrating unit drives the shell of an earphone device. The shell contacts the human skin (such as soft tissue, skull, etc.), and sound is transmitted through the vibration of the shell to the human bones and then to the auditory nerve in the cochlea. This entire transmission path is called the bone conduction path, and the vibrating unit can be called a bone conduction unit or bone conduction oscillator.

[0085] Air conduction: also known as air transmission, refers to the process by which sound travels through the air to the ear canal and then to the eardrum. This entire transmission path is called the air conduction path, and the loudspeaker (also known as a horn) that produces sound can be called an air conduction unit.

[0086] Uplink and downlink signals: In the audio field, uplink and downlink signals are typically used to describe the transmission direction of audio data.

[0087] Uplink signals refer to audio data transmitted from a terminal with audio capture capabilities (such as a mobile phone, microphone, or other audio recording device) to a server or other devices (such as a remote communication peer). For example, in a voice call scenario, a user's voice is captured through a microphone and sent as an uplink signal to the other party's device. In a video conferencing scenario, a user's audio input (such as speaking or discussing) is transmitted as an uplink signal to other participants in the meeting. In a speech recognition scenario, a user's voice is sent as an uplink signal to a server for recognition processing. Uplink signals typically involve the audio recording and encoding process to ensure that audio data can be transmitted efficiently and accurately.

[0088] Downlink signals refer to audio data sent from a server or other device (such as a remote communication peer, audio playback device, etc.) to a terminal with audio playback capabilities. For example, in a music playback scenario, the server sends music data as a downlink signal to the user's audio playback device. In a voice call scenario, the other party's voice is received and played by the user's audio playback device as a downlink signal. In an online education scenario, the teacher's audio lecture is transmitted to students as a downlink signal for them to listen to. Downlink signals typically involve the audio decoding and playback process to ensure that users can hear the audio content clearly.

[0089] Frequency response, also known as the frequency response function, is a mathematical tool used to describe the relationship between the input and output of a linear time-invariant system at different frequencies. It is a complex function, usually denoted by H(jω), where ω is the angular frequency and j is the imaginary unit. The frequency response function can provide information about the system's behavior in the frequency domain, including its gain, phase, and resonant frequencies.

[0090] A transfer function is a complex function describing the relationship between the input and output of a linear time-invariant system, usually denoted by H(s), where s is a complex frequency variable, which can be expressed as s = jω. Under zero initial conditions, the transfer function is the ratio of the Laplace transform (or z-transform) of the system response (i.e., the output) to the Laplace transform of the excitation (i.e., the input). The transfer function is a characteristic of the system's behavior in the frequency domain, encompassing both its dynamic and steady-state characteristics. It is one of the main tools for studying classical control theory. In the audio domain, the transfer function of a path describes the dynamic characteristics of an audio signal during propagation, including signal attenuation and phase changes. These characteristics are crucial for the transmission quality of audio signals; that is, the transfer function can be used to predict the output response of the path under different input signals, thereby adjusting system parameters to meet specific performance requirements.

[0091] The application scenarios involved in the embodiments of this application will be introduced next.

[0092] Open-back headphones are wearable audio devices available in both ear-hook and clip-on styles. They are also categorized by their sound-producing unit: air conduction headphones and bone conduction headphones. Air conduction headphones transmit audio signals through a speaker and via the air, while bone conduction headphones transmit signals through bone conduction elements. These signals are then conducted to the cochlea via contact between the headphone shell and the ear's cartilage / skin, thus producing the sensation of hearing, without requiring air transmission.

[0093] In noisy environments, even at high volumes, air-conduction headphones often struggle to clearly hear downstream audio. Furthermore, at higher volumes, sound from the speakers in air-conduction headphones travels not only into the ear but also outwards through the air, resulting in significant sound leakage. While bone conduction headphones, which generate sound through bone conduction and do not rely on air conduction, allow users to hear audio clearly even in noisy environments, some sound emitted by the bone conduction vibrators also radiates outwards through the air, still resulting in sound leakage. In other words, both air-conduction and bone conduction headphones exhibit some degree of sound leakage, posing a risk of user privacy breaches.

[0094] To address the issue of sound leakage, a special type of earphone with a unique cavity has been designed. This cavity is shaped to hang on both sides of the ear, with the speaker placed in the middle. The cavity is divided into two parts, allowing for bidirectional sound output when sound is emitted. These two sound sources automatically reverse phase after passing through the front and rear cavities, forming a dipole sound source—a synthesized sound source composed of two sound sources with opposite phases. This achieves directional sound emission and reduces sound leakage in specific directions.

[0095] However, for the above solution to achieve the desired effect, the earphone cavity needs to fit snugly against both sides of the ear. This places strict requirements on the shape and size of the earphone. If the shape or size of the earphone is not suitable for some users' earlobes, it may lead to discomfort or an ineffective fit, thus affecting the sound leakage reduction effect. In addition, this special cavity design may require more complex manufacturing processes and higher material costs, thus increasing the overall cost of the earphone. From a research and development perspective, since this solution relies on the precise design of the cavity structure to achieve anti-phase radiation and directional sound emission, a large number of tests are required to optimize the structural layout. After each prototyping and testing, the structure may need to be fine-tuned, resulting in a long design iteration cycle and high design costs, which cannot meet the needs of rapid iteration and rapid market launch.

[0096] Based on this, embodiments of this application provide various audio processing methods. In one such method, at least one microphone is integrated into the headphones to collect ambient sound in real time, obtaining an error signal. A target signal is then generated based on the error signal and at least one filter coefficient, and played through a first air-conducting unit. Since ambient sound includes leakage sound generated when the sound-emitting unit plays the downlink signal, playing the target signal generated based on the error signal can cancel out the leakage sound during the downlink signal playback process, thus reducing leakage sound. In another audio processing method, the downlink signal is processed using feedforward filter coefficients to generate a feedforward inverted signal. When this feedforward inverted signal is played through the first air-conducting speaker, it physically cancels out the original leakage sound generated by the sound-emitting unit, significantly reducing leakage sound. In other words, the various audio processing methods provided in this application all use a first air-conducting speaker to play the generated signal to eliminate leakage sound, effectively preventing others from hearing the content played in the headphones and protecting user privacy. Furthermore, compared to related technologies that reduce sound leakage through special cavities, the various methods provided in this application embodiment do not depend on cavity design, thus offering greater flexibility.

[0097] The implementation environment involved in the embodiments of this application will be described next.

[0098] Please refer to Figure 1, which is a schematic diagram of an implementation environment provided in an embodiment of this application. This implementation environment includes an earphone 01 and a terminal 02. The earphone 01 and the terminal 02 are connected via a wired or wireless means for communication. For example, the earphone 01 and the terminal 02 communicate via Bluetooth, or via any other communication protocol; this embodiment of the application does not limit this to any particular protocol.

[0099] The earphone 01 and the terminal 02 can transmit audio signals and control signals. For example, the terminal 02 sends audio signals such as music or voice (also known as downlink signals) to the earphone 01 for playback. Or, the terminal 02 sends control signals to the earphone 01 to control whether the earphone 01's anti-leakage function is enabled (i.e. whether to perform audio processing using the method provided in this application embodiment to achieve the anti-leakage effect), and so on.

[0100] Terminal 02 can be a mobile phone, computer (such as a laptop, desktop computer, handheld tablet, or in-vehicle tablet), etc. This application does not limit the type or structure of terminal 02.

[0101] In one possible implementation, the earphone 01 provided in this application embodiment can be wired or wireless. Furthermore, in terms of wearing method, the earphone 01 provided in this application embodiment can be a neckband type, ear hook / ear clip type, true wireless stereo (TWS), etc. In terms of appearance, the earphone 01 provided in this application embodiment can be in-ear, semi-open, open, over-ear, etc. This application embodiment does not limit the communication method, wearing method, or appearance of the earphone.

[0102] The hardware structure of an earphone provided in this application embodiment will be described next in conjunction with the wearing form of the earphone in the human ear. For example, please refer to Figure 2. Figure 2 is a schematic diagram of the hardware structure of an earphone provided in this application embodiment. The earphone includes a first air conduction speaker, a sound generating unit, at least one error microphone (the at least one error microphone is schematically represented by two error microphones in Figure 2) and a processor (not shown in Figure 2).

[0103] The sound-generating unit is used to play downlink signals (such as music, voice and other audio signals), at least one error microphone is used to collect ambient sound to obtain at least one error signal, the processor is used to determine the target signal based on at least one error signal and at least one filter coefficient, and the first air-conducting loudspeaker is used to play the target signal to reduce sound leakage during the downlink signal playback process of the sound-generating unit.

[0104] It should be noted that the filter coefficients in the embodiments of this application are actually a set of at least one coefficient. At least one filter coefficient corresponds one-to-one with at least one filter algorithm (also known as a filter). That is, each filter coefficient includes the coefficients corresponding to the corresponding filter. The number of coefficients corresponding to the filter can be one or more.

[0105] In some embodiments, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0106] In some embodiments, the first air-conducting loudspeaker can be placed facing inwards or outwards from the ear, or in other directions, and this application does not limit this.

[0107] In one possible implementation, the headphones further include a housing, comprising a front housing and a rear housing. The front housing is the portion of the housing that is close to (or in contact with) the skin and ear canal when the headphones are worn by the user, while the rear housing is the portion of the housing that is away from the skin and ear canal when the headphones are worn by the user. In this case, at least one error microphone can be deployed on the rear housing facing outwards to collect ambient sound. In this embodiment, the sound signal collected by the error microphone is referred to as an error signal.

[0108] In one possible implementation, the at least one filter coefficient includes at least one feedback (FB) filter coefficient, which corresponds one-to-one with at least one error microphone. In this case, based on the at least one feedback filter coefficient, the error signals acquired by the corresponding error microphones in the at least one error signal are processed to obtain at least one feedback inverted signal, which is out of phase with the corresponding error signal. Based on the at least one feedback inverted signal, the target signal is determined.

[0109] It should be noted that processing the error signals acquired by the corresponding error microphones in at least one error signal based on at least one feedback filter coefficient means filtering the error signals acquired by the corresponding error microphones in at least one error signal based on the coefficients corresponding to the feedback filter.

[0110] In one possible implementation, at least one filter coefficient also includes a feedforward (FF) filter coefficient; that is, at least one filter coefficient includes not only at least one feedback filter coefficient but also a feedforward filter coefficient. In this case, the downlink signal is processed based on the filter coefficient corresponding to the feedforward filter coefficient to obtain a feedforward inverted signal, which is out of phase with the leaked sound in the environment. The feedforward inverted signal is then superimposed with at least one feedback inverted signal to obtain the target signal.

[0111] It should be noted that processing the downlink signal based on the filter coefficients corresponding to the feedforward filter coefficients means filtering the downlink signal based on the coefficients corresponding to the feedforward filter.

[0112] In some embodiments, the number of the at least one error microphone is 1.

[0113] In some other embodiments, please refer to FIG3, which is a schematic diagram of the hardware structure of another type of earphone provided in the embodiments of this application. The earphone also includes at least one monitoring microphone (FIG3 schematically represents the at least one monitoring microphone). During the process of playing the target signal through the first air conduction speaker, the sound played by the first air conduction speaker is collected through the at least one monitoring microphone to obtain at least one monitoring signal. If it is determined based on the at least one monitoring signal that there is an abnormality in the first air conduction speaker, the playback of the target signal through the first air conduction speaker is stopped, or the coefficients corresponding to at least one filtering algorithm are cleared to zero (that is, at least one filtering coefficient is cleared to zero).

[0114] In one possible implementation, at least one detection microphone can be positioned in the front shell facing the ear canal and close to the first air conduction speaker to determine whether there is any abnormality in the signal output by the first air conduction speaker.

[0115] In some embodiments, the at least one monitoring microphone may be a microphone independent of the at least one error microphone described above. In other embodiments, when there is only one error microphone, the error microphone may integrate the functions of the monitoring microphone. When there are multiple error microphones, any one or more of the multiple error microphones may integrate the functions of the monitoring microphone. In other words, the at least one monitoring microphone and the at least one error microphone can be independent of each other, each performing its corresponding function. Of course, one or more of the at least one error microphones can also be reused to implement the functions of the monitoring microphone. This application does not limit this aspect.

[0116] It should be noted that the aforementioned at least one feedback filter coefficient corresponds one-to-one with at least one feedback filter, and the aforementioned feedforward filter coefficient corresponds to a feedforward filter. In some embodiments, the feedforward filtering algorithm can be implemented by a device independent of the processor, and the at least one feedback filtering algorithm can also be implemented by a device or circuit independent of the processor. Of course, the feedforward filtering algorithm and the at least one feedback filtering algorithm can also be implemented by the processor executing relevant software code, and this application embodiment does not limit this. When the feedforward filtering algorithm and the at least one feedback filtering algorithm are implemented by the processor, at least one monitoring signal collected by at least one monitoring microphone, at least one error signal collected by at least one error microphone, and the downlink signal can be used as inputs to the processor, and then the processor can process them according to the corresponding processing method. For example, the target signal can be determined based on at least one error signal and the filter coefficients corresponding to at least one filtering algorithm, and then the target signal can be played through the first air conduction speaker.

[0117] For example, the processor can be an active noise cancellation (ANC) chip (or ANC module), but this application embodiment does not limit this. Please refer to Figure 4, which is a schematic diagram of the hardware structure of another type of earphone provided in this application embodiment. The processor is an ANC module, and the above-mentioned feedforward filtering algorithm and at least one feedback filtering algorithm are implemented through the ANC module.

[0118] In some embodiments, the headphones may also include other components, such as a proximity sensor, for detecting whether the headphones are in the ear. If the headphones are wireless, they may also include a wireless communication module, which may be a Wi-Fi module or a Bluetooth module. This wireless communication module is used for the headphones to communicate with other devices.

[0119] It is understood that the structures illustrated in the embodiments of this application do not constitute a limitation on the headphones. In other embodiments, the headphones may include more or fewer components than illustrated, or combine some components, or separate some components, or have different component arrangements. The components illustrated can be implemented in hardware, software, or a combination of hardware and software.

[0120] The hardware structure of another type of headphone provided in this application embodiment will be described next in conjunction with the wearing form of the headphone in the human ear. For example, please refer to Figure 5. Figure 5 is a schematic diagram of the hardware structure of another type of headphone provided in this application embodiment. The headphone includes a first air conduction speaker, a sound generating unit and a processor (not shown in Figure 5).

[0121] The sound-generating unit is used to play downlink signals (such as music, voice and other audio signals), and the processor is used to process the downlink signals based on the feedforward filter coefficients to obtain a feedforward inverted signal. The feedforward inverted signal is opposite in phase to the leakage sound. The first air-conducting loudspeaker is used to play the feedforward inverted signal to reduce the leakage sound during the downlink signal playback process of the sound-generating unit.

[0122] It should be noted that the feedforward filter coefficients in this embodiment are actually a set of at least one coefficient. The feedforward filter coefficients correspond to a feedforward filtering algorithm (also known as a feedforward filter), meaning that the feedforward filter coefficients include the coefficients corresponding to the feedforward filter. The number of coefficients corresponding to the feedforward filter can be one or more. Processing the downlink signal based on the feedforward filter coefficients means filtering the downlink signal based on the coefficients corresponding to the feedforward filter.

[0123] In some embodiments, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0124] In some embodiments, the first air-conducting loudspeaker can be placed facing inwards or outwards from the ear, or in other directions, and this application does not limit this.

[0125] In other embodiments, the headphones further include at least one monitoring microphone. During the playback of the target signal through the first air conduction speaker, the sound played by the first air conduction speaker is collected by the at least one monitoring microphone to obtain at least one monitoring signal. If it is determined based on the at least one monitoring signal that there is an abnormality in the first air conduction speaker, the playback of the target signal through the first air conduction speaker is stopped, or the filter coefficients corresponding to at least one filtering algorithm are cleared to zero.

[0126] In one possible implementation, the headphones further include a housing, comprising a front housing and a rear housing. The front housing is the portion of the housing that is close to (or in contact with) the skin and ear canal when the headphones are worn by the user, while the rear housing is the portion of the housing that is away from the skin and ear canal when the headphones are worn by the user. In this case, at least one detection microphone can be positioned on the front housing facing the ear canal and close to the first air conduction speaker to determine whether there are any abnormalities in the signal output by the first air conduction speaker.

[0127] In some embodiments, the feedforward filtering algorithm described above can be implemented by a device independent of the processor. Of course, the feedforward filtering algorithm can also be implemented by the processor executing relevant software code, and this application embodiment does not limit this. When the feedforward filtering algorithm is implemented by a processor, at least one monitoring signal collected by at least one monitoring microphone and a downlink signal can be used as inputs to the processor, and then the processor can process them according to the corresponding processing method. For example, the downlink signal can be processed based on the filtering coefficients of the feedforward filtering algorithm to obtain a feedforward inverted signal, and then the feedforward inverted signal can be played through the first air-conducting loudspeaker.

[0128] For example, the processor may be an ANC chip (or ANC module), but this application embodiment does not limit this.

[0129] In some embodiments, the headphones may also include other components, such as a proximity sensor, for detecting whether the headphones are in the ear. If the headphones are wireless, they may also include a wireless communication module, which may be a Wi-Fi module or a Bluetooth module. This wireless communication module is used for the headphones to communicate with other devices.

[0130] It is understood that the structures illustrated in the embodiments of this application do not constitute a limitation on the headphones. In other embodiments, the headphones may include more or fewer components than illustrated, or combine some components, or separate some components, or have different component arrangements. The components illustrated can be implemented in hardware, software, or a combination of hardware and software.

[0131] It should be noted that the feedforward filtering algorithm in the embodiments of this application can also be called a feedforward filter, and the feedback filtering algorithm can also be called a feedback filter. The embodiments of this application do not limit this.

[0132] Please refer to Figure 6, which is a schematic diagram of another type of headset provided according to an embodiment of this application. The headset can be any of the headsets shown in Figures 1 to 5. The headset includes at least one processor 601, a communication bus 602, a memory 603, and at least one communication interface 604.

[0133] Processor 601 can be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solutions of this application, such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0134] The communication bus 602 is used to transmit information between the aforementioned components. The communication bus 602 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus.

[0135] The memory 603 may be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compressed optical disc, a laser disc, a digital versatile optical disc, a Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but not limited thereto. The memory 603 may exist independently and be connected to the processor 601 via a communication bus 602. The memory 603 may also be integrated with the processor 601.

[0136] Communication interface 604 uses any transceiver-like device for communicating with other devices or communication networks. Communication interface 604 includes a wired communication interface and may also include a wireless communication interface. The wired communication interface may be, for example, an Ethernet interface. The Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.

[0137] In a specific implementation, as one example, processor 601 may include one or more CPUs, such as CPU0 and CPU1 as shown in FIG6.

[0138] In a specific implementation, as one embodiment, the headphones may include multiple processors, such as processor 601 and processor 605 as shown in FIG6. Each of these processors may be a single-core processor or a multi-core processor. Here, "processor" may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0139] In a specific implementation, as one embodiment, the headphones may further include an output device 606 and an input device 607. The output device 606 communicates with the processor 601 and can display information in various ways. For example, the output device 606 may be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 607 communicates with the processor 601 and can receive user input in various ways. For example, the input device 607 may be a mouse, a keyboard, a touchscreen device, or a sensing device, etc.

[0140] In this embodiment, the output device 606 includes a sound-generating unit and a first air-conducting loudspeaker. The sound-generating unit is used to play a downlink signal, and the first air-conducting loudspeaker is used to play a target signal or a feedforward inverted signal obtained by the audio processing method provided in this embodiment. In some embodiments, the sound-generating unit may be a bone conduction vibrator or a second air-conducting loudspeaker. In some embodiments, the input device 607 includes at least one error microphone, which is used to collect ambient sound to obtain at least one error signal.

[0141] In some embodiments, memory 603 is used to store program code 610 for executing the scheme of this application, and processor 601 can execute program code 610 stored in memory 603. The program code 610 may include one or more software modules, and the headphones can implement the audio processing method provided in the embodiments of FIG7 and FIG12 below through processor 601 and program code 610 in memory 603.

[0142] Those skilled in the art should understand that the above-described headphones and terminals are merely examples, and other existing or future headphones and terminals that are applicable to the embodiments of this application should also be included within the scope of protection of the embodiments of this application, and are hereby incorporated by reference.

[0143] It should be noted that the application scenarios and implementation environments described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0144] Figure 7 is a flowchart of an audio processing method provided in an embodiment of this application. Optionally, the method shown in Figure 7 is applied to the headphones shown in Figure 2, the headphones including a first air-conducting speaker, a sound-emitting unit, and at least one error microphone. Referring to Figure 7, the method includes the following steps.

[0145] Step 701: Play downlink signal through sound unit.

[0146] In some embodiments, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0147] Based on the description of the above application scenarios, since bone conduction headphones generate hearing through bone conduction and do not rely on air conduction, users can hear audio clearly even in noisy environments. In contrast, air conduction headphones are still difficult to hear at high volumes in noisy environments. Therefore, if the sound-generating unit in this application embodiment is a bone conduction vibrator, the vibration of the bone conduction vibrator is less likely to be masked by ambient sound, thus enabling users to hear audio clearly even in noisy environments, thereby improving the user experience of the headphones. Based on this, combined with the subsequent processing steps of this application embodiment, it is possible to effectively reduce sound leakage outside the ear through bone conduction, enhance the signal-to-noise ratio inside the ear, and thus effectively improve the quality of audio playback.

[0148] Step 702: Acquire ambient sound using at least one error microphone to obtain at least one error signal.

[0149] In some embodiments, the number of at least one error microphone is 1.

[0150] Step 703: Determine the target signal based on at least one error signal and at least one filter coefficient.

[0151] It should be noted that the filter coefficients in the embodiments of this application are actually a set of at least one coefficient. At least one filter coefficient corresponds one-to-one with at least one filter algorithm (also known as a filter). That is, each filter coefficient includes the coefficients corresponding to the corresponding filter. The number of coefficients corresponding to the filter can be one or more.

[0152] In one possible implementation, at least one filter coefficient includes at least one feedback filter coefficient, which corresponds one-to-one with at least one error microphone. In this case, based on the at least one feedback filter coefficient, the error signals acquired by the corresponding error microphones in at least one error signal can be processed to obtain at least one feedback inverted signal, which is opposite in phase to the corresponding error signal. Based on the at least one feedback inverted signal, the target signal is determined.

[0153] In some embodiments, for any one of the at least one feedback filter coefficients, the error signal acquired by the error microphone corresponding to that feedback filter coefficient is processed to obtain the feedback inverted signal corresponding to that feedback filter coefficient. Processing each of the at least one feedback filter coefficients in the same manner yields at least one feedback inverted signal.

[0154] It should be noted that the above-mentioned at least one feedback filter coefficient corresponds one-to-one with at least one feedback filter (i.e., feedback filtering algorithm). Processing the error signal collected by the corresponding error microphone in the at least one error signal based on at least one feedback filter coefficient means that filtering is performed on the error signal collected by the corresponding error microphone in the at least one error signal based on the coefficients corresponding to the feedback filter.

[0155] There are several ways to determine the target signal based on at least one feedback inverted signal. Three of these methods will be introduced below.

[0156] The first implementation method involves superimposing at least one feedback inverted signal to obtain the target signal.

[0157] Since the ambient sound includes the leakage sound generated when the sound unit plays the downlink signal, the error signal in the environment collected by the error microphone can be filtered to generate a feedback inverse signal that is opposite in phase to the original error signal. By playing the feedback inverse signal, the leakage sound generated by the downlink signal can be canceled out when they meet, thereby reducing the leakage sound.

[0158] For example, please refer to Figure 8, which is a schematic diagram of an earphone provided in an embodiment of this application. The earphone includes a first air-conducting speaker, a sound-generating unit (i.e., a bone conduction vibrator in Figure 8), and two error microphones, namely error microphone 1 and error microphone 2. Error microphone 1 corresponds to feedback filter 1, and error microphone 2 corresponds to feedback filter 2. Based on the filtering coefficients corresponding to feedback filter 1, the error signal collected by error microphone 1 is processed by feedback filter 1. Based on the filtering coefficients corresponding to feedback filter 2, the error signal collected by error microphone 2 is processed by feedback filter 2. The signals processed by feedback filters 1 and 2 are superimposed to obtain the target signal.

[0159] The second implementation method is to superimpose the downlink signal with at least one feedback inverted signal to obtain the target signal.

[0160] Since the target signal contains a downlink signal, playing this target signal through the first air-conduction speaker enables both the first air-conduction speaker and the sound-generating unit to play audio simultaneously, thus effectively improving the in-ear signal-to-noise ratio. Furthermore, since ambient sound includes leakage sound generated when the sound-generating unit plays the downlink signal, filtering the error signal collected by the error microphone generates a feedback inverted signal with the opposite phase to the original error signal. Playing this feedback inverted signal allows it to cancel out the leakage sound generated by the downlink signal when they meet, thereby reducing leakage sound. In summary, the target signal contains not only a downlink signal but also at least one feedback inverted signal for eliminating leakage sound. Playing this target signal through the first air-conduction speaker can enhance the downlink signal, improve the in-ear signal-to-noise ratio, and simultaneously eliminate leakage sound.

[0161] For example, please refer to Figure 9, which is a schematic diagram of an earphone provided in an embodiment of this application. The earphone includes a first air-conducting speaker, a sound-generating unit (i.e., the bone conduction vibrator in Figure 9), and two error microphones, namely error microphone 1 and error microphone 2. Two feedback filtering algorithms are respectively feedback filter 1 and feedback filter 2. Error microphone 1 corresponds to feedback filter 1, and error microphone 2 corresponds to feedback filter 2. Based on the filtering coefficients corresponding to feedback filter 1, the error signal collected by error microphone 1 is processed by feedback filter 1. Based on the filtering coefficients corresponding to feedback filter 2, the error signal collected by error microphone 2 is processed by feedback filter 2. The downlink signal, the signal processed by feedback filter 1, and the signal processed by feedback filter 2 are superimposed to obtain the target signal.

[0162] In the third implementation, at least one filter coefficient also includes a feedforward filter coefficient. In this case, the downlink signal is processed based on the feedforward filter coefficient to obtain a feedforward inverted signal, which is out of phase with the leakage sound in the environment. The feedforward inverted signal is superimposed with at least one feedback inverted signal to obtain the target signal.

[0163] It should be noted that the above feedforward filter coefficients correspond one-to-one with the feedforward filter (i.e., the feedforward filtering algorithm). Processing the downlink signal based on the feedforward filter coefficients means filtering the downlink signal based on the coefficients corresponding to the feedforward filter.

[0164] Feedback filtering algorithms generate a feedback inverse signal based on the error signal (i.e., sound leakage in the environment) collected by the error microphone, thus accurately canceling out sound leakage. Feedforward filtering algorithms, on the other hand, predict potential sound leakage based on the downlink signal (i.e., the original audio played by the speaker unit), generating a corresponding feedforward inverse signal. By combining feedback and feedforward active signal control methods, they can more flexibly adapt to different sound environments and headphone usage, thereby improving the sound leakage prevention effect.

[0165] For example, please refer to Figure 10, which is a schematic diagram of an earphone provided in an embodiment of this application. The earphone includes a first air-conducting speaker, a sound-generating unit (i.e., the bone conduction vibrator in Figure 10), and two error microphones, namely error microphone 1 and error microphone 2. Error microphone 1 corresponds to feedback filter 1, and error microphone 2 corresponds to feedback filter 2. Based on the filtering coefficients corresponding to the feedforward filter, the downlink signal is processed by the feedforward filter to obtain a feedforward inverted signal. Based on the filtering coefficients corresponding to the feedback filter 1, the error signal collected by error microphone 1 is processed by the feedback filter 1. Based on the filtering coefficients corresponding to the feedback filter 2, the error signal collected by error microphone 2 is processed by the feedback filter 2. The feedforward inverted signal, the signal processed by feedback filter 1, and the signal processed by feedback filter 2 are superimposed to obtain the target signal.

[0166] In some embodiments, at least one filter coefficient may be determined before determining the target signal based on at least one error signal and at least one filter coefficient.

[0167] In one possible implementation, each error microphone in at least one error microphone corresponds to a first path and a second path. The first path is the air propagation path from the sound-emitting unit to the corresponding error microphone, and the second path is the air propagation path from the first air-conducting loudspeaker to the corresponding error microphone. In this case, the frequency response of at least one filter coefficient can be determined based on the transfer function of the first path and the transfer function of the second path corresponding to each error microphone. Based on the frequency response of at least one filter coefficient, at least one filter coefficient is obtained according to a relevant time-frequency conversion algorithm.

[0168] The first and second paths described above will be further explained by example. Please refer to Figure 11, which is a schematic diagram of an earphone provided in an embodiment of this application. The earphone includes a first air conduction speaker, a sound-generating unit (i.e., the bone conduction vibrator in Figure 11), two error microphones, and two feedback filtering algorithms, namely error microphone 1 and error microphone 2. The first path corresponding to error microphone 1 is FP11, the second path corresponding to error microphone 1 is FP21, the first path corresponding to error microphone 2 is FP12, and the second path corresponding to error microphone 1 is FP22.

[0169] Depending on how the target signal is determined, the methods for determining the frequency response of at least one filter coefficient will also differ, and these will be described separately below.

[0170] When the target signal is determined by the first implementation method, the at least one filter coefficient includes at least one feedback filter coefficient. The frequency response of each of the at least one filter coefficients satisfies the following formulas (1)-(2). |1+FB i ·FP 2i |>1 (2)

[0171] In the above formula (1), noise refers to the noise spectrum in the environment, DL refers to the signal spectrum played through the sound unit, and FP... 1i This refers to the transfer function of the first path corresponding to the i-th error microphone, Err. i It refers to the error signal spectrum acquired by the i-th error microphone, FB i This refers to the frequency response of the feedback filter coefficients corresponding to the i-th error microphone, FP. 2i This refers to the transfer function of the second path corresponding to the i-th error microphone, Err. j This refers to the error signal spectrum acquired by the j-th error microphone, FB. j This refers to the frequency response of the feedback filter coefficients corresponding to the j-th error microphone, FP. 2j It refers to the transfer function of the second path corresponding to the j-th error microphone, where j is any value from 1 to m except i, and m is the number of at least one error microphone.

[0172] In some embodiments, based on the above formula, with the goal of minimizing the residual, the frequency response of at least one filter coefficient can be determined by optimization algorithms such as the GA / PSO algorithm, adaptive iterative algorithm, or multi-objective FxLMS adaptive algorithm.

[0173] When the target signal is determined by the second implementation method, at least one filter coefficient includes at least one feedback filter coefficient, and the frequency response of each feedback filter coefficient in the at least one feedback filter coefficient satisfies the following formula (3) and the above formula (2).

[0174] In formula (3) above, noise refers to the noise spectrum in the environment, DL refers to the signal spectrum played through the sound-emitting unit, and FP... 1i This refers to the transfer function of the first path corresponding to the i-th error microphone, Err. i It refers to the error signal spectrum acquired by the i-th error microphone, FBi This refers to the frequency response of the feedback filter coefficients corresponding to the i-th error microphone, FP. 2i This refers to the transfer function of the second path corresponding to the i-th error microphone, Err. j This refers to the error signal spectrum acquired by the j-th error microphone, FB. j This refers to the frequency response of the feedback filter coefficients corresponding to the j-th error microphone, FP. 2j It refers to the transfer function of the second path corresponding to the j-th error microphone, where j is any value from 1 to m except i, and m is the number of at least one error microphone.

[0175] In some embodiments, based on the above formula, with the goal of minimizing the residual, the frequency response of at least one filter coefficient can be determined by optimization algorithms such as the GA / PSO algorithm, adaptive iterative algorithm, or multi-objective FxLMS adaptive algorithm.

[0176] When the target signal is determined by the third implementation method, at least one filter coefficient includes at least one feedback filter coefficient and a feedforward filter coefficient, and the frequency response of each filter coefficient in the at least one filter coefficient satisfies the following formula (4) and the above formula (2).

[0177] In formula (4) above, noise refers to the noise spectrum in the environment, DL refers to the signal spectrum played through the sound-emitting unit, and FP... 1i This refers to the transfer function of the first path corresponding to the i-th error microphone, Err. i This refers to the error signal spectrum acquired by the i-th error microphone, FF refers to the frequency response of the feedforward filter coefficients, and FB refers to... i This refers to the frequency response of the feedback filter coefficients corresponding to the i-th error microphone, FP. 2i This refers to the transfer function of the second path corresponding to the i-th error microphone, Err. j This refers to the error signal spectrum acquired by the j-th error microphone, FB. j This refers to the frequency response of the feedback filter coefficients corresponding to the j-th error microphone, FP. 2j It refers to the transfer function of the second path corresponding to the j-th error microphone, where j is any value from 1 to m except i, and m is the number of at least one error microphone.

[0178] In some embodiments, the frequency response of at least one filter coefficient can be determined based on the above formula using optimization algorithms such as the GA / PSO algorithm, adaptive iterative algorithm, or multi-objective FxLMS adaptive algorithm.

[0179] It should be noted that if the frequency response of the feedforward filter coefficients satisfies the following formula (5), the feedforward filter algorithm can be guaranteed to have the best sound leakage prevention capability.

[0180] Since the ratio of the transfer function of the first path to the transfer function of the second path is different for different error microphones, the frequency response of a single feedforward filter coefficient cannot simultaneously minimize the error signal of each error microphone. Therefore, the path ratio corresponding to at least one error microphone can be determined, and the path ratio is the ratio between the transfer function of the first path and the transfer function of the second path corresponding to the corresponding error microphone. The frequency response of the feedforward filter coefficient is determined based on the path ratio corresponding to at least one error microphone.

[0181] The average, maximum, or minimum path ratios corresponding to at least one error microphone are determined as the target ratios. The target ratios are then used as the design target of the feedforward filter, and the frequency response of the feedforward filter coefficients is determined by combining the above formulas (4) and (2).

[0182] In some embodiments, before determining the target signal based on at least one error signal and at least one filter coefficient, a target leakage signal-to-noise ratio can also be determined based on at least one error signal. The target leakage signal-to-noise ratio is used to describe the magnitude relationship between leakage sound and noise in the environment. If the leakage signal-to-noise ratio is greater than or equal to a signal-to-noise ratio threshold, the step of determining the target signal based on at least one error signal and at least one filter coefficient is performed.

[0183] For any one of the at least one error signals, the leakage signal-to-noise ratio (SNR) can be determined based on that error signal using a relevant algorithm. Processing each of the at least one error signal in the same way yields at least one leakage SNR. In this case, the maximum, minimum, median, mode, or average of the at least one leakage SNR can be determined as the target leakage SNR.

[0184] For example, the leaked sound signal and the noise signal can be separated based on the correlation between the downlink signal and the error signal (correlation algorithms such as Kalman / LMS), and the ratio of the leaked sound to the noise can be determined as the leaked sound signal-to-noise ratio corresponding to the error signal.

[0185] If the target sound leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, it indicates that the sound leakage of the headphones is large. Therefore, subsequent sound leakage prevention steps need to be performed. Thus, subsequent steps can be performed to determine the target signal based on at least one error signal and at least one filter coefficient.

[0186] If the target signal-to-noise ratio (SNR) of the leaked sound is less than the SNR threshold, then at least one filter coefficient is adjusted so that the average absolute value of the filter coefficients corresponding to the adjusted at least one filter algorithm is less than the average absolute value of the filter coefficients before the adjustment, and the step of determining the target signal based on at least one error signal and at least one filter coefficient is executed; or, the step of determining the target signal based on at least one error signal and at least one filter coefficient is not executed.

[0187] If the target sound leakage signal-to-noise ratio (SNR) is less than the SNR threshold, it indicates that the headphone's sound leakage is relatively small. The filter coefficients can be adjusted to be smaller to reduce the sound leakage prevention effect. Therefore, at least one filter coefficient can be adjusted so that the average absolute value of the adjusted filter coefficient is less than the average absolute value of the unadjusted filter coefficient. Then, the step of determining the target signal based on at least one error signal and at least one filter coefficient can be performed. Alternatively, the subsequent sound leakage prevention function can be omitted entirely, i.e., the step of determining the target signal based on at least one error signal and at least one filter coefficient can be skipped.

[0188] In some embodiments, a filter coefficient that needs to be adjusted is determined from at least one filter coefficient to obtain at least one target filter coefficient. For each filter coefficient in the at least one target filter coefficient, if the filter coefficient is positive, the target value is subtracted from the filter coefficient; if the filter coefficient is negative, the target value is added to the filter coefficient. In this way, the absolute value of the filter coefficient can be reduced.

[0189] For example, at least one filter coefficient can be directly determined as at least one target filter coefficient, or any one or more filter coefficients among at least one filter coefficient can be determined as the target filter coefficient. This application does not limit this.

[0190] Step 704: Play the target signal through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit.

[0191] In some embodiments, the headphones further include at least one monitoring microphone; in this case, during the process of playing the target signal through the first air conduction speaker, the sound played by the first air conduction speaker is collected by the at least one monitoring microphone to obtain at least one monitoring signal; if it is determined based on the at least one monitoring signal that there is an abnormality in the first air conduction speaker, the playback of the target signal through the first air conduction speaker is stopped, or at least one filter coefficient is cleared to zero.

[0192] Among the abnormal situations are uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space (i.e., the sound generated by the first loudspeaker has phenomena such as attenuation and divergence), and / or, the sound generated by the first air-conducting loudspeaker has a whistling phenomenon.

[0193] For the implementation method of determining the abnormality of the first air-conducting loudspeaker based on at least one monitoring signal, please refer to relevant howling detection and audio divergence detection technologies. This application embodiment will not elaborate on this.

[0194] In some embodiments, the at least one monitoring microphone may be a microphone independent of the at least one error microphone described above. In other embodiments, when there is only one error microphone, the error microphone may integrate the functions of the monitoring microphone. When there are multiple error microphones, any one or more of the multiple error microphones may integrate the functions of the monitoring microphone. In other words, the at least one monitoring microphone and the at least one error microphone can be independent of each other, each performing its corresponding function. Of course, one or more of the at least one error microphones can also be reused to implement the functions of the monitoring microphone. This application does not limit this aspect.

[0195] By integrating at least one microphone into the headphones to acquire ambient sound in real time, an error signal is obtained. A target signal is then generated based on the error signal and at least one filter coefficient, and played through a first air conduction unit. Since ambient sound includes sound leakage generated when the sound-emitting unit plays the downlink signal, playing the target signal generated based on the error signal can cancel out the sound leakage during the downlink signal playback process, thus reducing sound leakage. Furthermore, compared to related techniques that reduce sound leakage through special cavities, the various methods provided in this application embodiment do not rely on cavity design, thus offering greater flexibility.

[0196] Figure 12 is a flowchart of another audio processing method provided in an embodiment of this application. Optionally, the method shown in Figure 12 is applied to the headphones shown in Figure 5, the headphones including a first air-conducting speaker and a sound-emitting unit. Referring to Figure 12, the method includes the following steps.

[0197] Step 1201: Play downlink signal through sound unit.

[0198] In some embodiments, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0199] Based on the description of the above application scenarios, since bone conduction headphones generate hearing through bone conduction and do not rely on air conduction, users can hear audio clearly even in noisy environments. In contrast, air conduction headphones are still difficult to hear at high volumes in noisy environments. Therefore, if the sound-generating unit in this application embodiment is a bone conduction vibrator, the vibration of the bone conduction vibrator is less likely to be masked by ambient sound, thus enabling users to hear audio clearly even in noisy environments, thereby improving the user experience of the headphones. Based on this, combined with the subsequent processing steps of this application embodiment, it is possible to effectively reduce sound leakage outside the ear through bone conduction, enhance the signal-to-noise ratio inside the ear, and thus effectively improve the quality of audio playback.

[0200] Step 1202: Process the downlink signal based on the filter coefficients of the feedforward filtering algorithm to obtain the feedforward inverted signal, which is opposite in phase to the leaked sound signal.

[0201] It should be noted that the feedforward filter coefficients in the embodiments of this application are actually a set of at least one coefficient. The feedforward filter coefficients correspond to the feedforward filter algorithm (also known as the feedforward filter). That is to say, the feedforward filter coefficients include the coefficients corresponding to the feedforward filter, and the number of these coefficients can be one or more.

[0202] For example, please refer to Figure 13, which is a schematic diagram of an earphone provided in an embodiment of this application. The earphone includes a first air-conducting speaker and a sound-generating unit (i.e., the bone conduction vibrator in Figure 13). In this case, the downlink signal can be directly processed by the feedforward filter based on the filter coefficients corresponding to the feedforward filter to obtain the feedforward inverted signal.

[0203] In some embodiments, the feedforward filter coefficients may be determined before processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal.

[0204] In one possible implementation, the headphones further include at least one leakage sampling point, each leakage sampling point corresponding to a first path and a second path. The first path is the air propagation path from the sound-generating unit to the corresponding leakage sampling point, and the second path is the air propagation path from the first air-conducting loudspeaker to the corresponding leakage sampling point. In this case, the frequency response of the feedforward filter coefficients can be determined based on the transfer function of the first path and the transfer function of the second path corresponding to each leakage sampling point. Based on the frequency response of the feedforward filter coefficients, the feedforward filter coefficients are obtained according to the relevant time-frequency conversion algorithm.

[0205] The first and second paths described above will be further illustrated by examples. Please refer to Figure 14, which is a schematic diagram of an earphone provided in an embodiment of this application. The earphone includes a first air-conducting speaker, a sound-generating unit (i.e., the bone conduction vibrator in Figure 11), and two leakage sampling points, namely leakage sampling point 1 and leakage sampling point 2. The first path corresponding to leakage sampling point 1 is FP11, the second path corresponding to leakage sampling point 1 is FP21, the first path corresponding to leakage sampling point 2 is FP12, and the second path corresponding to leakage sampling point 1 is FP22.

[0206] In some embodiments, the frequency response of the feedforward filter coefficients satisfies the following formula (6). i =noise+DL·(FP) 1i -FF·FP 2i (6)

[0207] In the above formula (6), noise refers to the noise spectrum in the environment, DL refers to the signal spectrum played through the sound-emitting unit, and FP... 1i This refers to the transfer function of the first path corresponding to the i-th leaky sound sampling point, Err. i It refers to the error signal spectrum collected at the i-th leaky sound sampling point, FP 2i It refers to the transfer function of the second path corresponding to the i-th leaky sound sampling point, and FF refers to the frequency response of the feedforward filter coefficients.

[0208] In some embodiments, the frequency response of the feedforward filter coefficients can be determined based on the above formula, with the goal of minimizing the residual, using optimization algorithms such as the GA / PSO algorithm, adaptive iterative algorithm, or multi-objective FxLMS adaptive algorithm.

[0209] It should be noted that if the frequency response of the feedforward filter coefficients satisfies the following formula (7), the feedforward filter can be guaranteed to have the best anti-leakage capability.

[0210] Since the ratio of the transfer function of the first path to the transfer function of the second path is different for different leaky sound sampling points, the frequency response of a single feedforward filter coefficient cannot simultaneously minimize the error signal for each leaky sound sampling point. Therefore, at least one path ratio can be determined for each leaky sound sampling point. The path ratio is the ratio between the transfer function of the first path and the transfer function of the second path for the corresponding leaky sound sampling point. The frequency response of the feedforward filter coefficient is determined based on the path ratio for each leaky sound sampling point.

[0211] The average, maximum, or minimum value of the path ratio corresponding to at least one leaky sound sampling point is determined as the target ratio. The target ratio is then used as the design target of the feedforward filter, and the frequency response of the feedforward filter coefficient is determined by combining the above formula (6).

[0212] In some embodiments, each of the at least one leaky sound sampling point is equipped with an error microphone. In one possible implementation, the number of at least one leaky sound sampling point is 1.

[0213] In some embodiments, before processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal, at least one error signal can be obtained by acquiring ambient sound through at least one error microphone; based on the at least one error signal, the target leakage signal-to-noise ratio is determined, and the target sound energy is used to describe the magnitude relationship between the leakage sound and the noise in the environment; if the target leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, the step of processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal is performed.

[0214] For any one of the at least one error signals, the leakage signal-to-noise ratio (SNR) can be determined based on that error signal using a relevant algorithm. Processing each of the at least one error signal in the same way yields at least one leakage SNR. In this case, the maximum, minimum, median, mode, or average of the at least one leakage SNR can be determined as the target leakage SNR.

[0215] For example, the leaked sound signal and the noise signal can be separated based on the correlation between the downlink signal and the error signal (correlation algorithms such as Kalman / LMS), and the ratio of the leaked sound to the noise can be determined as the leaked sound signal-to-noise ratio corresponding to the error signal.

[0216] If the target sound leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, it indicates that the sound leakage of the headphones is large. Therefore, it is necessary to perform subsequent sound leakage prevention steps. Thus, the subsequent step of processing the downlink signal based on the feedforward filter coefficient to obtain the feedforward inverted signal can be performed.

[0217] If the target signal-to-noise ratio (SNR) of the leaky audio is less than the SNR threshold, the feedforward filter coefficients are adjusted so that the absolute value of the adjusted feedforward filter coefficients is less than the absolute value of the original feedforward filter coefficients. Then, the step of processing the downlink signal based on the feedforward filter coefficients to obtain a feedforward inverted signal is performed. Alternatively, the step of processing the downlink signal based on the feedforward filter coefficients to obtain a feedforward inverted signal is not performed.

[0218] If the target leakage signal-to-noise ratio (SNR) is less than the SNR threshold, it indicates that the headphone leakage is relatively small. The filter coefficient can be adjusted to a smaller value to reduce the leakage prevention effect. Therefore, the feedforward filter coefficient can be adjusted so that the absolute value of the adjusted feedforward filter coefficient is less than the absolute value of the original feedforward filter coefficient. Then, the step of processing the downlink signal based on the feedforward filter coefficient to obtain a feedforward inverted signal can be performed. Alternatively, the subsequent leakage prevention function can be omitted entirely, i.e., the step of processing the downlink signal based on the feedforward filter coefficient to obtain a feedforward inverted signal can be skipped.

[0219] In some embodiments, if the feedforward filter coefficient is positive, the target value is subtracted from the filter coefficient; if the filter coefficient is negative, the target value is added to the filter coefficient. In this way, the absolute value of the filter coefficient can be reduced.

[0220] Step 1203: Play a feedforward inverted signal through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit.

[0221] In some embodiments, the headphones further include at least one monitoring microphone; in this case, during the process of playing a feedforward inverted signal through the first air conduction speaker, the sound played by the first air conduction speaker is collected through the at least one monitoring microphone to obtain at least one monitoring signal; if it is determined based on the at least one monitoring signal that there is an abnormality in the first air conduction speaker, the playback of the feedforward inverted signal through the first air conduction speaker is stopped, or the feedforward filter coefficient is cleared to zero.

[0222] Among the abnormal situations are uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space (i.e., the sound generated by the first loudspeaker has phenomena such as attenuation and divergence), and / or, the sound generated by the first air-conducting loudspeaker has a whistling phenomenon.

[0223] For the implementation method of determining the abnormality of the first air-conducting loudspeaker based on at least one monitoring signal, please refer to relevant howling detection and audio divergence detection technologies. This application embodiment will not elaborate on this.

[0224] In some embodiments, the at least one monitoring microphone may be a microphone independent of the at least one error microphone described above. In other embodiments, when there is only one error microphone, the error microphone may integrate the functions of the monitoring microphone. When there are multiple error microphones, any one or more of the multiple error microphones may integrate the functions of the monitoring microphone. In other words, the at least one monitoring microphone and the at least one error microphone can be independent of each other, each performing its corresponding function. Of course, one or more of the at least one error microphones can also be reused to implement the functions of the monitoring microphone. This application does not limit this aspect.

[0225] By processing the downlink signal through feedforward filtering coefficients, a feedforward inverted signal can be generated. When this signal is played through the first air-conducting speaker, it physically cancels out the original sound leakage generated by the sound-generating unit, thus significantly reducing sound leakage. In other words, the various audio processing methods provided in this application all use the first air-conducting speaker to play the generated signal to eliminate sound leakage, effectively preventing others from hearing the content played through the headphones and protecting user privacy. Furthermore, compared to related technologies that reduce sound leakage through special cavities, the various methods provided in this application do not rely on cavity design, thus offering greater flexibility.

[0226] Figure 15 is a schematic diagram of an audio processing device provided in an embodiment of this application. This audio processing device can be implemented by software, hardware, or a combination of both as part or all of the aforementioned headphones. Referring to Figure 15, the device includes: a first playback module 1501, a acquisition module 1502, a determination module 1503, and a second playback module 1504.

[0227] The first playback module 1501 is used to play downlink signals through the sound unit. For detailed implementation details, please refer to the corresponding content in the above embodiments; they will not be repeated here.

[0228] The acquisition module 1502 is used to acquire ambient sound through at least one error microphone to obtain at least one error signal. For detailed implementation details, please refer to the corresponding contents in the above embodiments; they will not be repeated here.

[0229] The determination module 1503 is used to determine the target signal based on at least one error signal and at least one filter coefficient. For detailed implementation details, please refer to the corresponding contents in the above embodiments; they will not be repeated here.

[0230] The second playback module 1504 is used to play the target signal through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit. For detailed implementation details, please refer to the corresponding contents in the above embodiments; they will not be repeated here.

[0231] In one possible implementation, at least one filter coefficient includes at least one feedback filter coefficient, and the at least one feedback filter coefficient corresponds one-to-one with at least one error microphone;

[0232] Module 1503 is specifically used for:

[0233] Based on at least one feedback filter coefficient, the error signals collected by the corresponding error microphones in at least one error signal are processed to obtain at least one feedback inverted signal, which is opposite in phase to the corresponding error signal.

[0234] The target signal is determined based on at least one feedback inverted signal.

[0235] In one possible implementation, the determining module 1503 is specifically used for:

[0236] The target signal is obtained by superimposing at least one feedback inverted signal.

[0237] In one possible implementation, the determining module 1503 is specifically used for:

[0238] The target signal is obtained by superimposing the downlink signal with at least one feedback inverted signal.

[0239] In one possible implementation, at least one filter coefficient also includes a feedforward filter coefficient;

[0240] Module 1503 is specifically used for:

[0241] The downlink signal is processed based on the feedforward filter coefficients to obtain the feedforward inverted signal, which is opposite in phase to the leaked tone.

[0242] The target signal is obtained by superimposing the feedforward inverted signal with at least one feedback inverted signal.

[0243] In one possible implementation, each error microphone in at least one error microphone corresponds to a first path and a second path, the first path being the air propagation path from the sound-emitting unit to the corresponding error microphone, and the second path being the air propagation path from the first air-conducting loudspeaker to the corresponding error microphone.

[0244] Module 1503 is specifically used for:

[0245] Determine the path ratio corresponding to at least one error microphone, where the path ratio is the ratio between the transfer function of the first path and the transfer function of the second path corresponding to the respective error microphone.

[0246] The frequency response of the feedforward filter coefficients is determined based on the path ratios corresponding to at least one error microphone.

[0247] The feedforward filter coefficients are determined based on the frequency response of the feedforward filter coefficients.

[0248] In one possible implementation, each error microphone in at least one error microphone corresponds to a second path, which is the air propagation path from the first air-conducting loudspeaker to the corresponding error microphone.

[0249] Module 1503 is specifically used for:

[0250] For each of at least one feedback filter coefficient, the frequency response of the feedback filter coefficient is determined based on the transfer function of the second path corresponding to the feedback filter coefficient, such that the absolute value of the product of the frequency response of the feedback filter coefficient and the transfer function of the second path plus 1 is greater than 1.

[0251] The feedback filter coefficients are determined based on the frequency response of the feedback filter coefficients.

[0252] In one possible implementation, the number of at least one error microphone is 1.

[0253] In one possible implementation, the determining module 1503 is specifically used for:

[0254] Based on at least one error signal, the target leakage signal-to-noise ratio is determined, and the leakage signal-to-noise ratio is used to describe the relationship between the leakage sound and the noise in the environment.

[0255] If the target leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, then the step of determining the target signal based on at least one error signal and at least one filter coefficient is performed.

[0256] In one possible implementation, the determining module 1503 is specifically used for:

[0257] If the target signal-to-noise ratio (SNR) is less than the SNR threshold, at least one filter coefficient is adjusted so that the average absolute value of the adjusted at least one filter coefficient is less than the average absolute value of the unadjusted at least one filter coefficient, and the step of determining the target signal based on at least one error signal and at least one filter coefficient is performed.

[0258] In one possible implementation, the headset also includes at least one monitoring microphone;

[0259] The device also includes:

[0260] The monitoring module is used to collect the sound played by the first air conduction speaker through at least one monitoring microphone during the process of playing the target signal through the first air conduction speaker, and obtain at least one monitoring signal;

[0261] The stop-zeroing module is used to stop playing the target signal through the first air conduction speaker or to zero at least one filter coefficient if it is determined that there is an abnormality in the first air conduction speaker based on at least one monitoring signal.

[0262] In one possible implementation, the abnormal conditions include uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space, and / or, feedback phenomenon in the sound generated by the first air-conducting loudspeaker.

[0263] In one possible implementation, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0264] In this embodiment, at least one microphone is integrated into the earphone to collect ambient sound in real time to obtain an error signal. A target signal is then generated based on the error signal and at least one filter coefficient, and played through a first air conduction unit. Since ambient sound includes sound leakage generated when the sound-emitting unit plays the downlink signal, playing the target signal generated based on the error signal can cancel out the sound leakage during the downlink signal playback process, thus reducing sound leakage. Furthermore, compared to related techniques that reduce sound leakage through special cavities, the various methods provided in this embodiment do not rely on cavity design, thus offering greater flexibility.

[0265] It should be noted that the audio processing device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device and the audio processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0266] Figure 16 is a schematic diagram of an audio processing device provided in an embodiment of this application. This audio processing device can be implemented as part or all of the aforementioned headphones by software, hardware, or a combination of both. Referring to Figure 16, the device includes: a first playback module 1601, a processing module 1602, and a second playback module 1603.

[0267] The first playback module 1601 is used to play downlink signals through the sound unit. For detailed implementation details, please refer to the corresponding content in the above embodiments; they will not be repeated here.

[0268] Processing module 1602 is used to process the downlink signal based on the feedforward filter coefficients to obtain a feedforward inverted signal. For detailed implementation details, please refer to the corresponding contents in the above embodiments; they will not be repeated here.

[0269] The second playback module 1603 is used to play a feedforward inverted signal through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit. The feedforward inverted signal is out of phase with the leaked sound. For detailed implementation process, please refer to the corresponding content in the above embodiments, which will not be repeated here.

[0270] In one possible implementation, the headphones further include at least one leakage sampling point, each of the at least one leakage sampling point corresponds to a first path and a second path, the first path is the air propagation path from the sound unit to the corresponding leakage sampling point, and the second path is the air propagation path from the first air-conducting speaker to the corresponding leakage sampling point.

[0271] Processing module 1602 is specifically used for:

[0272] Determine the path ratio corresponding to at least one leaked sound sampling point. The path ratio is the ratio between the transfer function of the first path and the transfer function of the second path corresponding to the leaked sound sampling point.

[0273] The frequency response of the feedforward filter coefficients is determined based on the path ratio corresponding to at least one point of sound leakage.

[0274] The feedforward filter coefficients are determined based on the frequency response of the feedforward filter coefficients.

[0275] In one possible implementation, each of at least one leak sampling point is equipped with an error microphone.

[0276] In one possible implementation, the number of at least one leaky sound sampling point is 1.

[0277] In one possible implementation, the processing module 1602 is specifically used for:

[0278] At least one error signal is obtained by acquiring ambient sound through at least one error microphone;

[0279] Based on at least one error signal, the target leakage signal-to-noise ratio is determined, and the target sound energy is used to describe the relationship between the leakage sound and the noise in the environment.

[0280] If the target leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, then the downlink signal is processed based on the feedforward filter coefficients to obtain the feedforward inverted signal.

[0281] In one possible implementation, the processing module 1602 is specifically used for:

[0282] If the target leakage signal-to-noise ratio is less than the signal-to-noise ratio threshold, the feedforward filter coefficients are adjusted so that the absolute value of the adjusted feedforward filter coefficients is less than the absolute value of the feedforward filter coefficients before adjustment. Then, the downlink signal is processed based on the feedforward filter coefficients to obtain the feedforward inverted signal.

[0283] In one possible implementation, the headset also includes at least one monitoring microphone;

[0284] The device also includes:

[0285] The monitoring module is used to acquire the sound played by the first air-conducting speaker through at least one monitoring microphone during the process of playing the feedforward inverted signal through the first air-conducting speaker, and obtain at least one monitoring signal.

[0286] The stop-and-clear module is used to stop playing the feedforward inverted signal through the first air-conducting loudspeaker or clear the feedforward filter coefficient to zero if an abnormality is determined based on at least one monitoring signal.

[0287] In one possible implementation, the abnormal conditions include uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space, and / or, feedback phenomenon in the sound generated by the first air-conducting loudspeaker.

[0288] In one possible implementation, the sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

[0289] In this embodiment, the downlink signal is processed by a feedforward filter coefficient to generate a feedforward inverted signal. When this signal is played through the first air-conducting speaker, it physically cancels out the original sound leakage generated by the sound-generating unit, thus significantly reducing sound leakage. In other words, the various audio processing methods provided in this embodiment all use the first air-conducting speaker to play the generated signal to eliminate sound leakage, effectively preventing others from hearing the content played through the headphones and protecting user privacy. Furthermore, compared to related technologies that reduce sound leakage through special cavities, the various methods provided in this embodiment do not rely on cavity design, thus offering greater flexibility.

[0290] It should be noted that the audio processing device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the audio processing device and the audio processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0291] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform the steps of the audio processing method described in the above embodiments.

[0292] This application also provides a computer program product containing instructions that, when executed on a computer or processor, cause the computer or processor to perform the steps of the audio processing method described in the above embodiments. Alternatively, a computer program is provided that, when executed on a computer or processor, causes the computer or processor to perform the steps of the audio processing method described in the above embodiments.

[0293] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.

[0294] It should be understood that "multiple" as mentioned herein refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., do not necessarily imply that they are different.

[0295] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the downlink signals and error signals involved in the embodiments of this application were obtained under full authorization.

[0296] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, Applied to headphones, the headphones including a first air-conducting speaker, a sound-generating unit, and at least one error microphone, the method includes: The downlink signal is played through the sound-emitting unit; At least one error signal is obtained by acquiring ambient sound through the at least one error microphone; The target signal is determined based on the at least one error signal and at least one filter coefficient; The target signal is played through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit.

2. The method as described in claim 1, characterized in that, The at least one filter coefficient includes at least one feedback filter coefficient, and the at least one feedback filter coefficient corresponds one-to-one with the at least one error microphone; Determining the target signal based on the at least one error signal and the at least one filter coefficient includes: Based on the at least one feedback filter coefficient, the error signals collected by the corresponding error microphones in the at least one error signal are processed to obtain at least one feedback inverted signal, wherein the feedback inverted signal is opposite in phase to the corresponding error signal; The target signal is determined based on the at least one feedback inverted signal.

3. The method as described in claim 2, characterized in that, Determining the target signal based on the at least one feedback inverted signal includes: The target signal is obtained by superimposing the at least one feedback inverted signal.

4. The method as described in claim 2, characterized in that, Determining the target signal based on the at least one feedback inverted signal includes: The target signal is obtained by superimposing the downlink signal with the at least one feedback inverted signal.

5. The method as described in claim 2, characterized in that, The at least one filter coefficient also includes a feedforward filter coefficient; Determining the target signal based on the at least one feedback inverted signal includes: The downlink signal is processed based on the feedforward filter coefficients to obtain a feedforward inverted signal, which is opposite in phase to the leaked sound. The target signal is obtained by superimposing the feedforward inverted signal with the at least one feedback inverted signal.

6. The method as described in claim 5, characterized in that, Each error microphone in the at least one error microphone corresponds to a first path and a second path, the first path being the air propagation path from the sound-emitting unit to the corresponding error microphone, and the second path being the air propagation path from the first air-conducting loudspeaker to the corresponding error microphone. Before processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal, the method further includes: Determine the path ratio corresponding to each of the at least one error microphone, wherein the path ratio is the ratio between the transfer function of the first path and the transfer function of the second path corresponding to the respective error microphone; The frequency response of the feedforward filter coefficients is determined based on the path ratios corresponding to the at least one error microphone. The feedforward filter coefficients are determined based on the frequency response of the feedforward filter coefficients.

7. The method according to any one of claims 2-6, characterized in that, Each error microphone in the at least one error microphone corresponds to a second path, the second path being the air propagation path from the first air-conducting loudspeaker to the corresponding error microphone; Before processing the error signals acquired by the corresponding error microphones in the at least one error signal based on the at least one feedback filter coefficient, the method further includes: For each of the at least one feedback filter coefficients, the frequency response of the feedback filter coefficient is determined based on the transfer function of the second path corresponding to the feedback filter coefficient, such that the absolute value of the product of the frequency response of the feedback filter coefficient and the transfer function of the second path plus 1 is greater than 1. The feedback filter coefficients are determined based on the frequency response of the feedback filter coefficients.

8. The method according to any one of claims 1-7, characterized in that, The number of the at least one error microphone is 1.

9. The method according to any one of claims 1-8, characterized in that, Before determining the target signal based on the at least one error signal and the at least one filter coefficient, the method further includes: Based on the at least one error signal, a target leakage signal-to-noise ratio is determined, wherein the leakage signal-to-noise ratio is used to describe the relationship between the leakage sound and the noise in the environment; If the target leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, then the step of determining the target signal based on the at least one error signal and the at least one filter coefficient is performed.

10. The method as described in claim 9, characterized in that, The method further includes: If the target signal-to-noise ratio (SNR) of the leaked sound is less than the SNR threshold, then the at least one filter coefficient is adjusted so that the average absolute value of the at least one filter coefficient after adjustment is less than the average absolute value of the at least one filter coefficient before adjustment, and the step of determining the target signal based on the at least one error signal and the at least one filter coefficient is performed.

11. The method according to any one of claims 1-10, characterized in that, The headset also includes at least one monitoring microphone; The method further includes: During the process of playing the target signal through the first air-conducting speaker, the sound played by the first air-conducting speaker is collected through the at least one monitoring microphone to obtain at least one monitoring signal; If an abnormality is determined to exist in the first air-conducting loudspeaker based on the at least one monitoring signal, then the playback of the target signal through the first air-conducting loudspeaker shall be stopped, or the at least one filter coefficient shall be cleared to zero.

12. The method as described in claim 11, characterized in that, The abnormal conditions include uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space, and / or a whistling phenomenon in the sound generated by the first air-conducting loudspeaker.

13. The method according to any one of claims 1-12, characterized in that, The sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

14. An audio processing method, characterized in that, Applied to headphones, the headphones including a first air-conducting speaker and a sound-emitting unit, the method includes: The downlink signal is played through the sound-emitting unit; The downlink signal is processed based on the feedforward filter coefficients to obtain a feedforward inverted signal; The feedforward inverted signal is played through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit. The feedforward inverted signal is out of phase with the sound leakage.

15. The method as described in claim 14, characterized in that, The headphones also include at least one leakage sampling point, and each leakage sampling point in the at least one leakage sampling point corresponds to a first path and a second path. The first path is the air propagation path from the sound-generating unit to the corresponding leakage sampling point, and the second path is the air propagation path from the first air-conducting speaker to the corresponding leakage sampling point. Before processing the downlink signal based on the feedforward filter coefficients, the method further includes: Determine the path ratio corresponding to each of the at least one leaked sound sampling points, wherein the path ratio is the ratio between the transfer function of the first path and the transfer function of the second path corresponding to the corresponding leaked sound sampling point; The frequency response of the feedforward filter coefficients is determined based on the path ratios corresponding to the at least one leak point. The feedforward filter coefficients are determined based on the frequency response of the feedforward filter coefficients.

16. The method as described in claim 15, characterized in that, Each of the at least one leak sampling point is equipped with an error microphone.

17. The method as described in claim 15 or 16, characterized in that, The number of at least one leaky sound sampling point is 1.

18. The method as described in claim 16, characterized in that, Before processing the downlink signal based on the feedforward filter coefficients to obtain the feedforward inverted signal, the method further includes: At least one error signal is obtained by acquiring ambient sound through at least one error microphone; Based on the at least one error signal, the target leakage sound signal-to-noise ratio is determined, and the target sound energy is used to describe the relationship between the leakage sound and the noise in the environment. If the target leakage signal-to-noise ratio is greater than or equal to the signal-to-noise ratio threshold, then the step of processing the downlink signal based on the feedforward filter coefficient to obtain the feedforward inverted signal is performed.

19. The method as described in claim 18, characterized in that, The method further includes: If the target signal-to-noise ratio (SNR) of the leaky tone is less than the SNR threshold, the feedforward filter coefficients are adjusted so that the absolute value of the adjusted feedforward filter coefficients is less than the absolute value of the feedforward filter coefficients before adjustment, and the downlink signal is processed based on the feedforward filter coefficients to obtain a feedforward inverted signal.

20. The method according to any one of claims 14-19, characterized in that, The headset also includes at least one monitoring microphone; The method further includes: During the process of playing the feedforward inverted signal through the first air-conducting loudspeaker, the sound played by the first air-conducting loudspeaker is collected through the at least one monitoring microphone to obtain at least one monitoring signal; If an abnormality is determined to exist in the first air-conducting loudspeaker based on the at least one monitoring signal, then the playback of the feedforward inverted signal through the first air-conducting loudspeaker is stopped, or the feedforward filter coefficient is cleared to zero.

21. The method as described in claim 20, characterized in that, The abnormal conditions include uneven distribution or deformation of the sound waves generated by the first air-conducting loudspeaker in space, and / or, the sound generated by the first air-conducting loudspeaker exhibits a whistling phenomenon.

22. The method according to any one of claims 14-21, characterized in that, The sound-generating unit is a bone conduction vibrator or a second air conduction loudspeaker.

23. An audio processing apparatus, characterized in that, Applied to headphones, the headphones including a first air-conducting speaker, a sound-generating unit, and at least one error microphone, the device includes: The first playback module is used to play downlink signals through the sound-generating unit; The acquisition module is used to acquire ambient sound through the at least one error microphone to obtain at least one error signal; A determination module is used to determine a target signal based on the at least one error signal and at least one filter coefficient; The second playback module is used to play the target signal through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit.

24. An audio processing apparatus, characterized in that, Applied to headphones, the headphones including a first air-conducting speaker and a sound-emitting unit, the device includes: The first playback module is used to play downlink signals through the sound-generating unit; The processing module is used to process the downlink signal based on the feedforward filter coefficients to obtain a feedforward inverted signal; The second playback module is used to play the feedforward inverted signal through the first air-conducting loudspeaker to reduce sound leakage during the playback of the downlink signal by the sound-generating unit. The feedforward inverted signal is out of phase with the sound leakage.

25. An earphone, characterized in that, The headphones include a first air-conducting speaker, a sound-generating unit, at least one error microphone, and a processor; the processor is used to implement the steps of the method according to any one of claims 1-13.

26. An earphone, characterized in that, The headphones include a first air-conducting speaker, a sound-generating unit, and a processor; the processor is used to implement the steps of the method according to any one of claims 14-22.

27. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on a computer or processor, causes the computer or processor to perform the method as described in any one of claims 1-13, or to perform the method as described in any one of claims 14-22.

28. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a computer or processor, cause the steps of the method as described in any one of claims 1-13 to be performed, or cause the steps of the method as described in any one of claims 14-22 to be performed.