A pair of headphones

By using fixed structure and microphone array to generate noise reduction signals in open headphones, the problem of headphones blocking the ear canal and poor noise reduction effect is solved, and stability and comfort are improved.

CN115243137BActive Publication Date: 2025-07-22SHENZHEN SHOKZ CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111408328.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-29
Filing Date
2021-11-19
Publication Date
2025-07-22
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

Existing headphones are prone to block the user's ear canal during use, resulting in discomfort. The noise reduction effect of open headphones is not obvious when there is a high noise, affecting the user's hearing experience.

Method used

An open headphone is designed, and the headphones are fixed to the user's ears without blocking the ear canal. The first microphone array is used to pick up the ambient noise, and the sound field estimation is performed through the processor to generate a noise reduction signal, and the target signal is output from the speaker to reduce the ambient noise.

Benefits of technology

It effectively reduces environmental noise without blocking the ear canal, and improves the user's auditory experience and the stability and comfort of the headphones.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115243137B_ABST
    Figure CN115243137B_ABST
Patent Text Reader

Abstract

One or more embodiments of this specification relate to an earphone, which includes: a fixing structure configured to fix the earphone at a position near the user's ear without blocking the user's ear canal, and the fixing structure includes: a hook portion and a body portion; a first microphone array located in the body portion and configured to pick up ambient noise; a processor located in the hook portion or the body portion and configured to: estimate the sound field of a target spatial position by using the first microphone array, and generate a noise reduction signal based on the sound field estimation of the target spatial position; and a speaker located in the body portion and configured to: output a target signal according to the noise reduction signal, and the target signal is transmitted to the outside of the earphone through a sound outlet hole for reducing the ambient noise.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims the priority of International Application No. PCT / CN2021 / 109154 filed on July 29, 2021, the priority of International Application No. PCT / CN2021 / 089670 filed on April 25, 2021, and the priority of International Application No. PCT / CN2021 / 091652 filed on April 30, 2021, the contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of acoustics, and particularly to a headphone. Background Art

[0004] Active noise cancellation technology is a method of using the speakers of headphones to output sound waves opposite to the external environmental noise to cancel the environmental noise. Headphones can generally be divided into two categories: in-ear headphones and open headphones. In-ear headphones will block the user's ears during use, and users are prone to feelings of blockage, foreign objects, and swelling when wearing them for a long time. Open headphones can open the user's ears, which is beneficial for long-term wearing, but when the external noise is relatively large, their noise cancellation effect is not obvious, reducing the user's auditory experience.

[0005] Therefore, there is a need to provide a headphone and a noise cancellation method that can open the user's ears and improve the user's auditory experience. Summary of the Invention

[0006] An embodiment of this application provides a headphone, including: a fixing structure configured to fix the headphone near the user's ear without blocking the user's ear canal, the fixing structure including: a hook portion and a body portion, wherein when the user wears the headphone, the hook portion is hung between the first side of the user's ear and the head, and the body portion contacts the second side of the ear; a first microphone array located on the body portion and configured to pick up environmental noise; a processor located on the hook portion or the body portion and configured to: estimate the sound field of a target spatial position by using the first microphone array, the target spatial position being closer to the user's ear canal than any microphone in the first microphone array, and generate a noise cancellation signal based on the sound field estimation of the target spatial position; and a speaker located on the body portion and configured to: output a target signal according to the noise cancellation signal, the target signal being transmitted to the outside of the headphone through a sound outlet hole for reducing the environmental noise. Brief Description of the Drawings

[0007] This application will be further described in the form of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0008] Figure 1 is a framework diagram of an exemplary headset according to some embodiments of the present application;

[0009] Figure 2 is a schematic diagram of an exemplary ear according to some embodiments of the present application;

[0010] Figure 3 is a structural diagram of an exemplary earphone according to some embodiments of the present application;

[0011] Figure 4 is a wearing diagram of an exemplary headset according to some embodiments of the present application;

[0012] Figure 5 is a structural diagram of an exemplary earphone according to some embodiments of the present application;

[0013] Figure 6 is a wearing diagram of an exemplary headset according to some embodiments of the present application;

[0014] Figure 7 is a structural diagram of an exemplary earphone according to some embodiments of the present application;

[0015] Figure 8 is a wearing diagram of an exemplary headset according to some embodiments of the present application;

[0016] Figure 9A is a structural diagram of an exemplary earphone according to some embodiments of the present application;

[0017] Figure 9B is a structural diagram of an exemplary earphone according to some embodiments of the present application;

[0018] Figure 10 is a structural diagram of an exemplary earphone facing the ear according to some embodiments of the present application;

[0019] Figure 11 is a structural diagram of a side of an exemplary earphone facing away from the ear according to some embodiments of the present application;

[0020] Figure 12 is a top view of an exemplary headset according to some embodiments of the present application;

[0021] Figure 13 is a schematic cross-sectional structure diagram of an exemplary earphone according to some embodiments of the present application;

[0022] Figure 14 is an exemplary noise reduction flow chart of headphones according to some embodiments of the present application;

[0023] Figure 15 An exemplary flowchart for estimating the noise of a target spatial position as shown in some embodiments of the present application;

[0024] Figure 16 An exemplary flowchart for estimating the sound field and noise of a target spatial position as shown in some embodiments of the present application;

[0025] Figure 17 An exemplary flowchart for updating a noise reduction signal as shown in some embodiments of the present application;

[0026] Figure 18 An exemplary noise reduction flowchart for an earphone as shown in some embodiments of the present application;

[0027] Figure 19 An exemplary flowchart for estimating the noise of a target spatial position as shown in some embodiments of the present application. Detailed implementation manners

[0028] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the figures represent the same structure or operation.

[0029] It should be understood that the "system", "device", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the said words can be replaced by other expressions.

[0030] As shown in the present application and the claims, unless the context clearly indicates an exception, the words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0031] Flowcharts are used in the present application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations before or after do not necessarily need to be executed precisely in sequence. On the contrary, the steps can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several steps can be removed from these processes.

[0032] Some embodiments of this specification provide a pair of earphones. The earphones can be open earphones. The open earphones can fix the speakers near the user's ears through a fixing structure without blocking the user's ear canals. In some embodiments, the earphones can include a fixing structure, a first microphone array, a processor, and speakers. The fixing structure can be configured to fix the earphones near the user's ears without blocking the user's ear canals. The first microphone array, the processor, and the speakers can be located at the fixing structure to implement the active noise reduction function of the earphones. In some embodiments, the fixing structure can include a hook portion and a body portion. When the user wears the earphones, the hook portion can be hung between the first side of the user's ear and the head, and the body portion contacts the second side of the ear. In some embodiments, the body portion can include a connecting portion and a holding portion. When the user wears the earphones, the holding portion contacts the second side of the ear, and the connecting portion connects the hook portion and the holding portion. The connecting portion extends from the first side of the ear to the second side of the ear. The connecting portion cooperates with the hook portion to provide a pressing force on the second side of the ear for the holding portion, and the connecting portion cooperates with the holding portion to provide a pressing force on the first side of the ear for the hook portion, so that the earphones can clamp the user's ears and ensure the stability of the earphones during wearing. In some embodiments, the first microphone array can be located at the body portion of the earphones for picking up ambient noise. The processor is located at the hook portion or the body portion of the earphones for estimating the sound field at a target spatial position. The target spatial position can include a spatial position at a specific distance from the user's ear canal. For example, the target spatial position can be closer to the user's ear canal than any microphone in the first microphone array. It can be understood that the microphones in the first microphone array can be distributed at different positions near the user's ear canal, and the processor can estimate the sound field at a position near the user's ear canal (such as the target spatial position) according to the ambient noise collected by the microphones in the first microphone array. The speakers can be located at the body portion (holding portion) and output a target signal according to the noise reduction signal. The target signal can be transmitted to the outside of the earphones through the sound outlet holes on the holding portion to reduce the ambient noise heard by the user.

[0033] In some embodiments, in order to better reduce the ambient noise heard by the user, the body portion can include a second microphone. In comparison, the second microphone can be closer to the user's ear canal than the first microphone array, and the sound signal it collects is closer to and can reflect the sound heard by the user. The processor can update the above noise reduction signal according to the sound signal collected by the second microphone, so as to achieve a more ideal noise reduction effect.

[0034] It should be noted that the earphone provided in the embodiments of this specification can be fixed near the user's ear through a fixing structure without blocking the user's ear canal, opening the user's ears and improving the stability and comfort of the earphone during wearing. At the same time, the sound field near the user's ear canal (for example, the target spatial position) is estimated by using the first microphone array / second microphone located at the fixing structure (such as the body part) and the processor, and the ambient noise at the user's ear canal is reduced by the target signal output by the speaker, thereby realizing the active noise reduction of the earphone and improving the user's auditory experience during the use of the earphone.

[0035] Figure 1 It is a framework diagram of an exemplary earphone shown in some embodiments of the present application.

[0036] In some embodiments, the earphone 100 may include a fixing structure 110, a first microphone array 120, a processor 130, and a speaker 140. The first microphone array 120, the processor 130, and the speaker 140 may be located at the fixing structure 110. The earphone 100 can clamp the user's ear through the fixing structure 110 to fix the earphone 100 near the user's ear without blocking the user's ear canal. In some embodiments, the first microphone array 120 located at the fixing structure 110 (such as the body part) can pick up the ambient noise from the outside world and convert the ambient noise into an electrical signal and transmit it to the processor 130 for processing. The processor 130 is coupled (such as electrically connected) to the first microphone array 120 and the speaker 140. The processor 130 can receive the electrical signal transmitted by the first microphone array 120 and process it to generate a noise reduction signal, and transmit the generated noise reduction signal to the speaker 140. The speaker 140 can output a target signal according to the noise reduction signal. The target signal can be transmitted to the outside of the earphone 100 through the sound outlet hole on the fixing structure 110 (such as the holding part) and is used to reduce or cancel the ambient noise at the user's ear canal position (for example, the target spatial position), thereby realizing the active noise reduction of the earphone 100 and improving the user's auditory experience during the use of the earphone 100.

[0037] In some embodiments, the fixing structure 110 may include a hook part 111 and a body part 112. When the user wears the earphone 100, the hook part 111 can be hung between the first side of the user's ear and the head, and the body part 112 contacts the second side of the ear. The first side of the ear may be the dorsal side of the user's ear, and the second side of the user's ear may be the front side of the user's ear. The front side of the user's ear refers to the side where the user's ear includes parts such as the cymba conchae, triangular fossa, antihelix, scaphoid fossa, helix, etc. (the ear structure can be referred to Figure 2 ). The dorsal side of the user's ear refers to the side of the user's ear that faces away from the front side, that is, the side opposite to the front side.

[0038] In some embodiments, the body portion 112 may include a connecting portion and a holding portion. When the user wears the earphone 100, the holding portion contacts the second side of the ear, and the connecting portion connects the hook portion and the holding portion. The connecting portion extends from the first side of the ear to the second side of the ear. The connecting portion cooperates with the hook portion to provide a pressing force on the second side of the ear for the holding portion, and the connecting portion cooperates with the holding portion to provide a pressing force on the first side of the ear for the hook portion, so that the earphone 100 can be clamped near the user's ear by the fixing structure 110, ensuring the stability of the earphone 100 in wearing.

[0039] In some embodiments, the material of the hook portion 111 of the fixing structure 110 and / or the portion of the body portion 112 that contacts the user's ear can be selected according to specific circumstances. In some embodiments, a softer material can improve the comfort of the user wearing the earphone 100, and a harder material can improve the strength of the earphone 100. By reasonably configuring the materials of the various components of the earphone 100, the strength of the earphone 100 can be improved while improving the user's comfort.

[0040] The first microphone array 120 may be located in the body portion 112 (such as the connecting portion or the holding portion) of the fixing structure 110 for picking up ambient noise. In some embodiments, the ambient noise refers to a combination of various external sounds in the environment where the user is located. In some embodiments, by mounting the first microphone array 120 on the body portion 112 of the fixing structure 110, the first microphone array 120 can be located near the user's ear canal. Based on the ambient noise obtained in this way, the processor 130 can more accurately calculate the noise actually transmitted to the user's ear canal, which is more conducive to subsequent active noise reduction of the ambient noise heard by the user.

[0041] In some embodiments, the ambient noise may include the user's speaking voice. For example, the first microphone array 120 may pick up the ambient noise according to the working state of the earphone 100. The working state of the earphone 100 may refer to the usage state when the user wears the earphone 100. By way of example only, the working state of the earphone 100 may include but is not limited to a call state, a non-call state (e.g., a music playing state), a voice message sending state, etc. When the earphone 100 is in the non-call state, the sound generated by the user's own speech can be regarded as ambient noise, and the first microphone array 120 can pick up the sound of the user's own speech and other ambient noises. When the earphone 100 is in the call state, the sound generated by the user's own speech may not be regarded as ambient noise, and the first microphone array 120 can pick up the ambient noise other than the sound of the user's own speech. For example, the first microphone array 120 may pick up the noise emitted by a noise source at a certain distance (e.g., 0.5 meters, 1 meter) away from the first microphone array 120.

[0042] In some embodiments, the first microphone array 120 may include one or more air-conduction microphones. For example, when a user listens to music using the earphone 100, the air-conduction microphone can simultaneously acquire the noise of the external environment and the sound when the user speaks, and use the acquired noise of the external environment and the sound when the user speaks together as environmental noise. In some embodiments, the first microphone array 120 may further include one or more bone-conduction microphones. The bone-conduction microphone can be in direct contact with the user's skin. When the user speaks, the vibration signal generated by the bones or muscles can be directly transmitted to the bone-conduction microphone. Then, the bone-conduction microphone converts the vibration signal into an electrical signal and transmits the electrical signal to the processor 130 for processing. The bone-conduction microphone may also not be in direct contact with the human body. When the user speaks, the vibration signal generated by the bones or muscles can be first transmitted to the fixing structure 110 of the earphone 100, and then transmitted to the bone-conduction microphone by the fixing structure 110. In some embodiments, when the user is in a call state, the processor 130 may use the sound signal collected by the air-conduction microphone as environmental noise and perform noise reduction using the environmental noise, and use the sound signal collected by the bone-conduction microphone as a voice signal to be transmitted to the terminal device, thereby ensuring the call quality when the user is on the phone.

[0043] In some embodiments, the processor 130 may control the on / off states of the bone-conduction microphone and the air-conduction microphone based on the working state of the earphone 100. In some embodiments, when the first microphone array 120 picks up environmental noise, the on / off states of the bone-conduction microphone and the air-conduction microphone in the first microphone array 120 may be determined according to the working state of the earphone 100. For example, when the user wears the earphone 100 to play music, the on / off state of the bone-conduction microphone may be in a standby state, and the on / off state of the air-conduction microphone may be in a working state. For another example, when the user wears the earphone 100 to send a voice message, the on / off state of the bone-conduction microphone may be in a working state, and the on / off state of the air-conduction microphone may be in a working state. In some embodiments, the processor 130 may control the on / off states of the microphones (e.g., bone-conduction microphones, air-conduction microphones) in the first microphone array 120 by sending control signals.

[0044] In some embodiments, according to the working principle of the microphone, the first microphone array 120 may include a dynamic microphone, a ribbon microphone, a capacitive microphone, an electret microphone, an electromagnetic microphone, a carbon granule microphone, etc., or any combination thereof. In some embodiments, the arrangement mode of the first microphone array 120 may include a linear array (e.g., linear, curved), a planar array (e.g., cross-shaped, circular, annular, polygonal, mesh-shaped, etc., regular and / or irregular shapes), a three-dimensional array (e.g., cylindrical, spherical, hemispherical, polyhedral, etc.), etc., or any combination thereof.

[0045] The processor 130 may be located in the hook portion 111 or the body portion 112 of the fixed structure 110. The processor 130 may estimate the sound field at the target spatial position by using the first microphone array 120. The sound field at the target spatial position may refer to the distribution and variation of sound waves at or near the target spatial position (e.g., variation over time, variation over position). The physical quantities describing the sound field may include sound pressure level, sound frequency, sound amplitude, sound phase, sound source vibration velocity, or medium (e.g., air) density, etc. Generally, these physical quantities may be functions of position and time. The target spatial position may refer to a spatial position at a specific distance close to the user's ear canal. The specific distance here may be a fixed distance, e.g., 2 mm, 5 mm, 10 mm, etc. The target spatial position may be closer to the user's ear canal than any of the microphones in the first microphone array 120. In some embodiments, the target spatial position may be related to the number of microphones in the first microphone array 120 and their distribution positions relative to the user's ear canal. The target spatial position can be adjusted by adjusting the number of microphones in the first microphone array 120 and / or their distribution positions relative to the user's ear canal. For example, by increasing the number of microphones in the first microphone array 120, the target spatial position can be made closer to the user's ear canal. Also, for example, the target spatial position can be made closer to the user's ear canal by reducing the spacing between the microphones in the first microphone array 120. Further, for example, the target spatial position can be made closer to the user's ear canal by changing the arrangement of the microphones in the first microphone array 120.

[0046] In some embodiments, the processor 130 may be further configured to generate a noise reduction signal based on the sound field estimation at the target spatial position. Specifically, the processor 130 may receive the ambient noise acquired by the first microphone array 120 and process it to obtain the parameters of the ambient noise (e.g., amplitude, phase, etc.), and estimate the sound field at the target spatial position based on the parameters of the ambient noise. Further, the processor 130 generates a noise reduction signal based on the sound field estimation at the target spatial position. The parameters of the noise reduction signal (e.g., amplitude, phase, etc.) are related to the ambient noise at the target spatial position. By way of example only, the amplitude of the noise reduction signal may be approximately equal to the amplitude of the ambient noise at the target spatial position, and the phase of the noise reduction signal may be approximately opposite to the phase of the ambient noise at the target spatial position.

[0047] The speaker 140 may be located in the holding part of the fixed structure 110. When the user wears the earphone 100, the speaker 140 is located near the user's ear. The speaker 140 may output a target signal according to the noise reduction signal. The target signal may be transmitted to the user's ear through the sound outlet hole of the holding part to reduce or eliminate the environmental noise transmitted into the user's ear canal. In some embodiments, according to the working principle of the speaker, the speaker 140 may include one or more of an electro-dynamic speaker (e.g., a moving coil speaker), a magnetic speaker, an ion speaker, an electrostatic speaker (or a capacitive speaker), a piezoelectric speaker, etc. In some embodiments, according to the propagation mode of the sound output by the speaker, the speaker 140 may include an air-conduction speaker, a bone-conduction speaker. In some embodiments, the number of the speakers 140 may be one or more. When the number of the speakers 140 is one, the speaker may output a target signal to eliminate the environmental noise and simultaneously transmit effective sound information to the user (e.g., device media audio, remote call audio). For example, when the number of the speakers 140 is one and it is an air-conduction speaker, the air-conduction speaker may be used to output a target signal to eliminate the environmental noise. In this case, the target signal may be a sound wave (i.e., the vibration of air), and the sound wave may be transmitted through the air to a target spatial position and cancel out the environmental noise at the target spatial position. At the same time, the sound wave output by the air-conduction speaker also includes effective sound information. Another example, when the number of the speakers 140 is one and it is a bone-conduction speaker, the bone-conduction speaker may be used to output a target signal to eliminate the environmental noise. In this case, the target signal may be a vibration signal, and the vibration signal may be transmitted through the bone or tissue to the user's basilar membrane and cancel out the environmental noise at the user's basilar membrane. At the same time, the vibration signal output by the bone-conduction speaker also includes effective sound information. In some embodiments, when the number of the speakers 140 is multiple, a part of the multiple speakers 140 may be used to output a target signal to eliminate the environmental noise, and another part may be used to transmit effective sound information to the user (e.g., device media audio, remote call audio). For example, when the number of the speakers 140 is multiple and includes a bone-conduction speaker and an air-conduction speaker, the air-conduction speaker may be used to output a sound wave to reduce or eliminate the environmental noise, and the bone-conduction speaker may be used to transmit effective sound information to the user. Compared with the air-conduction speaker, the bone-conduction speaker can directly transmit mechanical vibration to the user's auditory nerve through the user's body (e.g., bones, skin tissue, etc.), and during this process, the interference to the air-conduction microphone for picking up environmental noise is relatively small.

[0048] In some embodiments, both the speaker 340 and the first microphone array 120 are located in the body portion 112 of the earphone 300. The target signal output by the speaker 340 may also be picked up by the first microphone array 120, and this target signal is not expected to be picked up, that is, the target signal should not be regarded as part of the ambient noise. In this case, in order to reduce the influence of the target signal output by the speaker 340 on the first microphone array 120, the first microphone array 120 can be disposed in a first target area. The first target area may be an area where the sound emitted by the speaker 340 has a relatively low or even the lowest intensity in space. For example, the first target area may be the acoustic null position of the radiation sound field of the acoustic dipole formed by the earphone 100 (e.g., sound outlet hole, pressure relief hole), or a position within a certain distance threshold range from the acoustic null position.

[0049] Figure 2 is a schematic diagram of an exemplary ear according to some embodiments of the present application.

[0050] See Figure 2 , the ear 200 may include the ear canal 201, the concha 202, the cymba conchae 203, the triangular fossa 204, the antihelix 205, the scaphoid fossa 206, the helix 207, the earlobe 208, and the crus of helix 209. In some embodiments, the wearing and stabilization of the earphone (e.g., earphone 100) can be achieved by means of one or more parts of the ear 200. In some embodiments, parts such as the ear canal 201, the concha 202, the cymba conchae 203, and the triangular fossa 204 have a certain depth and volume in three-dimensional space, which can be used to meet the wearing requirements of the earphone. In some embodiments, an open earphone (e.g., earphone 100) can be worn by means of parts such as the cymba conchae 203, the triangular fossa 204, the antihelix 205, the scaphoid fossa 206, the helix 207, or a combination thereof. In some embodiments, in order to improve the comfort and reliability of the earphone in terms of wearing, parts such as the user's earlobe 208 can be further utilized. By using other parts of the ear 200 except the ear canal 201 to achieve the wearing of the earphone and the propagation of sound, the user's ear canal 201 can be "liberated", reducing the impact of the earphone on the user's ear health. When the user wears the earphone on the road, the earphone does not block the user's ear canal 201, and the user can receive both the sound from the earphone and the sound from the environment (e.g., horn sounds, bicycle bell sounds, surrounding human voices, traffic command sounds, etc.), thereby reducing the probability of traffic accidents. For example, when the user wears the earphone, the whole or part of the structure of the earphone can be located in front of the crus of helix 209 (e.g., Figure 2The area J enclosed by the dashed line in the figure). For another example, when the user wears the earphone, the whole or part of the structure of the earphone can contact the upper part of the external auditory canal 201 (for example, the positions where one or more parts such as the helix crus 209, the cymba conchae 203, the triangular fossa 204, the antihelix 205, the scaphoid fossa 206, the helix 207 are located). For another example, when the user wears the earphone, the whole or part of the structure of the earphone can be located within one or more parts of the ear (for example, the cympa conchae 202, the cymba conchae 203, the triangular fossa 204, etc.) (for example, Figure 2 the area M enclosed by the dashed line in the figure).

[0051] Figure 3 is a structural diagram of an exemplary earphone shown according to some embodiments of the present application. Figure 4 is a wearing diagram of an exemplary earphone shown according to some embodiments of the present application.

[0052] Referring to Figures 3 - 4 , the earphone 300 can include a fixing structure 310, a first microphone array 320, a processor 330, and a speaker 340. Among them, the first microphone array 320, the processor 330, and the speaker 340 are located at the fixing structure 310. In some embodiments, the fixing structure 310 can be used to hang the earphone 300 near the user's ear without blocking the user's ear canal. In some embodiments, the fixing structure 310 can include a hook portion 311 and a body portion 312. In some embodiments, the hook portion 311 can include any shape suitable for the user to wear, for example, C-shaped, hook-shaped, etc. When the user wears the earphone 300, the hook portion 311 can be hung between the first side of the user's ear and the head. In some embodiments, the body portion 312 can include a connecting portion 3121 and a holding portion 3122, wherein the connecting portion 3121 is used to connect the hook portion 311 and the holding portion 3122. When the user wears the earphone 300, the holding portion 3122 contacts the second side of the ear, the connecting portion 3121 extends from the first side of the ear to the second side of the ear, and both ends of the connecting portion 3121 are respectively connected to the hook portion 311 and the holding portion 3122. The cooperation between the connecting portion 3121 and the hook portion 311 can provide a pressing force on the second side of the ear for the holding portion 3122, and the cooperation between the connecting portion 3121 and the holding portion 3122 can provide a pressing force on the first side of the ear for the connecting portion 3121.

[0053] In some embodiments, when the earphone 300 is in a non-wearing state (that is, the natural state), the connecting portion 3121 connects the hook portion 311 and the holding portion 3122, so that the fixing structure 310 is bent in three-dimensional space. It can also be understood that in three-dimensional space, the hook portion 311, the connecting portion 3121, and the holding portion 3122 are not coplanar. In this setting, when the earphone 300 is in a wearing state, as Figure 4As shown, the hook portion 311 can be hung between the first side of the user's ear 100 and the head, and the holding portion 3122 contacts the second side of the user's ear 100, so that the holding portion 3122 and the hook portion 311 cooperate to clamp the ear. In some embodiments, the connecting portion 3121 can extend from the head to the outside of the head (i.e., from the first side of the ear 100 to the second side of the ear), and then cooperate with the hook portion 311 to provide a pressing force on the second side of the ear 100 for the holding portion 3122. At the same time, according to the interaction of forces, when the connecting portion 3121 extends from the head to the outside of the head, it can also cooperate with the holding portion 3122 to provide a pressing force on the first side of the ear 100 for the hook portion 311, so that the fixing structure 310 can clamp the user's ear 100 to realize the wearing of the earphone 300.

[0054] In some embodiments, under the action of the pressing force, the holding portion 3122 can press against the ear. For example, it presses against the area where parts such as the cymba conchae, triangular fossa, and antihelix are located, so that the ear canal of the ear is not blocked when the earphone 300 is in the wearing state. Only for exemplary description, when the earphone 300 is in the wearing state, the projection of the holding portion 3122 on the user's ear can fall within the range of the helix of the ear; further, the holding portion 3122 can be located on the side of the ear canal of the ear close to the user's head top and contact the helix and / or antihelix. In this setting mode, on the one hand, it can avoid the holding portion 3122 blocking the ear canal, thus liberating the user's ears. At the same time, it can also increase the contact area between the holding portion 3122 and the ear, thereby improving the wearing comfort of the earphone 300. On the other hand, when the holding portion 3122 is located on the side of the ear canal of the ear close to the user's head top, the speaker 340 located at the holding portion 3122 can be closer to the user's ear canal, enhancing the auditory experience of the user when using the earphone 300.

[0055] In some embodiments, in order to improve the stability and comfort of the user wearing the earphone 300, the earphone 300 can also elastically clamp the ear. For example, in some embodiments, the hook portion 311 of the earphone 300 can include an elastic portion (not shown) connected to the connecting portion 3121. The elastic portion can have a certain elastic deformation ability, so that the hook portion 311 can deform under the action of an external force, and then generate a displacement relative to the holding portion 3122, allowing the hook portion 311 and the holding portion 3122 to cooperate to elastically clamp the ear. Specifically, during the process of the user wearing the earphone 300, the user can first apply force to make the hook portion 311 deviate from the holding portion 3122 so that the ear can extend between the holding portion 3122 and the hook portion 311; after the wearing position is appropriate, release the hand to allow the earphone 300 to elastically clamp the ear. The user can also further adjust the position of the earphone 300 on the ear according to the actual wearing situation.

[0056] In some embodiments, there may be significant differences among different users in terms of age, gender, gene-controlled trait expression, etc., resulting in different sizes and shapes of the ears and heads of different users. For this reason, in some embodiments, the hook portion 311 may be configured to be rotatable relative to the connecting portion 3121, or the holding portion 3122 may be rotatable relative to the connecting portion 3121, or a part of the connecting portion 3121 may be rotatable relative to another part, so that the relative positional relationship among the hook portion 311, the connecting portion 3121, and the holding portion 3122 in three-dimensional space is adjustable, facilitating the adaptation of the earphone 300 to different users, that is, increasing the applicable range of the earphone 300 for users in terms of wearing. At the same time, setting the relative positional relationship among the hook portion 311, the connecting portion 3121, and the holding portion 3122 in three-dimensional space to be adjustable can also adjust the positions of the first microphone array 320 and the speaker 340 relative to the user's ear (such as the external auditory canal), thereby improving the active noise reduction effect of the earphone 300. In some embodiments, the connecting portion 3121 may be made of deformable materials such as soft steel wires. The user bends the connecting portion 3121 to make a part of it rotate relative to another part, thereby adjusting the relative positions of the hook portion 311, the connecting portion 3121, and the holding portion 3122 in three-dimensional space, and further meeting their wearing requirements. In some embodiments, the connecting portion 3121 may also be provided with a rotating shaft mechanism 31211, and the user adjusts the relative positions of the hook portion 311, the connecting portion 3121, and the holding portion 3122 in three-dimensional space through the rotating shaft mechanism 31211, and further meets their wearing requirements.

[0057] It should be noted that considering the stability and comfort of the earphone 300 in terms of wearing, various changes and modifications can also be made to the earphone 300 (the fixing structure 310). For more descriptions of the earphone 300, reference can be made to the related application with the application number PCT / CN2021 / 109154, the content of which is incorporated into this application by reference.

[0058] In some embodiments, the earphone 300 can utilize the first microphone array 320 and the processor 330 to estimate the sound field at the user's ear canal (e.g., the target spatial position), and output a target signal through the speaker 340 to reduce the ambient noise at the user's ear canal, thereby achieving active noise reduction of the earphone 300. In some embodiments, the first microphone array 320 can be located in the body portion 312 of the fixed structure 310, such that when the user wears the earphone 300, the first microphone array 320 can be located near the user's ear canal. The first microphone array 320 can pick up the ambient noise near the user's ear canal, and the processor 330 can further estimate the ambient noise at the target spatial position based on the ambient noise near the user's ear canal, e.g., the ambient noise at the user's ear canal. In some embodiments, the target signal output by the speaker 340 is also picked up by the first microphone array 320. To reduce the influence of the target signal output by the speaker 340 on the ambient noise picked up by the first microphone array 320, the first microphone array 320 can be located in a region where the intensity of the sound emitted by the speaker 340 in space is small or even minimum, e.g., the acoustic null position of the radiation sound field of the acoustic dipole formed by the earphone 300 (e.g., the sound outlet hole and the pressure relief hole). For specific details regarding the position of the first microphone array 320, reference can be made to other parts of this specification, e.g., Figures 10 - 13 and its related descriptions.

[0059] In some embodiments, the processor 330 can be located in the hook portion 311 or the body portion 312 of the fixed structure 310. The processor 330 is electrically connected to the first microphone array 320. The processor 330 can estimate the sound field at the target spatial position based on the ambient noise picked up by the first microphone array 320, and generate a noise reduction signal based on the sound field estimation at the target spatial position. For specific details regarding the processor 330 using the first microphone array 320 to estimate the sound field at the target spatial position, reference can be made to this specification Figures 14 - 16 , and its related descriptions.

[0060] In some embodiments, the processor 330 can also be used to control the sound emission of the speaker 340. The processor 330 can control the sound emission of the speaker 340 according to the instruction input by the user. Alternatively, the processor 330 can generate an instruction for controlling the speaker 340 based on information of one or more components of the earphone 300. In some embodiments, the processor 330 can control other components of the earphone 300 (e.g., the battery). In some embodiments, the processor 330 can be disposed at any part of the fixed structure 310. For example, the processor 330 can be disposed in the holding portion 3122. In this case, the wiring distance between the processor 330 and other components (e.g., the speaker 340, the button switch, etc.) disposed on the holding portion 3122 can be shortened to reduce signal interference between the wirings and reduce the possibility of short circuit between the wirings.

[0061] In some embodiments, the speaker 340 may be located in the holding portion 3122 of the body portion 312 such that when the user wears the earphone 300, the speaker 340 may be located near the user's ear canal. The speaker 340 may output a target signal based on the noise reduction signal generated by the processor 330. The target signal may be transmitted to the outside of the earphone 300 through a sound outlet hole (not shown) on the holding portion 3122 for reducing the ambient noise at the user's ear canal. The sound outlet hole on the holding portion 3122 may be located on the side of the holding portion 3122 facing the user's ear. In this way, the sound outlet hole may be close enough to the user's ear canal, and the sound emitted therefrom can be better heard by the user.

[0062] In some embodiments, the earphone 300 may further include components such as a battery 350. The battery 350 may supply electrical energy to other components of the earphone 300 (such as the first microphone array 320, the speaker 340, etc.). In some embodiments, any two of the first microphone array 320, the processor 330, the speaker 340, and the battery 350 may communicate in various ways, for example, wired connection, wireless connection, or a combination thereof. In some embodiments, the wired connection may include a metal cable, an optical cable, or a hybrid cable of metal and optical, etc. The examples described above are only for convenience of illustration, and the medium of the wired connection may also be other types, for example, a transmission carrier of other electrical signals or optical signals, etc. The wireless connection may include radio communication, free space optical communication, acoustic communication, electromagnetic induction, etc.

[0063] In some embodiments, the battery 350 may be disposed at one end of the hook portion 311 away from the connecting portion 3121 and be located between the rear side of the user's ear and the head when the earphone 300 is in a worn state. In this setting mode, the capacity of the battery 350 can be increased, and the battery life of the earphone 300 can be improved. At the same time, the weight of the earphone 300 can also be balanced to facilitate overcoming the self-weight of the holding portion 3122 and the structures such as the processor 330 and the speaker 340 therein, thereby improving the stability and comfort of the earphone 300 in terms of wearing. In some embodiments, the battery 350 may also transmit its own status information to the processor 330 and receive instructions from the processor 330 to perform corresponding operations. The status information of the battery 350 may include on / off state, remaining power, remaining power usage time, charging time, etc., or a combination thereof.

[0064] To facilitate the description of the interrelationships of the various parts of the earphone (e.g., earphone 300) and the relationship between the earphone and the user, one or more coordinate systems are established in this specification. In some embodiments, three basic cutting planes, namely the Sagittal Plane, the Coronal Plane, and the Horizontal Plane, and three basic axes, namely the Sagittal Axis, the Coronal Axis, and the Vertical Axis, can be defined similar to those in the medical field. Refer to Figures 2 - 4 the coordinate axes in it. Among them, the sagittal plane refers to the cutting plane perpendicular to the ground along the front-back direction of the body, which divides the human body into left and right parts. In the embodiments of this specification, the sagittal plane can refer to the YZ plane, that is, the X axis is perpendicular to the user's sagittal plane; the coronal plane refers to the cutting plane perpendicular to the ground along the left-right direction of the body, which divides the human body into front and back parts. In the embodiments of this specification, the coronal plane can refer to the XZ plane, that is, the Y axis is perpendicular to the user's coronal plane; the horizontal plane refers to the cutting plane parallel to the ground along the up-down direction of the body, which divides the human body into upper and lower parts. In the embodiments of this specification, the horizontal plane can refer to the XY plane, that is, the Z axis is perpendicular to the user's horizontal plane. Correspondingly, the sagittal axis refers to the axis perpendicular to the coronal plane along the front-back direction of the body. In the embodiments of this specification, the sagittal axis can refer to the Y axis; the coronal axis refers to the axis perpendicular to the sagittal plane along the left-right direction of the body. In the embodiments of this specification, the coronal axis can refer to the X axis; the vertical axis refers to the axis perpendicular to the horizontal plane along the up-down direction of the body. In the embodiments of this specification, the vertical axis can refer to the Z axis.

[0065] Figure 5 is a structural diagram of an exemplary earphone shown in some embodiments of the present application. Figure 6 is a wearing diagram of an exemplary earphone shown in some embodiments of the present application.

[0066] Refer to Figures 5 - 6 , in some embodiments, the hook portion 311 can be close to the holding portion 3122 so that when the earphone 300 is in a worn state, as Figure 6 shown, the free end of the hook portion 311 facing away from the connecting portion 3121 acts on the first side (rear side) of the user's ear 100.

[0067] In some embodiments, refer to Figures 4 - 6, the connecting portion 3121 is connected to the hook portion 311, and the connecting portion 3121 and the hook portion 311 form a first connection point C. In the direction from the first connection point C between the hook portion 311 and the connecting portion 3121 to the free end of the hook portion 311, the hook portion 311 is bent backward toward the rear side of the ear portion 100 and forms a first contact point B with the rear side of the ear portion 100, and the holding portion 3122 forms a second contact point F with the second side (front side) of the ear portion 100. Among them, in the natural state (i.e., the non-wearing state), the distance between the first contact point B and the second contact point F in the extending direction of the connecting portion 3121 is smaller than the distance between the first contact point B and the second contact point F in the extending direction of the connecting portion 3121 in the wearing state. Furthermore, a pressing force on the second side (front side) of the ear portion 100 is provided for the holding portion 3122, and a pressing force on the first side (rear side) of the ear portion 100 is provided for the hook portion 311. It can also be understood that the distance between the first contact point B and the second contact point F in the extending direction of the connecting portion 3121 of the earphone 300 in the natural state is smaller than the thickness of the user's ear portion 100, so that the earphone 300 can be clamped on the user's ear portion 100 like a "clip" in the wearing state.

[0068] In some embodiments, the hook portion 311 can also extend in a direction away from the connecting portion 3121, that is, the overall length of the hook portion 311 is extended. When the earphone 300 is in the wearing state, the hook portion 311 can also form a third contact point A with the rear side of the ear portion 100. The first contact point B is located between the first connection point C and the third contact point A and is close to the first connection point C. Among them, in the natural state, the distance between the projections of the first contact point B and the third contact point A on a reference plane (such as the YZ plane) perpendicular to the extending direction of the connecting portion 3121 can be smaller than the distance between the projections of the first contact point B and the third contact point A on a reference plane (such as the YZ plane) perpendicular to the extending direction of the connecting portion 3121 in the wearing state. In this setting, the free end of the hook portion 311 presses against the rear side of the user's ear portion 100, which can make the third contact point A located in the area of the ear portion 100 close to the earlobe. Furthermore, the hook portion 311 can clamp the user's ear portion 100 in the vertical direction (Z-axis direction) to overcome the self-weight of the holding portion 3122. In some embodiments, after the overall length of the hook portion 311 is extended, while clamping the user's ear portion 100 in the vertical direction, the contact area between the hook portion 311 and the user's ear portion 100 can also be increased, that is, the friction between the hook portion 311 and the user's ear portion 100 is increased, thereby improving the stability of the earphone 300 in terms of wearing.

[0069] In some embodiments, a connecting portion 3121 is provided between the hook portion 311 and the holding portion 3122 of the earphone 300. When the earphone 300 is in a worn state, the cooperation between the connecting portion 3121 and the hook portion 311 can provide a pressing force on the first side of the ear for the holding portion 3122, so that when the earphone 300 is in a worn state, it can firmly adhere to the user's ear, thereby improving the wearing stability of the earphone 300 and the reliability of the earphone 300 in sound generation.

[0070] Figure 7 is a structural diagram of an exemplary earphone shown in some embodiments of the present application. Figure 8 is a wearing diagram of an exemplary earphone shown in some embodiments of the present application.

[0071] In some embodiments, Figures 7 - 8 the earphone 300 shown is substantially the same as Figures 5 - 6 the earphone 300 shown, except that the bending direction of the hook portion 311 is different. In some embodiments, referring to Figures 7 - 8 , in the direction from the first connection point C between the hook portion 311 and the connecting portion 3121 to the free end (the end far from the connecting portion 3121) of the hook portion 311, the hook portion 311 bends towards the user's head and forms a first contact point B and a third contact point A with the head. Among them, the first contact point B is located between the third contact point A and the first connection point C. With such a setting, the hook portion 311 can form a lever structure with the first contact point B as the fulcrum. At this time, the free end of the hook portion 311 presses against the user's head, and the user's head provides a force pointing outward from the head at the third contact point A. This force is converted into a force pointing towards the head at the first connection point C through the lever structure, and then provides a pressing force on the first side of the ear 100 for the holding portion 3122 through the connecting portion 3121.

[0072] In some embodiments, the magnitude of the force provided by the user's head pointing outward from the head at the third contact point A is positively correlated with the magnitude of the angle formed between the free end of the hook portion 311 and the YZ plane when the earphone 300 is in a non-worn state. Specifically, the larger the angle formed between the free end of the hook portion 311 and the YZ plane when the earphone 300 is in a non-worn state, the better the free end of the hook portion 311 can press against the user's head when the earphone 300 is in a worn state, and the correspondingly larger the force that the user's head can provide pointing outward from the head at the third contact point A. In some embodiments, in order to enable the free end of the hook portion 311 to press against the user's head when the earphone 300 is in a worn state and enable the user's head to provide a force pointing outward from the head at the third contact point A, the angle formed between the free end of the hook portion 311 and the YZ plane when the earphone 300 is in a non-worn state can be greater than the angle formed between the free end of the hook portion 311 and the YZ plane when the earphone 300 is in a worn state.

[0073] In some embodiments, when the free end of the hook portion 311 presses against the user's head, in addition to causing the user's head to provide a force pointing outward from the head at the third contact point A, it also causes at least another pressing force to be formed on the first side of the ear portion 100 by the hook portion 311, and can cooperate with the pressing force formed on the second side of the ear portion 100 by the holding portion 3122 to form a "front and back clamping" pressing effect on the user's ear portion 100, thereby improving the wearing stability of the earphone 300.

[0074] It should be noted that during actual wearing, due to differences in the physiological structures of the heads and ears of different users, there will be a certain impact on the actual wearing of the earphone 300, and the positions of the contact points between the earphone 300 and the user's head or ear (for example, the first contact point B, the second contact point F, the third contact point A, etc.) can change accordingly.

[0075] In some embodiments, when the speaker 340 is located in the holding portion 3122, due to differences in the physiological structures of the heads and ears of different users, there will be a certain impact on the actual wearing of the earphone 300. Therefore, when different users wear the earphone 300, the relative position between the speaker 340 and the user's ear will change. In some embodiments, the structure of the holding portion 3122 can be set to adjust the position of the speaker 340 in the overall structure of the earphone 300, thereby adjusting the distance between the speaker 340 and the user's ear canal.

[0076] Figure 9A is a structural diagram of an exemplary earphone shown according to some embodiments of the present application. Figure 9B is a structural diagram of an exemplary earphone shown according to some embodiments of the present application.

[0077] Refer to Figure 9A and Figure 9B , the holding part 3122 can be designed as a multi-segment structure to adjust the relative position of the speaker 340 on the overall structure of the earphone 300. In some embodiments, the holding part 3122 being a multi-segment structure enables the earphone 300 to be in a worn state, not blocking the external auditory canal of the ear while allowing the speaker 340 to be as close as possible to the external auditory canal, thereby improving the auditory experience of the user when using the earphone 300.

[0078] Refer to Figure 9A , in some embodiments, the holding part 3122 may include a first holding segment 3122-1, a second holding segment 3122-2, and a third holding segment 3122-3 that are sequentially connected end to end. Among them, one end of the first holding segment 3122-1 facing away from the second holding segment 3122-2 is connected to the connecting part 3121. The second holding segment 3122-2 is folded back relative to the first holding segment 3122-1, such that there is a spacing between the second holding segment 3122-2 and the first holding segment 3122-1. In some embodiments, the structure between the second holding segment 3122-2 and the first holding segment 3122-1 may be in a U-shaped configuration. The third holding segment 3122-3 is connected to one end of the second holding segment 3122-2 facing away from the first holding segment 3122-1, and the third holding segment 3122-3 can be used to arrange structural components such as the speaker 340.

[0079] In some embodiments, refer to Figure 9A , in this setting, by adjusting the spacing between the second holding segment 3122-2 and the first holding segment 3122-1, the folding length of the second holding segment 3122-2 folded back relative to the first holding segment 3122-1 (the length of the second holding segment 3122-2 along the Y-axis direction), etc., the position of the third holding segment 3122-3 on the overall structure of the earphone 300 can be adjusted, thereby adjusting the position or distance of the speaker 340 located on the third holding segment 3122-3 relative to the user's ear canal. In some embodiments, the spacing between the second holding segment 3122-2 and the first holding segment 3122-1, and the folding length of the second holding segment 3122-2 folded back relative to the first holding segment 3122-1 can be set accordingly according to the ear characteristics (such as shape, size, etc.) of different users, and no specific limitation is made here.

[0080] Refer to Figure 9B, in some embodiments, the holding portion 3122 may include a first holding segment 3122-1, a second holding segment 3122-2, and a third holding segment 3122-3 that are sequentially connected end to end. Among them, one end of the first holding segment 3122-1 facing away from the second holding segment 3122-2 is connected to the connecting portion 3121. The second holding segment 3122-2 is bent relative to the first holding segment 3122-1, such that there is a spacing between the third holding segment 3122-3 and the first holding segment 3122-1. The third holding segment 3122-3 can be used to arrange structural components such as the speaker 340.

[0081] In some embodiments, referring to Figure 9B , in this setting manner, by adjusting the spacing between the third holding segment 3122-3 and the first holding segment 3122-1, the bending length of the second holding segment 3122-2 bent relative to the first holding segment 3122-1 (the length of the second holding segment 3122-2 along the Z-axis direction), etc., the position of the third holding segment 3122-3 on the overall structure of the earphone 300 can be adjusted, so as to adjust the position or distance of the speaker 340 located on the third holding segment 3122-3 relative to the user's ear canal. In some embodiments, the spacing between the third holding segment 3122-3 and the first holding segment 3122-1, and the bending length of the second holding segment 3122-2 bent relative to the first holding segment 3122-1 can be correspondingly set according to the ear characteristics (such as shape, size, etc.) of different users, and no specific limitation is made here.

[0082] Figure 10 is a structural diagram of the exemplary earphone facing the ear side shown according to some embodiments of the present application.

[0083] In some embodiments, referring to Figure 10 , an acoustic hole 301 may be provided on the side of the holding portion 3122 facing the ear, and the target signal output by the speaker 340 can be transmitted to the user's ear through the acoustic hole 301. In some embodiments, the side of the holding portion 3122 facing the ear may include a first region 3122A and a second region 3122B. The second region 3122B is farther from the connecting portion 3121 than the first region 3122A, that is, the second region 3122B may be located at the free end of the holding portion 3122 away from the connecting portion 3121. In some embodiments, there may be a smooth transition between the first region 3122A and the second region 3122B. In some embodiments, the acoustic hole 301 may be provided in the first region 3122A, and the second region 3122B protrudes toward the ear compared to the first region 3122A, such that the second region 3122B contacts the ear to allow the acoustic hole 301 to be spaced from the ear in the wearing state.

[0084] In some embodiments, the free end of the holding portion 3122 may be configured as a convex hull structure. On the side of the holding portion 3122 close to the user's ear, the convex hull structure protrudes outward (i.e., towards the user's ear) relative to this side. Since the speaker 340 can generate sound (e.g., a target signal) transmitted to the ear through the sound outlet hole 301, the convex hull structure can prevent the ear from blocking the sound outlet hole 301, which may cause the sound generated by the speaker 340 to weaken or even be unable to be output. In some embodiments, in the thickness direction (X-axis direction) of the holding portion 3122, the protrusion height of the convex hull structure may be represented by the maximum protrusion height of the second region 3122B relative to the first region 3122A. In some embodiments, the maximum protrusion height of the second region 3122B relative to the first region 3122A may be greater than or equal to 1 mm. In some embodiments, in the thickness direction of the holding portion 3122, the maximum protrusion height of the second region 3122B relative to the first region 3122A may be greater than or equal to 0.8 mm. In some embodiments, in the thickness direction of the holding portion 3122, the maximum protrusion height of the second region 3122B relative to the first region 3122A may be greater than or equal to 0.5 mm.

[0085] In some embodiments, by setting the structure of the holding portion 3122, when the user wears the earphone 300, the distance between the sound outlet hole 301 and the user's ear canal can be made less than 10 mm. In some embodiments, in order to ensure that the sound transmitted to the ear canal through the sound outlet hole 301 heard by the user is clearer, by setting the structure of the holding portion 3122, when the user wears the earphone 300, the distance between the sound outlet hole 301 and the user's ear canal can be made less than 8 mm. In some embodiments, considering the performance of the earphone 300 and the comfort during wearing, by setting the structure of the holding portion 3122, when the user wears the earphone 300, the distance between the sound outlet hole 301 and the user's ear canal can be made less than 7 mm. In some embodiments, in order to further improve the auditory experience of the user when using the earphone 300, by setting the structure of the holding portion 3122, when the user wears the earphone 300, the distance between the sound outlet hole 301 and the user's ear canal can be made less than 6 mm.

[0086] It should be noted that, if the purpose is only to space the sound hole 301 from the ear when worn, then the raised area facing the ear compared to the first area 3122A may also be located in other areas of the retaining portion 3122, such as the area between the sound hole 301 and the connecting portion 3121. In some embodiments, since the concha cavity and the hymena concha have a certain depth and are connected to the ear hole, the orthographic projection of the sound hole 301 on the ear along the thickness direction of the retaining portion 3122 may at least partially fall within the concha cavity and / or the hymena concha. As an exemplary description only, when the user wears the earphone 300, the retaining portion 3122 may be located on the side of the ear hole close to the top of the user's head and in contact with the antihelix, and at this time, the orthographic projection of the sound hole 301 on the ear along the thickness direction of the retaining portion 3122 may at least partially fall within the hymena concha.

[0087] Figure 11 It is a structural diagram of the side of an exemplary earphone facing away from the ear according to some embodiments of the present application. Figure 12 is a top view of an exemplary headset according to some embodiments of the present application.

[0088] See also Figures 11 - 12 , a pressure relief hole 302 may be provided on one side of the retaining portion 3122 along the vertical axis (Z axis) direction and close to the top of the user's head, and the pressure relief hole 302 is further away from the user's ear canal than the sound outlet 301. In some embodiments, the opening direction of the pressure relief hole 302 may be toward the top of the user's head, and a specific angle may be provided between the opening direction of the pressure relief hole 302 and the vertical axis (Z axis) to allow the pressure relief hole 302 to be further away from the user's ear canal, thereby making it difficult for the user to hear the sound output through the pressure relief hole 302 and transmitted to the user's ear. In some embodiments, the angle between the opening direction of the pressure relief hole 302 and the vertical axis (Z axis) may be 0° to 10°. In some embodiments, in order to make the pressure relief hole 302 further away from the user's ear canal, the angle between the opening direction of the pressure relief hole 302 and the vertical axis (Z axis) may be 0° to 8°. In some embodiments, in order to ensure that the sound output through the pressure relief hole 302 heard by the user is smaller, the angle between the opening direction of the pressure relief hole 302 and the vertical axis (Z axis) may be 0° to 5°.

[0089] In some embodiments, by setting the structure of the holding portion 3122 and the angle between the opening direction of the pressure relief hole 302 and the vertical axis (Z-axis), when the user wears the earphone 300, the distance between the pressure relief hole 302 and the user's ear canal can be within a suitable range. In some embodiments, to ensure that the sound output through the pressure relief hole 302 and transmitted to the user's ear canal is small enough for the user to hear, when the user wears the earphone 300, the distance between the pressure relief hole 302 and the user's ear canal can be 5 millimeters to 20 millimeters. In some embodiments, based on the structure and wearing method of the earphone 300, when the user wears the earphone 300, the distance between the pressure relief hole 302 and the user's ear canal can be 5 millimeters to 18 millimeters. In some embodiments, to further improve the auditory experience when the user wears the earphone 300, when the user wears the earphone 300, the distance between the pressure relief hole 302 and the user's ear canal can be 5 millimeters to 15 millimeters.

[0090] Figure 13 is a schematic cross-sectional structure diagram of an exemplary earphone shown in some embodiments of the present application.

[0091] Figure 13 shows the acoustic structure formed by the holding portion (e.g., the holding portion 3122) of the earphone (e.g., the earphone 300), including: the sound outlet hole 301, the pressure relief hole 302, the sound tuning hole 303, the front cavity 304, and the rear cavity 305.

[0092] In some embodiments, in combination with Figure 11 and Figure 13 , the holding portion 3122 can form a front cavity 304 and a rear cavity 305 on the opposite sides of the speaker 340 respectively. The front cavity 304 is communicated with the outside of the earphone 300 through the sound outlet hole 301 and outputs sound (e.g., the target signal, the audio signal, etc.) to the ear. The rear cavity 305 is communicated with the outside of the earphone 300 through the pressure relief hole 302, and the pressure relief hole 302 is farther from the user's ear canal than the sound outlet hole 301. In some embodiments, the pressure relief hole 302 can allow air to freely enter and exit the rear cavity 305, so that the change in air pressure in the front cavity 304 can be blocked by the rear cavity 305 as little as possible, thereby improving the sound quality of the sound output to the ear through the sound outlet hole 301.

[0093] In some embodiments, by setting the positions of the pressure relief hole 302 and the sound outlet hole 301 on the holding portion 3122, the included angle between the line connecting the pressure relief hole 302 and the sound outlet hole 301 and the thickness direction (X-axis direction) of the holding portion 3122 can be made to be 0° to 50°. In some embodiments, considering the overall size (e.g., thickness) of the holding portion 3122, the included angle between the line connecting the pressure relief hole 302 and the sound outlet hole 301 and the thickness direction of the holding portion 3122 can be 5° to 45°. It should be noted that the included angle between the line connecting the pressure relief hole 302 and the sound outlet hole 301 and the thickness direction of the holding portion 3122 can be the included angle between the line connecting the center of the pressure relief hole 302 and the center of the sound outlet hole 301 and the thickness direction of the holding portion 3122.

[0094] In some embodiments, in combination with Figure 11 and Figure 13 , the sound outlet hole 301 and the pressure relief hole 302 can be regarded as two sound sources that radiate sound outward, with the same amplitude and opposite phases of the radiated sound. The two sound sources can approximately form an acoustic dipole or something similar to an acoustic dipole, and thus the sound radiated outward by them has obvious directivity, forming an "8"-shaped sound radiation region. In the direction of the straight line where the connection line of the two sound sources is located, the sound radiated by the two sound sources is the largest, and the sound radiated in the remaining directions is significantly reduced, and the sound radiated at the perpendicular bisector of the connection line of the two sound sources is the smallest. That is, in the direction of the straight line where the connection line of the pressure relief hole 302 and the sound outlet hole 301 is located, the sound radiated by the pressure relief hole 302 and the sound outlet hole 301 is the largest, the sound radiated in the remaining directions is significantly reduced, and the sound radiated at the perpendicular bisector of the connection line of the pressure relief hole 302 and the sound outlet hole 301 is the smallest. In some embodiments, the acoustic dipole formed by the pressure relief hole 302 and the sound outlet hole 301 can reduce the sound leakage of the speaker 340.

[0095] In some embodiments, in combination with Figure 11 and Figure 13 , the holding portion 3122 can also be provided with a sound tuning hole 303 communicating with the rear cavity 305. The sound tuning hole 303 can be used to break the high-pressure area in the sound field of the rear cavity 305, so that the wavelength of the standing wave in the rear cavity 305 becomes shorter, and further make the resonance frequency of the sound output to the outside of the earphone 300 through the pressure relief hole 302 as high as possible, such as greater than 4 kHz, thereby reducing the sound leakage of the speaker 340. In some embodiments, the sound tuning hole 303 and the pressure relief hole 302 can be respectively located on opposite sides of the speaker 340, for example, arranged in opposite directions in the Z-axis direction, to break the high-pressure area in the sound field of the rear cavity 305 to the greatest extent. In some embodiments, the sound tuning hole 303 can be farther away from the sound outlet hole 301 than the pressure relief hole 302, so as to increase the distance between the sound tuning hole 303 and the sound outlet hole 301 as much as possible, and then weaken the anti-phase cancellation between the sound output to the outside of the earphone 300 through the sound tuning hole 303 and the sound transmitted to the ear through the sound outlet hole 301.

[0096] In some embodiments, the target signal output by the speaker 340 through the sound outlet hole 301 and / or the pressure relief hole 302 may also be picked up by the first microphone array 320, and this target signal will affect the processor 330's estimation of the sound field in the target space position, that is, the target signal output by the speaker 340 is not expected to be picked up. In this case, in order to reduce the influence of the target signal output by the speaker 340 on the first microphone array 320, the first microphone array 320 may be disposed in a first target area where the sound output by the speaker 340 is as small as possible. In some embodiments, the first target area may be the acoustic null position of the radiation sound field of the acoustic dipole formed by the pressure relief hole 302 and the sound outlet hole 301 or a position near it. In some embodiments, the first target area may be Figure 10 the area G shown in. When the user wears the earphone 300, the area G is located in front of the sound outlet hole 301 and / or the pressure relief hole 302 (the front here refers to the direction the user is facing), that is, the area G is closer to the user's eyes. Optionally, the area G may be a partial area on the connecting portion 3121 of the fixed structure 310. That is to say, the first microphone array 320 may be located at the connecting portion 3121. For example, the first microphone array 320 may be located at a position on the connecting portion 3121 close to the holding portion 3122. In some alternative embodiments, the area G may also be located behind the sound outlet hole 301 and / or the pressure relief hole 302 (the front here refers to the opposite direction of the direction the user is facing). For example, the area G may be located at the end of the holding portion 3122 far from the connecting portion 3121.

[0097] In some embodiments, referring to Figures 10 - 11 , in order to reduce the influence of the target signal output by the speaker 340 on the first microphone array 320 and improve the active noise reduction effect of the earphone 300, the relative positions between the first microphone array 320, the sound outlet hole 301, and the pressure relief hole 302 may be reasonably set. The position of the first microphone array 320 mentioned here may be the position where any microphone in the first microphone array 320 is located. In some embodiments, the line connecting the first microphone array 320 and the sound outlet hole 301 forms a first included angle with the line connecting the sound outlet hole 301 and the pressure relief hole 302, and the line connecting the first microphone array 320 and the pressure relief hole 302 forms a second included angle with the line connecting the sound outlet hole 301 and the pressure relief hole 302. In some embodiments, in order to make the first microphone array 320 less affected by the speaker 340, the difference between the first included angle and the second included angle may not be greater than 30°. In some embodiments, considering the overall size of the holding portion 3122 of the earphone 300, the difference between the first included angle and the second included angle may not be greater than 20°. In some embodiments, in order to ensure better noise reduction effect of the earphone 300, the difference between the first included angle and the second included angle may not be greater than 10°.

[0098] In some embodiments, there is a first distance between the first microphone array 320 and the sound outlet hole 301, and a second distance between the first microphone array 320 and the pressure relief hole 302. To ensure that the target signal output by the speaker 340 has a relatively small impact on the first microphone array 320, the difference between the first distance and the second distance may not be greater than 6 millimeters. In some embodiments, considering the overall size of the holding portion 3122 of the earphone 300, the difference between the first distance and the second distance may not be greater than 5 millimeters. In some embodiments, to ensure better noise reduction effect of the earphone 300, the difference between the first distance and the second distance may not be greater than 3 millimeters.

[0099] It can be understood that the positional relationship between the first microphone array 320 described herein and the sound outlet hole 301 and the pressure relief hole 302 may refer to the positional relationship between any microphone in the first microphone array 320 and the centers of the sound outlet hole 301 and the pressure relief hole 302. For example, the first included angle formed by the connection line between the first microphone array 320 and the sound outlet hole 301 and the connection line between the sound outlet hole 301 and the pressure relief hole 302 may refer to the first included angle formed by the connection line between any microphone in the first microphone array 320 and the center of the sound outlet hole 301 and the connection line between the center of the sound outlet hole 301 and the center of the pressure relief hole 302. Another example, the first distance between the first microphone array 320 and the sound outlet hole 301 may refer to the first distance between any microphone in the first microphone array 320 and the center of the sound outlet hole 301.

[0100] In some embodiments, the first microphone array 320 is located at the acoustic null position of the acoustic dipole formed by the sound outlet hole 301 and the pressure relief hole 302, which can minimize the influence of the target signal output by the speaker 340 on the first microphone array 320, and further enable the first microphone array 320 to pick up the ambient noise near the user's ear canal more accurately. Further, the processor 330 can more accurately estimate the ambient noise at the user's ear canal based on the ambient noise picked up by the first microphone array 320 and generate a noise reduction signal, so as to better achieve the active noise reduction of the earphone 300. For the specific description of using the first microphone array 320 to achieve the active noise reduction of the earphone 300, reference can be made to Figures 14 - 16 , and its related description.

[0101] Figure 14 is an exemplary noise reduction flowchart of the earphone shown in some embodiments of the present application. In some embodiments, the process 1400 can be executed by the earphone 300. As Figure 14 shown, the process 1400 may include:

[0102] In step 1410, pick up the ambient noise. In some embodiments, this step can be executed by the first microphone array 320.

[0103] In some embodiments, ambient noise may refer to a combination of various external sounds in the environment where the user is located (e.g., traffic noise, industrial noise, construction noise, social noise). In some embodiments, the first microphone array 320 may be located in the body portion 312 of the earphone 300 near the user's ear canal for picking up the ambient noise at a position near the user's ear canal. Further, the first microphone array 320 may convert the picked-up ambient noise signal into an electrical signal and transmit it to the processor 330 for processing.

[0104] In step 1420, the noise of the target spatial position is estimated based on the picked-up ambient noise. In some embodiments, this step may be executed by the processor 330.

[0105] In some embodiments, the processor 330 may perform signal separation on the picked-up ambient noise. In some embodiments, the ambient noise picked up by the first microphone array 320 may include various sounds. The processor 330 may perform signal analysis on the ambient noise picked up by the first microphone array 320 to separate various sounds. Specifically, the processor 330 may adaptively adjust the parameters of the filter according to the statistical distribution characteristics and structural characteristics of various sounds in different dimensions such as space, time domain, and frequency domain, estimate the parameter information of each sound signal in the ambient noise, and complete the signal separation process according to the parameter information of each sound signal. In some embodiments, the statistical distribution characteristics of the noise may include probability distribution density, power spectral density, autocorrelation function, probability density function, variance, mathematical expectation, etc. In some embodiments, the structural characteristics of the noise may include noise distribution, noise intensity, global noise intensity, noise rate, etc., or any combination thereof. The global noise intensity may refer to the average noise intensity or the weighted average noise intensity. The noise rate may refer to the degree of dispersion of the noise distribution. Only by way of example, the ambient noise picked up by the first microphone array 320 may include a first signal, a second signal, and a third signal. The processor 330 obtains the differences of the first signal, the second signal, and the third signal in space (e.g., the position where the signal is located), time domain (e.g., delay), and frequency domain (e.g., amplitude, phase), and separates the first signal, the second signal, and the third signal according to the differences in the three dimensions to obtain relatively pure first signal, second signal, and third signal. Further, the processor 330 may update the ambient noise according to the parameter information (e.g., frequency information, phase information, amplitude information) of the separated signals. For example, the processor 330 may determine that the first signal is the user's call voice according to the parameter information of the first signal and remove the first signal from the ambient noise to update the ambient noise. In some embodiments, the removed first signal may be transmitted to the far end of the call. For example, when the user wears the earphone 300 for a voice call, the first signal may be transmitted to the far end of the call.

[0106] The target spatial position is a position located in or near the user's ear canal, determined based on the first microphone array 320. The target spatial position may refer to a spatial position at a specific distance (e.g., 2 mm, 3 mm, 5 mm, etc.) from the user's ear canal (e.g., the ear canal opening). In some embodiments, the target spatial position is closer to the user's ear canal than any of the microphones in the first microphone array 320. In some embodiments, the target spatial position is related to the number of microphones in the first microphone array 320 and their distribution positions relative to the user's ear canal, and the target spatial position can be adjusted by adjusting the number of microphones in the first microphone array 320 and / or their distribution positions relative to the user's ear canal. In some embodiments, estimating the noise at the target spatial position based on the picked-up ambient noise (or updated ambient noise) may further include determining one or more spatial noise sources related to the picked-up ambient noise and estimating the noise at the target spatial position based on the spatial noise sources. The ambient noise picked up by the first microphone array 320 can be spatial noise sources from different directions and of different types. The parameter information (e.g., frequency information, phase information, amplitude information) corresponding to each spatial noise source is different.

[0107] In some embodiments, the processor 330 may separate and extract the noise at the target spatial position according to the statistical distribution and structural characteristics of different types of noise in different dimensions (e.g., spatial domain, time domain, frequency domain, etc.), so as to obtain different types of noise (e.g., different frequencies, different phases, etc.) and estimate the parameter information (e.g., amplitude information, phase information, etc.) corresponding to each type of noise.

[0108] In some embodiments, the processor 330 may also determine the overall parameter information of the noise at the target spatial position according to the parameter information corresponding to different types of noise at the target spatial position. More content regarding estimating the noise at the target spatial position based on one or more spatial noise sources can be found elsewhere in this specification, for example, Figure 15 and its corresponding description.

[0109] In some embodiments, estimating the noise at the target spatial position based on the picked-up ambient noise (or updated ambient noise) may further include constructing a virtual microphone based on the first microphone array 320 and estimating the noise at the target spatial position based on the virtual microphone. More content regarding estimating the noise at the target spatial position based on the virtual microphone can be found elsewhere in this specification, for example Figure 16 and its corresponding description.

[0110] In step 1430, a noise reduction signal is generated based on the noise at the target spatial position. In some embodiments, this step may be executed by the processor 330.

[0111] In some embodiments, the processor 330 may generate a noise reduction signal based on the parameter information of the noise at the target spatial position obtained in step 1420 (e.g., amplitude information, phase information, etc.). In some embodiments, the phase difference between the phase of the noise reduction signal and the phase of the noise at the target spatial position may be less than or equal to a preset phase threshold. The preset phase threshold may be in the range of 90 - 180 degrees. The preset phase threshold may be adjusted within this range according to the user's needs. For example, when the user does not want to be disturbed by the sounds of the surrounding environment, the preset phase threshold may be a larger value, such as 180 degrees, that is, the phase of the noise reduction signal is opposite to the phase of the noise at the target spatial position. For another example, when the user wants to be sensitive to the surrounding environment, the preset phase threshold may be a smaller value, such as 90 degrees. It should be noted that the more sounds of the surrounding environment the user wants to receive, the closer the preset phase threshold may be to 90 degrees, and the less sounds of the surrounding environment the user wants to receive, the closer the preset phase threshold may be to 180 degrees. In some embodiments, when the phase of the noise reduction signal and the phase of the noise at the target spatial position are in a certain relationship (e.g., opposite phases), the amplitude difference between the amplitude of the noise at the target spatial position and the amplitude of the noise reduction signal may be less than or equal to a preset amplitude threshold. For example, when the user does not want to be disturbed by the sounds of the surrounding environment, the preset amplitude threshold may be a smaller value, such as 0 dB, that is, the amplitude of the noise reduction signal is equal to the amplitude of the noise at the target spatial position. For another example, when the user wants to be sensitive to the surrounding environment, the preset amplitude threshold may be a larger value, such as approximately equal to the amplitude of the noise at the target spatial position. It should be noted that the more sounds of the surrounding environment the user wants to receive, the closer the preset amplitude threshold may be to the amplitude of the noise at the target spatial position, and the less sounds of the surrounding environment the user wants to receive, the closer the preset amplitude threshold may be to 0 dB.

[0112] In some embodiments, the speaker 340 may output a target signal based on the noise reduction signal generated by the processor 330. For example, the speaker 340 may convert the noise reduction signal (e.g., an electrical signal) into a target signal (i.e., a vibration signal) based on its vibration component, and the target signal is transmitted to the user's ear through the sound outlet hole 301 on the earphone 300 and cancels out the ambient noise at the user's ear canal. In some embodiments, when the noise at the target spatial position is from multiple spatial noise sources, the speaker 340 may output target signals corresponding to the multiple spatial noise sources based on the noise reduction signal. For example, when the multiple spatial noise sources include a first spatial noise source and a second spatial noise source, the speaker 340 may output a first target signal with a noise phase approximately opposite and an amplitude approximately equal to that of the first spatial noise source to cancel the noise of the first spatial noise source, and a second target signal with a noise phase approximately opposite and an amplitude approximately equal to that of the second spatial noise source to cancel the noise of the second spatial noise source. In some embodiments, when the speaker 340 is an air-conduction speaker, the position where the target signal cancels out the ambient noise may be the target spatial position. The distance between the target spatial position and the user's ear canal is small, and the noise at the target spatial position can be approximately regarded as the noise at the user's ear canal position. Therefore, the cancellation of the noise reduction signal and the noise at the target spatial position can be approximately regarded as the elimination of the ambient noise transmitted to the user's ear canal, realizing the active noise reduction of the earphone 300. In some embodiments, when the speaker 340 is a bone-conduction speaker, the position where the target signal cancels out the ambient noise may be the basilar membrane. The target signal and the ambient noise are cancelled out at the user's basilar membrane, thereby realizing the active noise reduction of the earphone 300.

[0113] In some embodiments, when the position of the earphone 300 changes, for example, when the head of the user wearing the earphone 300 turns, the ambient noise (such as noise direction, amplitude, phase) changes accordingly, and it is difficult for the earphone 300 to execute noise reduction at a speed that can keep up with the change speed of the ambient noise, resulting in a weakened active noise reduction function of the earphone 300. For this reason, the earphone 300 may further include one or more sensors, and the one or more sensors can be located at any position of the earphone 300. For example, the hook portion 311 and / or the connecting portion 3121 and / or the holding portion 3122. The one or more sensors can be electrically connected to other components of the earphone 300 (such as the processor 330). In some embodiments, the one or more sensors can be used to obtain the physical position and / or motion information of the earphone 300. By way of example only, the one or more sensors can include an Inertial Measurement Unit (IMU), a Global Position System (GPS), a radar, etc. The motion information can include a motion trajectory, a motion direction, a motion speed, a motion acceleration, a motion angular velocity, motion-related time information (such as a motion start time, an end time), etc., or any combination thereof. Taking the IMU as an example, the IMU can include a Micro electro Mechanical System (MEMS). The micro electro mechanical system can include a multi-axis accelerometer, a gyroscope, a magnetometer, etc., or any combination thereof. The IMU can be used to detect the physical position and / or motion information of the earphone 300 to enable control of the earphone 300 based on the physical position and / or motion information.

[0114] In some embodiments, the processor 330 can update the noise at the target spatial position and the sound field estimation at the target spatial position based on the motion information (such as a motion trajectory, a motion direction, a motion speed, a motion acceleration, a motion angular velocity, motion-related time information) of the earphone 300 obtained by one or more sensors of the earphone 300. Further, based on the updated noise at the target spatial position and the sound field estimation at the target spatial position, the processor 330 can generate a noise reduction signal. The one or more sensors can record the motion information of the earphone 300, and then the processor 330 can quickly update the noise reduction signal, which can improve the noise tracking performance of the earphone 300, so that the noise reduction signal can more accurately cancel the ambient noise, further improving the noise reduction effect and the user's auditory experience.

[0115] It should be noted that the above description of process 1400 is only for illustration and example, and does not limit the scope of application of the present application. Those skilled in the art can make various modifications and changes to process 1400 under the guidance of the present application. For example, steps in process 1400 can also be added, omitted, or combined. These modifications and changes are still within the scope of the present application.

[0116] Figure 15 is an exemplary flowchart for estimating the noise of a target spatial position shown in some embodiments of the present application. As Figure 15 shown, process 1500 can include:

[0117] In step 1510, one or more spatial noise sources related to the ambient noise picked up by the first microphone array 320 are determined. In some embodiments, this step can be executed by the processor 330. As described herein, determining the spatial noise source refers to determining the spatial noise source related information, for example, the position of the spatial noise source (including the azimuth of the spatial noise source, the distance between the spatial noise source and the target spatial position, etc.), the phase of the spatial noise source, and the amplitude of the spatial noise source, etc.

[0118] In some embodiments, the spatial noise source related to the ambient noise refers to a noise source whose sound wave can be transmitted to or near the user's ear canal (for example, the target spatial position). In some embodiments, the spatial noise source can be noise sources in different directions of the user's body (for example, in front, behind, etc.). For example, there is a noisy crowd in front of the user's body and a vehicle horn noise on the left side of the user's body. In this case, the spatial noise sources include the noisy crowd noise source in front of the user's body and the vehicle horn noise source on the left side of the user's body. In some embodiments, the first microphone array 320 can pick up the spatial noise in all directions of the user's body, convert the spatial noise into an electrical signal and transmit it to the processor 330. The processor 330 can analyze the electrical signal corresponding to the spatial noise to obtain the parameter information (for example, frequency information, amplitude information, phase information, etc.) of the picked spatial noise in all directions. The processor 330 determines the information of the spatial noise sources in all directions according to the parameter information of the spatial noise in all directions, for example, the azimuth of the spatial noise source, the distance of the spatial noise source, the phase of the spatial noise source, and the amplitude of the spatial noise source, etc. In some embodiments, the processor 330 can determine the spatial noise source based on the spatial noise picked up by the first microphone array 320 through a noise localization algorithm. The noise localization algorithm can include one or more of a beamforming algorithm, a super-resolution spatial spectrum estimation algorithm, a time difference of arrival algorithm (which can also be called a time delay estimation algorithm), etc.

[0119] In some embodiments, the processor 330 may divide the picked-up ambient noise into multiple frequency bands according to a specific frequency bandwidth (e.g., every 500 Hz as a frequency band), each frequency band may respectively correspond to a different frequency range, and determine a spatial noise source corresponding to the frequency band on at least one frequency band. For example, the processor 330 may perform signal analysis on the frequency bands into which the ambient noise is divided, obtain parameter information of the ambient noise corresponding to each frequency band, and determine the spatial noise source corresponding to each frequency band according to the parameter information.

[0120] In step 1520, based on the spatial noise source, estimate the noise at the target spatial position. In some embodiments, this step may be performed by the processor 330. As described herein, estimating the noise at the target spatial position refers to estimating the parameter information of the noise at the target spatial position, for example, frequency information, amplitude information, phase information, etc.

[0121] In some embodiments, the processor 330 may estimate the parameter information of the noise transmitted from each spatial noise source to the target spatial position based on the parameter information (such as frequency information, amplitude information, phase information, etc.) of the spatial noise sources located in various directions of the user's body obtained in step 1510, so as to estimate the noise at the target spatial position. For example, there is a spatial noise source in the first direction (e.g., the front) and the second direction (e.g., the back) of the user's body. The processor 330 may estimate the frequency information, phase information, or amplitude information of the first-direction spatial noise source when the noise of the first-direction spatial noise source is transmitted to the target spatial position according to the position information, frequency information, phase information, or amplitude information of the first-direction spatial noise source. The processor 330 may estimate the frequency information, phase information, or amplitude information of the second-direction spatial noise source when the noise of the second-direction spatial noise source is transmitted to the target spatial position according to the position information, frequency information, phase information, or amplitude information of the second-direction spatial noise source. Further, the processor 330 may estimate the noise information at the target spatial position based on the frequency information, phase information, or amplitude information of the first-direction spatial noise source and the second-direction spatial noise source, so as to estimate the information of the noise at the target spatial position. Merely by way of example, the processor 330 may use virtual microphone technology or other methods to estimate the noise information at the target spatial position. In some embodiments, the processor 330 may extract the parameter information of the noise of the spatial noise source from the frequency response curve of the spatial noise source picked up by the microphone array by means of feature extraction. In some embodiments, the method for extracting the parameter information of the noise of the spatial noise source may include, but is not limited to, principal components analysis (PCA), independent component analysis (ICA), linear discriminant analysis (LDA), singular value decomposition (SVD), etc.

[0122] It should be noted that the above description of process 1500 is merely for illustration and example, and does not limit the scope of application of the present application. Those skilled in the art can make various modifications and changes to process 1500 under the guidance of the present application. For example, process 1500 may further include steps such as locating the spatial noise source and extracting the parameter information of the noise of the spatial noise source. These modifications and changes are still within the scope of the present application.

[0123] Figure 16 is an exemplary flowchart of estimating the sound field and noise at the target spatial position according to some embodiments of the present application. As Figure 16 shown, process 1600 may include:

[0124] In step 1610, a virtual microphone is constructed based on the first microphone array 320. In some embodiments, this step may be performed by the processor 330.

[0125] In some embodiments, the virtual microphone may be used to represent or simulate the audio data that would be collected by a microphone if it were placed at the target spatial position. That is, the audio data obtained through the virtual microphone can be approximated or equivalent to the audio data that would be collected by a physical microphone if it were placed at the target spatial position.

[0126] In some embodiments, the virtual microphone may include a mathematical model. This mathematical model can reflect the relationship between the noise or sound field estimation at the target spatial position and the parameter information (such as frequency information, amplitude information, phase information, etc.) of the ambient noise picked up by the microphone array (e.g., the first microphone array 320) and the parameters of the microphone array. The parameters of the microphone array may include one or more of the arrangement mode of the microphone array, the spacing between each microphone, the number and position of the microphones in the microphone array, etc. The mathematical model can be obtained through calculation based on the initial mathematical model, the parameters of the microphone array, and the parameter information (such as frequency information, amplitude information, phase information, etc.) of the sound (such as ambient noise) picked up by the microphone array. For example, the initial mathematical model may include the parameters corresponding to the parameters of the microphone array and the parameter information of the ambient noise picked up by the microphone array, as well as the model parameters. Substitute the parameters of the microphone array, the parameter information of the sound picked up by the microphone array, and the initial values of the model parameters into the initial mathematical model to obtain the predicted noise or sound field at the target spatial position. Then compare the predicted noise or sound field with the data (noise and sound field estimation) obtained by the physical microphone set at the target spatial position to adjust the model parameters of the mathematical model. Based on the above adjustment method, through a large amount of data (such as the parameters of the microphone array and the parameter information of the ambient noise picked up by the microphone array), multiple adjustments are made to obtain this mathematical model.

[0127] In some embodiments, the virtual microphone may include a machine learning model. The machine learning model can be obtained through training based on the parameters of the microphone array and the parameter information (such as frequency information, amplitude information, phase information, etc.) of the sound picked up by the microphone array (such as ambient noise). For example, the parameters of the microphone array and the parameter information of the sound picked up by the microphone array are used as training samples to train an initial machine learning model (such as a neural network model) to obtain the machine learning model. Specifically, the parameters of the microphone array and the parameter information of the sound picked up by the microphone array can be input into the initial machine learning model, and a prediction result (such as noise and sound field estimation at the target spatial position) can be obtained. Then, the prediction result is compared with the data (noise and sound field estimation) obtained by the physical microphone set at the target spatial position to adjust the parameters of the initial machine learning model. Based on the above adjustment method, through a large amount of data (such as the parameters of the microphone array and the parameter information of the ambient noise picked up by the microphone array), after multiple iterations, the parameters of the initial machine learning model are optimized until the prediction result of the initial machine learning model is the same as or approximately the same as the data obtained by the physical microphone set at the target spatial position, and the machine learning model is obtained.

[0128] Virtual microphone technology can move the physical microphone away from positions where it is difficult to place the microphone (such as the target spatial position). For example, in order to achieve the purpose of opening the user's ears without blocking the ear canals, the physical microphone cannot be set at the position of the user's ear holes (such as the target spatial position). At this time, the microphone array can be set at a position close to the user's ear and not blocking the ear canal through virtual microphone technology, and then a virtual microphone at the position of the user's ear holes can be constructed through the microphone array. The virtual microphone can use the physical microphone at the first position (such as the first microphone array 320) to predict the sound data (such as amplitude, phase, sound pressure, sound field, etc.) at the second position (such as the target spatial position). In some embodiments, the sound data at the second position (which can also be referred to as a specific position, such as the target spatial position) predicted by the virtual microphone can be adjusted according to the distance between the virtual microphone and the physical microphone (the first microphone array 320), the type of the virtual microphone (such as a mathematical model virtual microphone, a machine learning virtual microphone), etc. For example, the closer the distance between the virtual microphone and the physical microphone, the more accurate the sound data at the second position predicted by the virtual microphone. Another example is that in some specific application scenarios, the sound data at the second position predicted by the machine learning virtual microphone is more accurate than that of the mathematical model virtual microphone. In some embodiments, the position corresponding to the virtual microphone (i.e., the second position, such as the target spatial position) can be near the first microphone array 320 or far from the first microphone array 320.

[0129] In step 1620, the noise and sound field at the target spatial position are estimated based on the virtual microphone. In some embodiments, this step may be performed by the processor 330.

[0130] In some embodiments, when the virtual microphone is a mathematical model, the processor 330 can input the parameter information of the ambient noise picked up by the first microphone array (e.g., the first microphone array 320) (such as frequency information, amplitude information, phase information, etc.) and the parameters of the first microphone array (e.g., the arrangement mode of the first microphone array, the spacing between each microphone, the number of microphones in the first microphone array) into the mathematical model as the parameters of the mathematical model in real time to estimate the noise and sound field at the target spatial position.

[0131] In some embodiments, when the virtual microphone is a machine learning model, the processor 330 can input the parameter information of the ambient noise picked up by the first microphone array (such as frequency information, amplitude information, phase information, etc.) and the parameters of the first microphone array (e.g., the arrangement mode of the first microphone array, the spacing between each microphone, the number of microphones in the first microphone array) into the machine learning model in real time and estimate the noise and sound field at the target spatial position based on the output of the machine learning model.

[0132] It should be noted that the above description of the process 1600 is only for illustration and explanation, and does not limit the scope of application of the present application. For those skilled in the art, various modifications and changes can be made to the process 1600 under the guidance of the present application. For example, step 1620 can be divided into two steps to estimate the noise and sound field at the target spatial position respectively. These modifications and changes are still within the scope of the present application.

[0133] In some embodiments, the speaker 340 outputs a target signal based on the noise reduction signal. After the target signal cancels out the ambient noise, there may still be a part of the sound signal that has not been cancelled out near the user's ear canal. These uncancelled sound signals can be residual ambient noise and / or residual target signals. Therefore, there is still a certain amount of noise at the user's ear canal. Based on this, in some embodiments, Figure 1 the headphone 100 shown, Figures 3 - 12 the headphone 300 shown may further include a second microphone 360. The second microphone 360 may be located in the body part 312 (such as the holding part 3122). The second microphone 360 may be configured to pick up the ambient noise and the target signal.

[0134] In some embodiments, the number of the second microphones 360 may be one or more. When the number of the second microphones 360 is one, the second microphone can be used to pick up the ambient noise and the target signal at the user's ear canal to monitor the sound field at the user's ear canal after the target signal and the ambient noise are cancelled. When the number of the second microphones 360 is multiple, the multiple second microphones can be used to pick up the ambient noise and the target signal at the user's ear canal, and the relevant parameter information of the sound signals picked up by the multiple microphones at the user's ear canal can estimate the noise at the user's ear canal by means of an averaging or weighting algorithm, etc. In some embodiments, when the number of the second microphones 360 is multiple, some of the multiple microphones can be used to pick up the ambient noise and the target signal at the user's ear canal, and the remaining microphones can be used as the microphones in the first microphone array 320. At this time, there is an overlap or intersection between the microphones in the first microphone array 320 and the microphones in the second microphones 360.

[0135] In some embodiments, referring to Figure 10 , the second microphone 360 can be disposed in a second target area, and the second target area can be an area on the holding portion 3122 close to the user's ear canal. In some embodiments, the second target area can be Figure 10 area H in. Area H can be a partial area on the holding portion 3122 close to the user's ear canal. That is to say, the second microphone 360 can be located on the holding portion 3122. For example, area H can be a partial area in the first area 3122A on the side of the holding portion 3122 facing the user's ear. By disposing the second microphone 360 in the second target area H, the second microphone 360 can be located near the user's ear canal and be closer to the user's ear canal than the first microphone array 320, so as to ensure that the sound signals picked up by the second microphone 360 (such as the residual ambient noise, the residual target signal, etc.) are closer to the sound heard by the user. The processor 330 further updates the noise reduction signal according to the sound signals picked up by the second microphone 360, so as to achieve a more ideal noise reduction effect.

[0136] In some embodiments, to ensure that the second microphone 360 can pick up the remaining ambient noise at the user's ear canal more accurately, the position of the second microphone 360 on the holding part 3122 can be adjusted so that the distance between the second microphone 360 and the user's ear canal is within a suitable range. In some embodiments, to ensure that the second microphone 360 can pick up the remaining ambient noise at the user's ear canal more accurately, when the user wears the earphone 300, the distance between the second microphone 360 and the user's ear canal can be less than 10 millimeters. In some embodiments, to ensure that the sound signal picked up by the second microphone 360 is closer to the sound heard by the user, when the user wears the earphone 300, the distance between the second microphone 360 and the user's ear canal can be less than 9 millimeters. In some embodiments, to further improve the active noise cancellation effect of the earphone 300, when the user wears the earphone 300, the distance between the second microphone 360 and the user's ear canal can be less than 8 millimeters. In some embodiments, based on the active noise cancellation effect of the earphone 300 and its comfort in wearing, when the user wears the earphone 300, the distance between the second microphone 360 and the user's ear canal can be less than 7 millimeters.

[0137] In some embodiments, the second microphone 360 needs to pick up the remaining target signal after the target signal output by the speaker 340 through the sound outlet hole 301 is cancelled with the ambient noise. To ensure that the second microphone 360 can pick up the remaining target signal more accurately, the distance between the second microphone 360 and the sound outlet hole 301 can be reasonably set. In some embodiments, to ensure that the second microphone 360 can pick up the remaining target signal more accurately, in the sagittal plane (YZ plane) of the user, the distance between the second microphone 360 and the sound outlet hole 301 along the sagittal axis (Y axis) direction can be less than 10 millimeters. In some embodiments, considering the overall size of the holding part 3122 of the earphone 300 (for example, the length of the holding part 3122 along the Y axis), in the sagittal plane (YZ plane) of the user, the distance between the second microphone 360 and the sound outlet hole 301 along the sagittal axis (Y axis) direction can be less than 9 millimeters. In some embodiments, to ensure the active noise cancellation effect of the earphone 300, in the sagittal plane (YZ plane) of the user, the distance between the second microphone 360 and the sound outlet hole 301 along the sagittal axis (Y axis) direction can be less than 8 millimeters.

[0138] In some embodiments, on the sagittal plane of the user, to ensure that the second microphone 360 can pick up the residual target signal more accurately, the distance between the second microphone 360 and the sound outlet 301 in the direction of the vertical axis (Z-axis) can be 3 millimeters to 6 millimeters. In some embodiments, considering the overall size of the holding portion 3122 of the earphone 300 (for example, the length of the holding portion 3122 in the Z-axis direction), on the sagittal plane of the user, the distance between the second microphone 360 and the sound outlet 301 in the direction of the vertical axis (Z-axis) can be 2.5 millimeters to 5.5 millimeters. In some embodiments, to ensure the active noise reduction effect of the earphone 300, on the sagittal plane of the user, the distance between the second microphone 360 and the sound outlet 301 in the direction of the vertical axis (Z-axis) can be 3 millimeters to 5 millimeters.

[0139] In some embodiments, to ensure the active noise reduction effect of the earphone 300, on the sagittal plane of the user, the distance between the second microphone 360 and the first microphone array 320 in the direction of the vertical axis (Z-axis) can be 2 millimeters to 8 millimeters. In some embodiments, considering the overall size of the holding portion 3122 of the earphone 300 (for example, the length of the holding portion 3122 in the Z-axis direction), on the sagittal plane of the user, the distance between the second microphone 360 and the first microphone array 320 in the direction of the vertical axis (Z-axis) can be 3 millimeters to 7 millimeters. In some embodiments, based on the positional relationship between the second microphone 360 and the sound outlet 301 and the positional relationship between the first microphone array 320 and the sound outlet 301, on the sagittal plane of the user, the distance between the second microphone 360 and the first microphone array 320 in the direction of the vertical axis (Z-axis) can be 4 millimeters to 6 millimeters.

[0140] In some embodiments, to ensure the active noise reduction effect of the earphone 300, on the sagittal plane of the user, the distance between the second microphone 360 and the first microphone array 320 in the direction of the sagittal axis (Y-axis) can be 2 millimeters to 20 millimeters. In some embodiments, considering the overall size of the holding portion 3122 of the earphone 300 (for example, the length of the holding portion 3122 in the Y-axis direction), on the sagittal plane of the user, the distance between the second microphone 360 and the first microphone array 320 in the direction of the sagittal axis (Y-axis) can be 4 millimeters to 18 millimeters. In some embodiments, based on the positional relationship between the second microphone 360 and the sound outlet 301 and the positional relationship between the first microphone array 320 and the sound outlet 301, on the sagittal plane of the user, the distance between the second microphone 360 and the first microphone array 320 in the direction of the sagittal axis (Y-axis) can be 5 millimeters to 15 millimeters.

[0141] In some embodiments, to ensure the active noise reduction effect of the earphone 300, on the cross-section (XY plane) of the user, the distance between the second microphone 360 and the first microphone array 320 in the coronal axis (X-axis) direction can be less than 3 millimeters. In some embodiments, considering the overall size of the holding portion 3122 of the earphone 300 (e.g., the length of the holding portion 3122 in the X-axis direction), on the cross-section (XY plane) of the user, the distance between the second microphone 360 and the first microphone array 320 in the coronal axis (X-axis) direction can be less than 2.5 millimeters. In some embodiments, based on the positional relationship between the second microphone 360 and the sound outlet 301 and the positional relationship between the first microphone array 320 and the sound outlet 301, on the cross-section (XY plane) of the user, the distance between the second microphone 360 and the first microphone array 320 in the coronal axis (X-axis) direction can be less than 2 millimeters. It can be understood that the distance between the second microphone 360 and the first microphone array 320 can be the distance between the second microphone 360 and any microphone in the first microphone array 320.

[0142] In some embodiments, the second microphone 360 is configured to pick up ambient noise and target signals. Further, the processor 330 can update the noise reduction signal based on the sound signal picked up by the second microphone 360, thereby further improving the active noise reduction effect of the earphone 300. For a specific description of using the second microphone 360 to update the noise reduction signal, reference can be made to Figure 17 , and its related description.

[0143] Figure 17 is an exemplary flowchart of updating the noise reduction signal shown in some embodiments of the present application. As Figure 17 shown, the process 1700 may include:

[0144] In step 1710, based on the sound signal picked up by the second microphone 360, the sound field at the user's ear canal is estimated.

[0145] In some embodiments, this step may be executed by the processor 330. In some embodiments, the sound signal picked up by the second microphone 360 includes ambient noise and the target signal output by the speaker 340. In some embodiments, after the ambient noise and the target signal output by the speaker 340 are canceled out, there may still be a part of the sound signal that has not been canceled out near the user's ear canal. These uncanceled sound signals can be residual ambient noise and / or residual target signals. Therefore, there is still a certain amount of noise at the user's ear canal after the ambient noise and the target signal are canceled out. The processor 330 can process the sound signal (e.g., ambient noise, target signal) picked up by the second microphone 360 to obtain parameter information of the sound field at the user's ear canal, such as frequency information, amplitude information, and phase information, etc., so as to realize the estimation of the sound field at the user's ear canal.

[0146] In step 1720, the noise reduction signal is updated according to the sound field at the user's ear canal.

[0147] In some embodiments, step 1720 may be executed by the processor 330. In some embodiments, the processor 330 may adjust the parameter information of the noise reduction signal (e.g., frequency information, amplitude information, and / or phase information) according to the parameter information of the sound field at the user's ear canal obtained in step 1710, so that the amplitude information and frequency information of the updated noise reduction signal are more consistent with the amplitude information and frequency information of the ambient noise at the user's ear canal, and the phase information of the updated noise reduction signal is more consistent with the anti-phase information of the ambient noise at the user's ear canal, thereby enabling the updated noise reduction signal to more accurately cancel the ambient noise.

[0148] It should be noted that the above description of the process 1700 is only for illustration and explanation, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the process 1700 under the guidance of this specification. However, these modifications and changes are still within the scope of this specification. For example, the microphone for picking up the sound field at the user's ear canal is not limited to the second microphone 360, and may also include other microphones, such as the third microphone, the fourth microphone, etc. The relevant parameter information of the sound field at the user's ear canal picked up by multiple microphones can be used to estimate the sound field at the user's ear canal by means of an average or weighted algorithm, etc.

[0149] In some embodiments, in order to more accurately obtain the sound field at the user's ear canal, the second microphone 360 may include a microphone that is closer to the user's ear canal than any microphone in the first microphone array 320. In some embodiments, the sound signal picked up by the first microphone array 320 is ambient noise, and the sound signal picked up by the second microphone 360 is ambient noise and the target signal. In some embodiments, the processor 330 may estimate the sound field at the user's ear canal according to the sound signal picked up by the second microphone 360 to update the noise reduction signal. The second microphone 360 needs to monitor the sound field at the user's ear canal after the noise reduction signal and the ambient noise are canceled. The second microphone 360 including a microphone that is closer to the user's ear canal than any microphone in the first microphone array 320 can more accurately represent the sound signal heard by the user. By estimating the sound field through the second microphone 360 to update the noise reduction signal, the noise reduction effect and the user's auditory experience can be further improved.

[0150] In some embodiments, the earphone 300 may also not include the above-mentioned first microphone array, and only use the second microphone 360 for active noise reduction. At this time, the processor 330 may regard the ambient noise picked up by the second microphone 360 as the noise at the user's ear canal and generate a feedback signal therefrom to adjust the noise reduction signal, so as to cancel or reduce the ambient noise at the user's ear canal. For another example, when the number of the second microphones 360 is multiple, some of the multiple microphones may be used to pick up the ambient noise near the user's ear canal, and the remaining microphones are used to pick up the ambient noise and the target signal at the user's ear canal, so that the processor 330 can update the noise reduction signal according to the sound signal at the user's ear canal after the target signal cancels the ambient noise, further improving the active noise reduction effect of the earphone 300.

[0151] Figure 18 is an exemplary noise reduction flowchart of an earphone shown in some embodiments of the present application. As Figure 18 shown, the process 1800 may include:

[0152] In step 1810, the picked-up ambient noise is divided into multiple frequency bands, and the multiple frequency bands correspond to different frequency ranges.

[0153] In some embodiments, this step may be executed by the processor 330. The ambient noise picked up by the microphone array (such as the first microphone array 320) contains different frequency components. In some embodiments, when processing the ambient noise signal, the processor 330 may divide the ambient noise frequency band into multiple frequency bands, and each frequency band corresponds to a different frequency range. The frequency range corresponding to each frequency band here may be a preset frequency range, for example, 20 - 100 Hz, 100 Hz - 1000 Hz, 3000 Hz - 6000 Hz, 9000 Hz - 20000 Hz, etc.

[0154] In step 1820, based on at least one of the multiple frequency bands, a noise reduction signal corresponding to each of the at least one frequency band is generated.

[0155] In some embodiments, this step may be executed by the processor 330. The processor 330 may analyze the frequency bands into which the ambient noise is divided to obtain parameter information of the ambient noise corresponding to each frequency band (such as frequency information, amplitude information, phase information, etc.). The processor 330 generates a noise reduction signal corresponding to each of at least one of the frequency bands according to the parameter information. For example, in the frequency band of 20 Hz - 100 Hz, the processor 330 may generate a noise reduction signal corresponding to the frequency band of 20 Hz - 100 Hz based on the parameter information of the ambient noise corresponding to the frequency band of 20 Hz - 100 Hz (such as frequency information, amplitude information, phase information, etc.). Further, the speaker 340 outputs a target signal based on the noise reduction signal of the frequency band of 20 Hz - 100 Hz. For example, the speaker 340 may output a target signal with an approximately opposite phase and an approximately equal amplitude to the noise of the frequency band of 20 Hz - 100 Hz to cancel the noise of this frequency band.

[0156] In some embodiments, generating a noise reduction signal corresponding to each of at least one of the multiple frequency bands based on at least one of the multiple frequency bands may include obtaining the sound pressure levels corresponding to the multiple frequency bands, and generating a noise reduction signal corresponding to only some of the frequency bands based on the sound pressure levels corresponding to the multiple frequency bands and the frequency ranges corresponding to the multiple frequency bands. In some embodiments, the sound pressure levels of the ambient noise in different frequency bands picked up by the microphone array (such as the first microphone array 320) may be different. The processor 330 analyzes the frequency bands into which the ambient noise is divided to obtain the sound pressure level corresponding to each frequency band. In some embodiments, considering the structural differences of the open earphone (such as the earphone 300), and the change in the transfer function caused by the different wearing positions of the earphone due to the differences in the ear structures of the users, the earphone 300 may select some of the frequency bands in the ambient noise frequency band for active noise reduction. The processor 330 generates a noise reduction signal corresponding to only some of the frequency bands based on the sound pressure levels and frequency ranges of the multiple frequency bands. For example, when the low-frequency (such as 20 Hz - 100 Hz) noise in the ambient noise is relatively large (such as the sound pressure level is greater than 60 dB), the open earphone may not be able to emit a large enough noise reduction signal to cancel this low-frequency noise. In this case, the processor 330 may generate a noise reduction signal corresponding to only some of the higher-frequency bands (such as 100 Hz - 1000 Hz, 3000 Hz - 6000 Hz) in the ambient noise frequency band. Another example is that the change in the transfer function caused by the different wearing positions of the earphone due to the differences in the ear structures of the users makes it difficult for the open earphone to perform active noise reduction on the ambient noise of high-frequency signals (such as greater than 2000 Hz). In this case, the processor 330 may generate a noise reduction signal corresponding to only some of the lower-frequency bands (such as 20 Hz - 100 Hz) in the ambient noise frequency band.

[0157] It should be noted that the above description of process 1800 is only for illustration and example, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to process 1800 under the guidance of this specification. For example, steps 1810 and 1820 can be combined. Another example is adding other steps to process 1800. However, these modifications and changes are still within the scope of this specification.

[0158] Figure 19 is an exemplary flowchart for estimating the noise of the target spatial position as shown in some embodiments of the present application. As Figure 19 shown, process 1900 may include:

[0159] In step 1910, components associated with the signal picked up by the bone conduction microphone are removed from the picked-up ambient noise to update the ambient noise.

[0160] In some embodiments, this step may be executed by the processor 330. In some embodiments, when the microphone array (e.g., the first microphone array 320) picks up the ambient noise, the user's own speaking voice will also be picked up by the microphone array, that is, the user's own speaking voice is also regarded as part of the ambient noise. In this case, the target signal output by the speaker (e.g., the speaker 340) will cancel out the user's own speaking voice. In some embodiments, in specific scenarios, the user's own speaking voice needs to be retained, such as in scenarios where the user makes a voice call, sends a voice message, etc. In some embodiments, the earphone (e.g., earphone 300) may include a bone conduction microphone. When the user wears the earphone to make a voice call or record voice information, the bone conduction microphone can pick up the voice signal of the user by picking up the vibration signal generated by the facial bones or muscles when the user speaks, and transmit it to the processor 330. The processor 330 obtains the parameter information of the voice signal picked up by the bone conduction microphone, and removes the voice signal components associated with the voice signal picked up by the bone conduction microphone from the ambient noise picked up by the microphone array. The processor 330 updates the ambient noise according to the parameter information of the remaining ambient noise. The updated ambient noise no longer contains the user's own speaking voice signal, that is, the user can hear the user's own speaking voice signal when making a voice call.

[0161] In step 1920, the noise of the target spatial position is estimated based on the updated ambient noise.

[0162] In some embodiments, this step may be executed by the processor 330. Step 1920 may be executed in a manner similar to step 1420, and the relevant description will not be repeated here.

[0163] It should be noted that the above description of process 1900 is only for illustration and example, and does not limit the scope of application of the present application. For those skilled in the art, various modifications and changes can be made to process 1900 under the guidance of the present application. For example, the components associated with the signal picked up by the bone conduction microphone can also be pre-processed, and the signal picked up by the bone conduction microphone can be transmitted to the terminal device as an audio signal. These modifications and changes are still within the scope of the present application.

[0164] In some embodiments, the noise reduction signal can also be updated according to the user's manual input. For example, in some embodiments, due to differences in ear structures or different wearing states of the earphone 300 among different users, the active noise reduction effect of the earphone 300 will be different, resulting in an unsatisfactory auditory experience. At this time, the user can manually adjust the parameter information of the noise reduction signal (such as frequency information, phase information, or amplitude information) according to their own auditory effect, so as to match the wearing position of different users wearing the earphone 300 and improve the active noise reduction performance of the earphone 300. For another example, during the use of the earphone 300 by special users (such as hearing-impaired users or older users), there are differences in hearing ability compared with that of ordinary users, and the noise reduction signal generated by the earphone 300 itself does not match the hearing ability of special users, resulting in a poor auditory experience for special users. In such a case, the special user can manually adjust the frequency information, phase information, or amplitude information of the noise reduction signal according to their own auditory effect, so as to update the noise reduction signal to improve the auditory experience of special users. In some embodiments, the way for the user to manually adjust the noise reduction signal can be to manually adjust it through the key positions on the earphone 300. In some embodiments, key positions for the user to adjust can be provided at any position of the fixing structure 310 of the earphone 300 (such as the side surface of the holding part 3122 facing away from the ear) to adjust the active noise reduction effect of the earphone 300, thereby improving the auditory experience of the user when using the earphone 300. In some embodiments, the way for the user to manually adjust the noise reduction signal can also be to manually input and adjust it through the terminal device. In some embodiments, the sound field at the user's ear canal can be displayed on the earphone 300 or on electronic products such as mobile phones, tablet computers, and computers that are communicatively connected to the earphone 300, and the range of frequency information, amplitude information, or phase information of the recommended noise reduction signal can be fed back to the user. The user can manually input according to the parameter information of the recommended noise reduction signal, and then fine-tune the parameter information according to their own auditory experience.

[0165] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is merely an example and does not constitute a limitation to this application. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are proposed in this application, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this application.

Claims

1. A kind of earphone, characterized in that, Comprising: A fixing structure configured to fix the earphone near the user's ear without blocking the user's ear canal. The fixing structure includes a hook portion and a body portion. When the user wears the earphone, the hook portion is hung between the first side of the user's ear and the head, and the body portion contacts the second side of the ear; A first microphone array located on the body portion and configured to pick up ambient noise; A processor located on the hook portion or the body portion and configured to: Estimate the sound field of a target spatial position by using the first microphone array, where the target spatial position is closer to the user's ear canal than any microphone in the first microphone array, and Generate a noise reduction signal based on the sound field estimation of the target spatial position; and A speaker located on the body portion and configured to output a target signal according to the noise reduction signal, and the target signal is transmitted to the outside of the earphone through a sound outlet hole for reducing the ambient noise.

2. The earphone according to claim 1, wherein, The body portion includes a connecting portion and a holding portion. When the user wears the earphone, the holding portion contacts the second side of the ear, and the connecting portion connects the hook portion and the holding portion.

3. The earphone according to claim 2, wherein The speaker is disposed on the holding portion, and the holding portion is a multi-segment structure to adjust the relative position of the speaker in the overall structure of the earphone.

4. The earphone according to claim 2, wherein The side of the holding portion facing the ear is provided with the sound outlet hole so that the target signal output by the speaker is transmitted to the ear through the sound outlet hole.

5. The earphone according to claim 4, wherein The side of the holding portion facing the ear includes a first region and a second region. The first region is provided with a sound outlet hole. The second region is farther from the connecting portion than the first region and protrudes toward the ear compared to the first region to allow the sound outlet hole to be spaced from the ear in the wearing state.

6. The earphone according to claim 5, characterized in that, When the user wears the earphone, the distance between the sound outlet hole and the user's ear canal is less than 10 millimeters.

7. The earphone according to claim 2, characterized in that, The holding portion is provided with a pressure relief hole on the side along the vertical axis and near the user's head top, and the pressure relief hole is farther from the user's ear canal than the sound outlet hole.

8. The earphone according to claim 7, wherein, When the user wears the earphone, the distance between the pressure relief hole and the user's ear canal is 5 millimeters to 15 millimeters.

9. The earphone according to claim 7, characterized in that, The included angle between the line connecting the pressure relief hole and the sound outlet hole and the thickness direction of the holding portion is 0° to 50°.

10. The earphone according to claim 7, wherein The pressure relief hole and the sound outlet hole form an acoustic dipole, and the first microphone array is disposed in a first target area, and the first target area is the acoustic zero position of the dipole radiation sound field.

11. The earphone according to claim 7, characterized in that, The first microphone array is located on the connecting portion.

12. The earphone according to claim 7, wherein, The line connecting the first microphone array and the sound outlet hole and the line connecting the sound outlet hole and the pressure relief hole have a first included angle, and the line connecting the first microphone array and the pressure relief hole and the line connecting the sound outlet hole and the pressure relief hole have a second included angle, and the difference between the first included angle and the second included angle is not greater than 30°.

13. The earphone according to claim 7, characterized in that, There is a first distance between the first microphone array and the sound outlet hole, and a second distance between the first microphone array and the pressure relief hole, and the difference between the first distance and the second distance is not greater than 6 millimeters.

14. The earphone according to claim 1, wherein The generating of the noise reduction signal based on the sound field estimation of the target spatial position includes: estimating the noise at the target spatial position based on the picked-up ambient noise; and generating the noise reduction signal based on the noise at the target spatial position and the sound field estimation of the target spatial position.

15. The earphone according to claim 1, wherein The estimating of the sound field of the target spatial position by using the first microphone array includes: constructing a virtual microphone based on the first microphone array, the virtual microphone including a mathematical model or a machine learning model for representing the audio data collected by the microphone if the microphone is included at the target spatial position; and estimating the sound field of the target spatial position based on the virtual microphone.

16. The earphone according to claim 1, wherein The earphone includes a second microphone located in the body part, and the second microphone is configured to pick up the ambient noise and the target signal; and The processor is configured to update the target signal based on the sound signal picked up by the second microphone.

17. The earphone according to claim 16, wherein, The second microphone includes at least one microphone that is closer to the user's ear canal than any microphone in the first microphone array.

18. The earphone according to claim 16, characterized in that, The second microphone is disposed in a second target area, and the second target area is an area on the body part close to the user's ear canal.

19. The earphone according to claim 17, characterized in that, On the sagittal plane of the user, the distance between the second microphone and the sound outlet hole in the sagittal axis direction is less than 10 millimeters, and the distance between the second microphone and the sound outlet hole in the vertical axis direction is 2 millimeters to 5 millimeters.

Citation Information

Patent Citations

  • Acoustically open headphone with active noise reduction

    CN109565626A

  • Active noise reduction method and device, electronic equipment and chip

    CN111935589A

  • Wireless earphone and ear hook assembly thereof

    CN209767789U

  • Open Audio Device

    US20210076118A1